Compare commits

...
7 Commits
Author SHA1 Message Date
JesseMarkowitzandClaude Opus 5 db7b309e3d v1.1 closeout: accept integrated release validation
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
Release validation of candidate 87a4032, not a work package. No product code
changed, no requirement or acceptance test changed, no schema or bundle format
changed, and nothing is tagged or merged by it.

V1.1 RELEASE VALIDATION: PASS

What was run, on this candidate:

- v1 contract: 82 REQUIRED tests — 81 PASS, H09 NOT APPLICABLE, 0 waived,
  0 weakened, 0 reclassified.
- Suites: backend 1,723 passed / 17 skipped / 0 failed / 0 xfailed; frontend
  175 passed; lint 0 errors (15 documented warnings); production build clean.
- Docker: docker build --no-cache; the image's SPA is file-for-file identical
  to the local build (16 files, same combined sha256).
- Offline: 23/23 against the candidate image with no network and a fresh volume.
- Browser: 101 passed / 0 failed / 0 skipped (M11 38, WP-C 53, WP-E 10) over
  trusted-LAN HTTPS with a private CA; every narrator turn "fits".
- Long run: 102 accepted turns at a verified 16,384 window with memory on,
  3 process restarts, M01-M04 pass, 0 post-turn failures, 0 database locks.
- A1: every turn "fits"; the ten largest prompts re-counted against the server
  keep the documented reserve, smallest margin 879 tokens against v1's 23-42.
- A2: release-gate leak count 0 across 105 stored replies.
- Identity: 0 signals and 0 stored protocol shapes, with memory on; the
  scripted detector still fires on an injected defect.
- Recovery: 16/16 on the long run's own bundle, into a database and directory
  that never existed.
- Upgrade: a campaign built and played by the v1.0.0 application compares
  identical on all 15 census fields, schema parity at user_version 94, and both
  bundle directions import.
- Release smoke: 15/15 from the shipped image — loopback only, private CA
  verified, public endpoint refused, a real turn, restart, persistence, and
  Firefox rendering the reopened campaign.

Carried residuals, stated rather than summarised away:

- WP-B: deterministic independent-memory recovery PASS; reference-model
  independent-memory recovery FAIL at memory creation — the owner-accepted
  limitation, unchanged and not a new regression.
- The mid-reply instruction echo A2's trailing cleanup does not remove is still
  reproducible on the stored WP-B.1 fixture (1 of 105), and did not recur in
  release evidence.
- The doubled full stop in the memory-search scene text.
- K1 ("Correct" on an Important Facts row is refused) is classified v1.2
  backlog, reproduced and not fixed during validation.

Three harness corrections were made during validation — the identity diagnostic
did not enable memory, the smoke test needed hostname resolution inside the
container, and the first upgrade campaign was too short to write memories. All
harness-only; each corrected harness repeated its own check, and no product
evidence became stale.

Docs: README, V1.1-PLAN, planning/README and VERSION now say v1.0.0 remains the
released version, that v1.1 is implemented and validated, and that no v1.1.0 tag
exists. WP-E's report records OWNER SCREENSHOT APPROVAL: APPROVED, sourced to
the owner's brief. New harness tools: v11_upgrade_check.py, v11_release_smoke.py.

Still the owner's to do: sign the release commit, update main, tag v1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-16 07:12:23 -04:00
JesseMarkowitz 87a40326a2 v1.1: harden recovery and control boundaries
WP-D and WP-E complete the planned v1.1 implementation packages.

WP-D — recovery honesty:
- backups verify the completed copy with PRAGMA integrity_check
- corruption missed by quick_check is detected by the full check
- existing good backups remain protected
- oversized exports are still delivered but declare whether this version can
  import them, while the 20 MB import limit remains unchanged
- backup was exercised through the real browser UI on both the normal campaign
  database and a campaign-shaped database over 100 MB

WP-E — control-boundary contrast:
- interactive control boundaries meet the WCAG 1.4.11 3:1 target
- the contrast audit is now a failing gate rather than an advisory
- rendered browser measurements pass for the composer, controls, tabs and nav
- text contrast and focus visibility remain intact
- owner reviewed and approved the before/after screenshots

Reports:
- planning/reports/v1.1/V1.1-WP-D-REPORT.md
- planning/reports/v1.1/V1.1-WP-E-REPORT.md

All planned v1.1 work packages A-E are now complete. Release validation has not
yet begun.
2026-09-16 05:37:13 -04:00
JesseMarkowitzandClaude Opus 5 59b5ebc2d8 v1.1 WP-C: browser release coverage
Drives in a real browser the reader workflows v1 proved only through the API
or the component suite, including an export that leaves the browser as a
file. Final run: 91 checks (the 38 existing M11 checks plus 53 new), 0 failed,
0 skipped, on the production build over trusted-LAN HTTPS.

- tools/m11_browser.py: scenarios for Retry and takes, Save Point create /
  restore / Redo, state correction (accepted, and a refused correction with
  its reason), narration length reaching each turn's prompt, failed
  generation (an unserved model blocked up front; a listed model that cannot
  narrate failing in the open) and recovery, and export download from the
  library and from campaign settings, imported into a fresh application.
  Rows are tagged M11 / WP-C and counted separately; --only for development.
  The M11 checks now wait on conditions instead of sleeping.
- tools/m11_webdriver.py: Firefox download preferences, a $HOME-only
  download folder, a download wait that ignores partial, empty, pre-existing
  and still-growing files, centred real clicks, tabs, and condition waits.
- tests/test_v11_c_browser_helpers.py: the download wait, prefs and $HOME
  guard, without a browser.
- frontend: a correction the story refused was presented as "Generation
  failed" with a Retry offer and a typed-input claim. It is now "That
  correction was not applied", not retryable, with the reason kept
  (errors.js, FailureNotice.jsx; 3 regression tests).
- DEVELOPMENT.md: the harness command, download profile and $HOME rule,
  what counts as a finished download, and the no-sleep rule.
- docs: V1.1-PLAN, VERSION v4.4, planning README,
  reports/v1.1/V1.1-WP-C-REPORT.md.

Open for the owner: "Correct" on an Important Facts row is always refused
(K1), and a partly refused correction is not reachable from the reader UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 18:29:45 -04:00
JesseMarkowitzandClaude Opus 5 0c1ba836ba v1.1 WP-B.2: independent long-term memory retention
Corrects the memory mechanisms WP-B.1 diagnosed, one at a time, each
verified before the next. Accepted by the owner with a documented
reference-model limitation. No schema, bundle format, setting default,
lineage, authority or protocol-cleanup change.

- B2.1 ranking: the retrieval query is the player's input plus a bounded
  scene context (state scene + end of the newest narration), embedded in
  one call. final = semantic (0.6 input / 0.4 context) + 0.15 x lexical,
  where lexical is a rarity-weighted share of the input's words, computed
  per turn over the candidates with no index. Scores and the query are
  recorded per used memory; pins and redundancy suppression unchanged.
- B2.2 coverage-aware eviction (memorybank.eviction_order): the earliest
  and newest memories are kept, the smallest coverage hole goes first,
  least-recently-used breaks ties and remains the fallback. Bounded; pins
  never evicted; frozen-bank protection kept; reads no text or vectors.
- B2.3 bounded memory creation: a block longer than 2,000 tokens is shown
  to the summariser as head + tail with an omission marker, inside the
  same budget; shorter blocks unchanged; the marker is never stored.
- The memory summariser prompt is unchanged from v1.0.0. A B2.4 prompt
  experiment was measured on the reference model, showed no reliable
  improvement for the target failure (0/5 under both prompts, with new
  "Memory:"-prefix, second-person and length regressions), and was
  reverted. memorybank.memory_user_prompt is kept as a behaviour-neutral
  helper.
- tools/memory_fidelity.py (diagnostic only): genre-neutral fixtures plus
  the failed block, a deterministic fidelity checker, and a real-model
  shipped-vs-experiment measurement.
- tools/memory_diagnostic.py: ranking replica uses production scoring;
  ranking_crowded, ranking_context_dependent and independent_full
  fixtures; per-turn isolation and provenance.
- tests: B.1's two strict xfails are now ordinary passes; ranking,
  eviction and excerpt tests; summariser acceptance tests kept apart from
  diagnostic-measurement tests.
- DEVELOPMENT.md: the GPU-host kernel/Ollama watch used `-k -u ollama`,
  which matches nothing; now the OR form.
- docs: CONTEXT-AND-MEMORY 15/18/20/21 as shipped, V1.1-PLAN (status and
  release criteria 12-13), planning README, VERSION v4.3,
  reports/v1.1/V1.1-WP-B2-REPORT.md.

Deterministic independent-memory recovery: PASS (independent_full fails
on v1.0.0 at creation and returns recovered_through_memory_independent
here). Reference-model independent recovery: FAILED on the
precondition-valid attempt, at memory creation: the summariser omitted a
player-established fact from a block it received whole. Accepted as a
documented v1.1 residual and carried into the release gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-15 11:21:53 -04:00
JesseMarkowitzandClaude Opus 5 beb17ada10 v1.1 WP-B.1: diagnose independent long-term memory retention
Diagnostic only; no memory behaviour changes.

- tools/memory_diagnostic.py: planted-fact isolation checks, the four-stage
  diagnosis (created / retained / ranked / injected) with a verdict, a
  production-ranking replica, deterministic summariser/embedder/narrator
  stubs and seven scenarios (default, past capacity, pinned, low top_k,
  long-block early/late, lineage control)
- tools/v11_b1_memory.py: CLI for the scenarios and for diagnosing a copy of
  a finished real campaign
- tools/m11_long_run.py: opt-in --independent-fact mode with per-turn
  isolation tracking and the recovered_through_memory_independent verdict;
  M04 verdicts unchanged
- tests: diagnostic stages, eviction, creation window, ranking, lineage and
  authority controls; two strict xfails record the diagnosed retention and
  creation defects for WP-B.2 to flip
- planning/reports/v1.1/V1.1-WP-B1-REPORT.md

First failing stage: ranking (real model); retention past capacity and
creation for early facts in long blocks (deterministic, same on v1.0.0).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 20:50:05 -04:00
JesseMarkowitzandClaude Opus 5 d63804f22e v1.1: harden context window and narrator protocol boundary
WP-A1 and WP-A2, implemented in sequence, plus the corrective work the owner
asked for at review. Reported in
planning/reports/v1.1/V1.1-WP-A1-A2-REPORT.md (corrective addendum §R).
Planning package v4.2.

WP-A1: context-window safety reserve
- The prompt leaves max(256, ceil(5% of the effective window)) tokens free
  beside the reply. That is 256 at 4,096 and 820 at 16,384. The value is fixed,
  not a setting, and not calibrated per model.
- M6's 64-token margin is gone. Separators and the chat hint are priced
  exactly; tokenizer drift is the reserve's job.
- Protected context that cannot fit raises ContextOverflow before the model
  is called.
- Streams set stream_options.include_usage. Measured on Ollama 0.33, a stream
  sent no usage without it.
- Each sent turn records fits, exceeded, truncation_suspected or unknown.
  The status is returned on the done event, logged when bad, and shown in the
  context inspector. The turn is always kept.
- Accounting is per-attempt data (attempts.ATTEMPT_KEYS).
- Corrective: a cold model is loaded before its turn is built. When the
  window is unverified but the server answered, contextwindow.ensure_window
  makes one bounded POST /api/generate naming only the model. It sends no
  prompt, generates nothing and writes nothing. It then probes again, and the
  turn is built to that answer. If the load fails, or the window is still
  unknown, the turn falls back to the old behaviour.
- Real host, 4,096 window:
  - v1 cold turn: sent 13,875, the server read 2,050.
  - Same turn after the correction: the window was verified, 3,082 sent,
    3,097 read, fits, 499 tokens left beside the reply.
  - Verified turns elsewhere left 275-2,297 tokens against v1's 23-42.

WP-A2: protocol echo and genre-neutral state prompting
- The vocabulary is shown as the JSON object the model sends, not as
  name(field, ...). This costs 121 tokens.
- The example uses character-1, item-1 and location-1.
- The extractor removes shapes anchored to application-owned text:
  - a vocabulary call line;
  - an echoed length hint;
  - the renderer's scene line left last;
  - an empty fence opener.
- Corrective R5: the echoed continue hint is recognised by its own sentence
  ("Output only story text"). A Hard-limit-opened bracket is removed only
  directly above an echo already cut from the same reply.
- Replay of all 518 real v1 replies: 9 changed, 0 flagged, and no story prose
  removed. That is unchanged by R5.
- Replay of 64 v1.1 replies: 3 changed, 0 flagged. The depth-16 instruction
  tail is removed.
- Identity diagnostic after the correction:
  - 0 identity signals;
  - 0 prompt example identifiers proposed;
  - 0/10 stored turns with protocol or instruction shapes.
- 50-turn run: 51 accepted, 0 of 54 stored turns carry protocol.
- SPECS, render.py and validate.py are identical to v1.0.0.

Compatibility: a real v1.0.0 database reads identically on v1.0.0 and v1.1,
field for field, with schema and user_version 94 unchanged. Undo, redo, Save
Point restore, export and import all work on it. There is no schema,
migration or bundle-format change.

Verification: the backend suite passes 1,534 with 17 skipped and 0 failed.
The frontend passes 165/165, and lint and the build are clean. The offline
container and the browser regression were re-run on this tree (see §R.3).

One test was re-calibrated, not weakened: test_history_block_trim's prefix
test had assumed which turn holds the floor at a 2,048 budget.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 16:35:05 -04:00
JesseMarkowitzandClaude Opus 5 ac465ed867 Planning v4.1: record the v1.0.0 release, and plan v1.1
Documentation only. No product code, requirement, acceptance test or
schema changes.

Post-release correction. v4.0 was written before the closeout commit was
signed (432f041), main was fast-forwarded to it, and the signed v1.0.0 tag
was pushed. Current-state wording now says so in README.md,
planning/README.md, BUILD-MILESTONES.md and VERSION.md. BUILD-MILESTONES.md's
header had been stale since M8. The M11 report is not edited: its §T
records the state at closeout.

v1.1 plan. planning/V1.1-PLAN.md triages the post-v1 backlog and the other
recorded v1 residual risks, and orders them into work packages, not
milestones:
- A1: a context-window safety reserve, plus reporting a turn the server
  truncated
- A2: removing protocol echoes from stored narration, and a genre-neutral
  state rule
- B: long-term memory retention that holds without help from state
- C: browser coverage of Retry, Save Points, correction, length, failure
  and export download
- D: integrity_check on backups, and a warning when an export exceeds the
  import limit
- E: WCAG 1.4.11 control-boundary contrast

Scheduled backups, the import limit, identity detectors and duplication
suppression move to v1.2; media adapters are future work. The first brief
to write is A1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VvegagkhuCZoFPdv4M1egY
2026-09-14 08:19:31 -04:00
65 changed files with 16464 additions and 367 deletions
+106 -5
View File
@@ -408,9 +408,17 @@ AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
# What that campaign is worth on a machine that has never seen it (I01-I07).
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
# The browser release regression and the accessibility measurements. Needs
# `frontend/dist` built and geckodriver on PATH.
.venv/bin/python -m tools.m11_browser --out "$HOME/m11-evidence/browser"
# The browser release regression and the accessibility measurements: M11's 38
# checks plus v1.1 WP-C's reader workflows (Retry, Save Points, state correction,
# narration length, failed generation, export download). Needs `frontend/dist`
# built, geckodriver on PATH, and --out under $HOME (the downloads land inside
# it). Release evidence needs the narrator over trusted-LAN HTTPS.
AIDND_TEST_ENDPOINT=https://... AIDND_TEST_MODEL=qwen2.5:3b-instruct \
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/browser"
# Without a narrator (a partial smoke run, not evidence), or one scenario while
# developing (--only takes: shell, history, markdown, hidden, context, csp, a11y,
# retry, savepoint, state, length, failure, export).
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/smoke" --no-narrator
# A container with no network at all: the offline run and the packaging path.
.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline"
@@ -473,8 +481,11 @@ nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,p
# The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling
nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/gpu-dmon-$(date +%F-%H%M).log"
# Kernel and Ollama messages, live
journalctl -f -k -u ollama | tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log"
# Kernel and Ollama messages, live. The `+` is an OR: `journalctl -k -u ollama`
# asks for messages that are both kernel messages and the ollama unit's, which
# is none, and writes an empty log.
journalctl -f -o short-iso _TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service \
| tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log"
```
If the GPU faults, find the moment and then read what the card was doing just
@@ -496,6 +507,33 @@ It exists so browser evidence needs no Selenium in the dependency surface, and
it documents the one environment quirk that matters here: a snap Firefox will
not open a file the driver names under `/tmp`, but will under `$HOME`.
**Downloads in the browser harness (v1.1 WP-C).** The export checks click the
real Export controls and wait for the file on disk, so the browser has to save
without asking. `m11_webdriver.firefox_download_prefs` gives the WebDriver
session a profile that does that:
- `browser.download.folderList` 2, `browser.download.dir` the run's
`downloads/` folder, `browser.download.useDownloadDir` true;
- no "always ask", and `application/json` saved to disk.
It works on the snap Firefox this machine has (155.0.1, geckodriver 0.37.1), and
no separate Firefox is needed. The same sandbox rule applies as for opening
files: the download folder must be under `$HOME`, and the harness refuses one
that is not.
A download counts as finished only when all of these hold at once
(`m11_webdriver.wait_for_download`):
- a new name has appeared;
- no `*.part` file is left;
- the file is more than zero bytes;
- its size is the same across consecutive polls.
The toast that says "Campaign exported." is not evidence.
**Waiting.** Nothing in the harness sleeps before an assertion. Every wait is on
something the page, the browser or the filesystem shows. A condition that
already holds before the action it waits for does not count as waiting for that
action; the harness defects found in M8, M11 and WP-C were all of that shape.
## Backing up, and getting a campaign back
There are two recovery tools and they answer different questions. Using the
@@ -510,6 +548,40 @@ together.
| Taken from | Export, on a campaign | Settings → *Back up everything on this machine* |
| Restored by | Import campaign, on the library screen | replacing the database file, below |
### How large an export can get
The importer accepts a request body up to **20 MB**
(`backend/app/limits.py`, `MAX_IMPORT_BODY_BYTES`), and v1.1 does not change it.
What that means for a campaign, measured rather than guessed:
- the M11 evidence campaign came to roughly **13 kB per action** in its bundle;
- M9's conservative estimate from that figure is about **279 turns** before a
bundle approaches the limit.
Both are measurements of particular campaigns, **not a turn limit**. What a
campaign actually weighs depends on how long its turns are, how much imported
knowledge travels with it, and how many attempts each turn kept. A campaign of
400 short turns can be well inside the limit; one of 200 long ones with a large
library may not be.
**v1.1 (WP-D) makes the individual case visible.** Every export reports its own
serialised size and whether this version could import it back:
```text
X-Export-Bytes the bundle's size, as the importer would weigh it
X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
X-Importable-By-This-Version true / false
X-Export-Warning present only when it is false
```
The export always succeeds and the file is always delivered — it is complete and
undamaged; what it exceeds is this version's import ceiling. Both Export
controls show the warning when there is one. The size compared is the compact
serialisation the browser would POST back, which is smaller than the
pretty-printed file on disk.
Raising the limit, or streaming an import past it, is deferred to v1.2.
### Exporting and importing a campaign
Export is on each campaign in the library, and in the campaign's own Settings
@@ -620,6 +692,35 @@ campaign gets less history than the setting asks for, which is a visible,
explicable loss rather than a silent one, and Settings' **Test connection**
reports the window it found or says plainly that it could not check.
**It also keeps a margin, and checks the server's own count (v1.1).** The
application counts tokens with `cl100k_base`, and your model counts them with
its own tokenizer. The two disagree slightly, so the prompt is built to leave
`max(256, 5% of the window)` tokens free on top of the reply: 256 at 4,096, and
820 at 16,384. After each turn the server's reported prompt-token count is
compared with what was sent. The context inspector shows the result for any
past turn:
- **The server read the whole prompt:** the ordinary case.
- **The server did not say how much it read:** the server reported no usage.
Nothing is wrong, and nothing is confirmed either.
- **The server may have cut the start of the prompt:** it read far fewer tokens
than were sent. Ollama does this, silently, to a prompt larger than the window
the model was loaded with. The turn is kept. Check the window with the
commands above.
- **The prompt was larger than the server allowed for:** its count and the reply
together exceed the window. The reply may have been cut short. The turn is
kept.
The last two also appear in the server log as a warning.
**A model that is not loaded yet is loaded first.** Before a turn, if the
application cannot read the window because your model isn't in memory, it asks
the same Ollama to load it once. That is a `POST /api/generate` naming only the
model, which generates no text. It then reads the window again, so the first
turn of a session is built to the window the model really has rather than to
your setting. If loading fails, or the window still can't be read, the turn goes
ahead exactly as before, unverified, and the check above still applies.
That does not make the window *bigger*, and the rest of this section is still
how you do that.
+28 -8
View File
@@ -80,6 +80,14 @@ that isn't the live one starts a new branch.
in settings lets you state the window so the prompt is still capped. A window the server
itself reported always wins over that, and a declared one is never reported as verified.
Since v1.1 the prompt also stops short of that window on purpose. It leaves
`max(256, 5% of the window)` tokens free, because your model counts tokens differently from
the application, and the v1 evidence came within 23 tokens of the edge. After each turn,
the server's own count of what it read is compared with what was sent. A turn the server
appears to have truncated is kept, flagged and shown in the context inspector, not left to
pass silently. A model that isn't loaded yet, and so cannot report its window, is loaded
once before the turn is built, so the first turn of a session gets the real window too.
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
card used to arrive in front of it as a world fact with no class, no visibility, no source and
@@ -179,9 +187,9 @@ that isn't the live one starts a new branch.
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
removed rather than left standing as a picture of a product that no longer exists. The M4
closeout drove the real application in a real browser, so the screens exist and work; taking
presentable screenshots of them is a job for the UI pass in M8.
removed rather than left standing as a picture of a product that no longer exists. The
screens exist and are driven in a real browser by the release harness
(`backend/tools/m11_browser.py`). Presentable screenshots of them have not been taken.
## Quick start
@@ -285,7 +293,7 @@ player input
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (94 and counting)
├─ endpoints.py the inference-endpoint address policy
├─ contextwindow.py what the server will actually accept, and the cap
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
@@ -310,7 +318,8 @@ development, Vite proxies `/api` to FastAPI.
## Tests
1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with
1,524 backend tests (1,507 run everywhere, 17 need a real local model and skip without one): unit
tests plus full HTTP integration through the real turn engine, with
the model provider mocked. They run with no route to the Internet, which is a requirement
rather than a convenience — an offline claim proved on a machine that has been online once
proves nothing. A further handful need a real local model and skip without one; they exist
@@ -346,9 +355,20 @@ most interesting engineering in the repo.
## Repo notes
- **Status:** milestones M1-M11 are complete. The v1 release gate passed on
2026-09-14 (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
§T). Passing the gate is not a release: there is no `v1.0.0` tag yet.
- **Status:** **v1.0.0 remains the released version.** Milestones M1-M11 are
complete, and the v1 release gate passed (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
§T). The signed tag `v1.0.0` and `main` both point at the signed release
commit `432f041`.
**v1.1 is implemented and validated, but not yet released.** All six work
packages (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) are complete and accepted on
the `v1.1-development` branch, and integrated release validation passed on
candidate `87a4032` — see
[`planning/reports/v1.1/V1.1-RELEASE-REPORT.md`](planning/reports/v1.1/V1.1-RELEASE-REPORT.md).
WP-B ships with a documented reference-model memory limitation, recorded in
that report. **No `v1.1.0` tag exists and `main` is unchanged**; the release
commit, `main` and the tag are the owner's to make. The plan is
[`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md).
- `planning/` is this fork's own package: the product specification, the architecture
decisions, the milestone plan, the acceptance contract, and a review report for every
+8 -1
View File
@@ -46,7 +46,14 @@ from .narrative import model as narrative_model
# token accounting. Each attempt is its own API call, and a retry is the call
# most likely to read the prompt back out of cache. Everything else in a snapshot
# is the prompt, which is assembled once per turn.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage")
#
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
# for *that* call with the turn's estimate. Left out of this tuple, it was
# treated as part of the shared prompt, so moving the live flag handed the
# superseded attempt's accounting to the new live one and threw the new one's
# away. Found by the A2 long run: two retries and one take selection left three
# attempts reporting no accounting, or another attempt's.
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
# ------------------------------------------------------------------ reading
+20 -9
View File
@@ -39,8 +39,9 @@ turn is blocked.
never leaves a half-written file wearing a backup's name. `os.replace` is
atomic on the same filesystem, which is why the temporary sits in the
destination's own directory rather than in `/tmp`.
3. `PRAGMA quick_check` runs against the finished copy, opened as its own
database, before it is renamed. A backup nobody verified is a belief.
3. `PRAGMA integrity_check` runs against the finished copy, opened as its own
database, before it is renamed. A backup nobody verified is a belief. v1.1
WP-D made this the full check rather than `quick_check`; see `_verify`.
4. An existing file is never overwritten. Each run writes a new name stamped
with the time, so yesterday's backup survives today's mistake — which is most
of what a backup is for.
@@ -189,18 +190,28 @@ def _copy(source_path: Path, working: Path) -> int:
def _verify(working: Path) -> str:
"""Runs `PRAGMA quick_check` against the finished copy.
"""Runs `PRAGMA integrity_check` against the finished copy.
Opened as its own connection, so what is checked is the file on disk rather
than any page cache the copy left behind. `quick_check` rather than
`integrity_check` because it does the structural work — every page reachable,
every record readable — without the full index cross-check, which on a large
database is minutes rather than moments. A backup nobody verified is a
belief; a backup verified slowly enough that nobody takes one is worse.
than any page cache the copy left behind.
**v1.1 WP-D: the full check, not `quick_check`.** M9 chose `quick_check` for
its speed, on the argument that a backup verified slowly enough that nobody
takes one is worse than a fast one. The measurements say the trade was not
needed here: `quick_check` omits the cross-check between a table and its
indexes, and that is a real class of damage it reports as `ok`. A copy whose
index disagrees with its table restores into a database that answers queries
with rows that are not there — the failure a backup exists to prevent.
The cost is small at the sizes this application produces: on the 100-turn
evidence campaign both checks are a few milliseconds, and on a synthetic
database two orders of magnitude larger the difference is still short of a
second (WP-D report §E). A backup nobody verified is a belief; this is the
check that makes it a fact.
"""
connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True)
try:
rows = connection.execute("PRAGMA quick_check").fetchall()
rows = connection.execute("PRAGMA integrity_check").fetchall()
finally:
connection.close()
result = ", ".join(str(row[0]) for row in rows) if rows else "no result"
+51 -20
View File
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
from .. import contextwindow, derived, models, narrative, summaries, worldstate
from ..knowledge import inject as knowledge_inject
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
from ..knowledge import records as knowledge_records
from . import encoding, history
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
# Built from the table vendored in `encoding.py`, not fetched: the upstream
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
# is called on every turn.
# M6: added to the configured reply budget when reserving output space. It
# absorbs the section separators added after budgeting and the drift between
# this tokenizer and the serving model's. Fixed rather than proportional: what
# it covers does not grow with the size of the budget.
OUTPUT_SAFETY_MARGIN = 64
#
# v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
# budget to absorb two unrelated things, and v1.1 separates them:
#
# * **Text the application adds after pricing.** The separators between
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
# request and nothing counted. That is not drift, it is our own text, so it is
# now priced exactly (`transport` below).
# * **The drift between this tokenizer and the narrator's.** That is what the
# 64 tokens were really for, and the v1 evidence showed it was too small. It
# is now `contextwindow.safety_reserve`, sized to the window.
#
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
#: author's note, recent history, summary, lore, memories, state, front memory,
#: length hint, refusals, reminder. Knowledge and history rows price their own.
STORY_SECTION_SLOTS = 11
class ContextOverflow(RuntimeError):
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
words = min(words, band_ceiling)
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
tail = (
" Finish the narration and append the state block well inside the limit."
" " + narrative.extract.LENGTH_HINT_TAIL
)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
tail = " Finish the narration and append the state block well inside the limit."
tail = " " + narrative.extract.LENGTH_HINT_TAIL
# State the number as a ceiling, never as a budget. In measurements, the
# wording "keep this turn under about N words" read to the model as a target
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
# Both numbers are bounds, and the wording is deliberately asymmetric. The
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
# so a terse model reading the same clause stops at the floor rather than at
# forty words.
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
@@ -625,12 +637,19 @@ def build_context(
# truncated turn on a model whose window is the budget
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
#
# The margin covers what is added after this arithmetic — the separators
# between sections, and the difference between our tokenizer's count and the
# serving model's. It is small and fixed rather than proportional, because
# what it absorbs does not scale with the budget.
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
protected = reserved + output_reserve
# v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
# application adds after pricing — separators, and the chat hint the
# provider appends — is counted as `transport`. What neither can know, the
# narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
# which is sized to the window and taken before any history is chosen.
output_reserve = max(0, settings.max_output_tokens)
separator_tokens = count_tokens(SEPARATOR)
transport = (
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
+ count_tokens(CHAT_CONTINUE_HINT)
)
safety = contextwindow.safety_reserve(budget)
protected = reserved + transport + output_reserve + safety
if protected >= budget:
# Failing here is the point. The alternative — carrying on with a token
# or two of history — builds a prompt that is known to overflow, and
@@ -638,8 +657,9 @@ def build_context(
# gracefully if protected context alone is too large."
raise ContextOverflow(
f"The protected context needs {protected} tokens "
f"({reserved} of prompt plus {output_reserve} reserved for the "
f"reply) but the context budget is {budget}. "
f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
f"reserved for the reply and a {safety}-token safety margin) but the "
f"context budget is {budget}. "
+ (
"That budget is what this server was found to accept, so raising "
"the setting alone will not help — load the model with a larger "
@@ -827,6 +847,7 @@ def build_context(
story_text = SEPARATOR.join(s.text for s in story_sections)
all_sections = [s for s in system_sections if s.text] + story_sections
total_tokens = count_tokens(system_text) + count_tokens(story_text)
report = {
"sections": [
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
@@ -837,13 +858,23 @@ def build_context(
# what the history was actually allowed to spend after everything
# protected was subtracted.
"tokens": {
"total": count_tokens(system_text) + count_tokens(story_text),
"total": total_tokens,
"budget": budget,
"configured_budget": settings.context_token_budget,
"output_reserve": output_reserve,
"protected": reserved,
"available_for_history": available,
"history_spent": spent,
# v1.1 WP-A1. `transport` is the formatting priced in above;
# `estimate` is what this application believes it actually sent,
# the assembled text plus what the provider adds to it, and is what
# the server's own count is compared against after the reply.
"transport": transport,
"safety_reserve": safety,
"estimate": total_tokens + (
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
else separator_tokens
),
},
# M11: what the server was found to accept, and how. `verified` false
# means nobody could check — the prompt was built to the configured
+208
View File
@@ -86,6 +86,7 @@ one.
from __future__ import annotations
import logging
import math
import re
import time
from dataclasses import dataclass
@@ -132,6 +133,11 @@ class Window:
model_max: int | None = None
#: Why the window is unknown, or how it was found. Shown to the user.
detail: str = ""
#: v1.1: the server answered a discovery request at all, whatever it said.
#: A server that answered but could not report a window may simply not have
#: the model loaded yet, which `ensure_window` can fix; one that did not
#: answer cannot be helped by asking it to load anything.
reachable: bool = False
@property
def verified(self) -> bool:
@@ -180,6 +186,116 @@ def effective_budget(configured: int, window: Window | int | None) -> int:
return min(configured, tokens)
#: v1.1 WP-A1: the tokens kept free below the effective window, beyond the reply.
#:
#: The builder counts with `cl100k_base`; the narrator counts with its own
#: tokenizer. The v1 evidence put the largest prompts 23-42 real tokens from the
#: edge of a 16,384 window, and Ollama does not refuse a prompt past the edge —
#: measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt came back 200
#: with `prompt_tokens` 2,050. So the reserve is deliberate and sized to the
#: window: the larger of a floor and a share, **rounded up to a whole token**.
#:
#: 4,096 -> 256 8,192 -> 410 16,384 -> 820
#:
#: A fixed, documented tolerance, owner-chosen for v1.1. It is not a setting and
#: it is not calibrated per model.
SAFETY_RESERVE_FLOOR = 256
SAFETY_RESERVE_PERCENT = 5
def safety_reserve(effective_window: int) -> int:
"""`max(256, ceil(5% of the effective window))`, in tokens.
The effective window is the budget the prompt is actually built to — the
verified or declared window when there is one, the configured budget
otherwise — so a 16,384 setting against a 4,096 server reserves 256, not 820.
Integer arithmetic, so the rounding is exact rather than a float's.
"""
share = math.ceil(max(0, effective_window) * SAFETY_RESERVE_PERCENT / 100)
return max(SAFETY_RESERVE_FLOOR, share)
#: v1.1 WP-A1: what the server's own count says about a turn that was sent.
FITS = "fits"
EXCEEDED = "exceeded"
TRUNCATION_SUSPECTED = "truncation_suspected"
#: `UNKNOWN` above: the server reported no usable count.
def classify_usage(usage: dict | None, *, estimate: int, budget: int,
max_output_tokens: int, window_verified: bool) -> dict:
"""Sets the server's reported prompt count against what the application sent.
The order of the checks is the order of what they prove:
``unknown``
No positive integer `prompt_tokens`. Nothing can be said, and nothing
is claimed: an absent count is never read as a prompt that fitted.
``truncation_suspected``
The server read fewer tokens than were sent by more than the safety
reserve. A tokenizer thriftier than `cl100k_base` may honestly count a
little less; a shortfall larger than the tolerance the application keeps
for drift is the signature of a server that cut the prompt — the real
shape was 6,316 sent and 2,050 read.
``exceeded``
The server's count plus the reply allocation is more than the window
the prompt was built for. The drift was larger than the whole reserve,
so the reply may have been cut short.
``fits``
Otherwise.
`observed_margin` is what was left beside the reply by the server's count:
`budget - max_output_tokens - server_prompt_tokens`. The safety reserve is
the tolerance, so a margin between 0 and the reserve is still `fits`.
A discrepancy is recorded, never acted on: the reply has already streamed
to the reader and is accepted story.
"""
prompt = usage.get("prompt_tokens") if isinstance(usage, dict) else None
reserve = safety_reserve(budget)
verified_note = "" if window_verified else (
" The window itself was not verified for this turn.")
record = {
"status": UNKNOWN,
"server_prompt_tokens": None,
"estimate": estimate,
"difference": None,
"budget": budget,
"max_output_tokens": max_output_tokens,
"safety_reserve": reserve,
"observed_margin": None,
"window_verified": bool(window_verified),
"detail": "",
}
if type(prompt) is not int or prompt <= 0:
record["detail"] = ("The server reported no prompt token count, so nothing "
"confirms the whole prompt was read." + verified_note)
return record
record["server_prompt_tokens"] = prompt
record["difference"] = prompt - estimate
record["observed_margin"] = budget - max_output_tokens - prompt
if prompt + reserve < estimate:
record["status"] = TRUNCATION_SUSPECTED
record["detail"] = (
f"The server read {prompt:,} prompt tokens of the {estimate:,} sent, a "
f"shortfall larger than the {reserve:,}-token safety reserve. A server "
"that cuts an over-window prompt reports exactly this, and what it cuts "
"is the start: the narrator's rules and the canon." + verified_note)
elif prompt + max_output_tokens > budget:
record["status"] = EXCEEDED
record["detail"] = (
f"The server counted {prompt:,} prompt tokens; with {max_output_tokens:,} "
f"for the reply that is more than the {budget:,}-token window the prompt "
"was built for, so the reply may have been cut short." + verified_note)
else:
record["status"] = FITS
record["detail"] = (
f"The server read {prompt:,} prompt tokens, leaving "
f"{record['observed_margin']:,} beside the reply." + verified_note)
return record
def cache_clear() -> None:
"""Forgets what was learned. Called when the endpoint or model changes."""
_cache.clear()
@@ -223,6 +339,7 @@ def _declared_or(declared: int | None, discovered: Window) -> Window:
declared, DECLARED, discovered.model_max,
f"{declared:,} tokens, declared in settings — the server was not able "
f"to say ({discovered.detail})",
reachable=discovered.reachable,
)
@@ -266,6 +383,7 @@ async def _ask(endpoint_url: str, model: str) -> Window:
return Window(
tokens, LOADED, ceiling,
f"{tokens:,} tokens, reported by the running model",
reachable=True,
)
return await _declared_window(client, base, model)
except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc:
@@ -293,6 +411,7 @@ async def _declared_window(client, base: str, model: str) -> Window:
return Window(
None, UNKNOWN,
detail=f"the server did not describe the model (HTTP {resp.status_code})",
reachable=True,
)
body = resp.json() or {}
ceiling = _architecture_ceiling(body.get("model_info") or {})
@@ -304,14 +423,103 @@ async def _declared_window(client, base: str, model: str) -> Window:
"the model sets no num_ctx, so the server will load it at its own "
"default — which is 4,096 where there is no VRAM"
),
reachable=True,
)
tokens = min(declared, ceiling) if ceiling else declared
return Window(
tokens, PARAMETERS, ceiling,
f"{tokens:,} tokens, from the model's own num_ctx",
reachable=True,
)
#: v1.1 WP-A1 corrective: loading the configured model so its window can be read.
#:
#: The first real turn of the A1 evidence found a cold model: `/api/ps` knew
#: nothing, `/api/show` found no `num_ctx`, so the window was unverified and the
#: prompt was built to the configured 16,384. Ollama loaded the model at its own
#: 4,096 default, kept 2,050 of 13,875 tokens and answered 200. That case is
#: preventable, because the window becomes readable the moment the model is
#: resident. Ollama's native `POST /api/generate` with a model and **no prompt**
#: loads the model and generates nothing — measured on Ollama 0.33: HTTP 200,
#: `"response": ""`, `"done_reason": "load"`, and `/api/ps` then reported the
#: window. The OpenAI-compatible request that followed did not reload it.
WARM_PATH = "/api/generate"
async def warm(endpoint_url: str, model: str, *, timeout: float) -> tuple[bool, str]:
"""Asks the configured server to load `model`. One request, no story text.
Held to the same endpoint policy and TLS trust as inference and the probe, and
sent to the same host the probe asks. The body names the model and nothing
else: no prompt, so nothing is generated, and no `options` or `keep_alive`, so
the model loads the way the server would load it for the turn itself.
Returns `(loaded, detail)`. Every failure is `(False, why)` and never raises:
a server that will not load the model on request will fail the turn's own
call the ordinary way, which is where that failure belongs.
"""
reason = endpoints.rejection_reason(endpoint_url)
if reason is not None:
return False, f"endpoint not allowed — {reason}"
base = native_base(endpoint_url)
try:
async with httpx.AsyncClient(
timeout=httpx.Timeout(timeout, connect=CONNECT_TIMEOUT),
verify=tlstrust.ssl_context(),
) as client:
resp = await client.post(f"{base}{WARM_PATH}", json={"model": model})
except httpx.HTTPError as exc:
log.debug("model warm-up failed for %s: %s", base, exc)
return False, f"could not ask the server to load the model ({type(exc).__name__})"
if resp.status_code != 200:
return False, f"the server did not load the model (HTTP {resp.status_code})"
try:
body = resp.json() or {}
except ValueError:
return False, "the server answered the load request with something that was not JSON"
return True, f"the server loaded the model ({body.get('done_reason') or 'done'})"
async def ensure_window(endpoint_url: str, model: str, *, declared: int | None = None,
warm_timeout: float = 300.0) -> tuple[Window, dict]:
"""The window for a turn about to be generated, loading the model once if that is what it takes.
1. Probe as before.
2. If the window is not verified, the server answered, and there is a model to
load: one bounded `warm` request.
3. If the model loaded, probe again, bypassing the cache that still holds the
unverified answer.
Whatever the second probe says is the answer. There is no retry loop, no
guessed window, and no hard-coded 4,096: a window still unverified leaves the
configured budget standing, exactly as before, and the turn's accounting
still catches a server that cut the prompt.
Returns the window and a `preflight` record for the turn's provenance.
Not used by the context dry run: loading a model is a side effect, and
opening a panel should not cause one.
"""
window = await probe(endpoint_url, model, declared=declared)
preflight = {"attempted": False, "loaded": None, "verified_before": window.verified,
"verified_after": window.verified, "detail": ""}
if window.verified:
preflight["detail"] = "the window was already verified"
return window, preflight
if not (endpoint_url and model):
preflight["detail"] = "no endpoint or model configured"
return window, preflight
if not window.reachable:
preflight["detail"] = "the server did not answer, so no model was loaded"
return window, preflight
loaded, detail = await warm(endpoint_url, model, timeout=warm_timeout)
preflight.update(attempted=True, loaded=loaded, detail=detail)
if loaded:
window = await probe(endpoint_url, model, declared=declared, use_cache=False)
preflight["verified_after"] = window.verified
return window, preflight
def _num_ctx(parameters) -> int | None:
"""Reads `num_ctx` out of the plain-text parameter block Ollama returns."""
if not isinstance(parameters, str):
+28
View File
@@ -129,6 +129,34 @@ MAX_BODY_BYTES = 2 * 1024 * 1024
MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024
def import_limit_label(limit: int | None = None) -> str:
"""The import ceiling as a reader would say it, e.g. "20 MB".
Derived from the constant rather than written beside it, so the refusal, the
export warning and the documentation cannot drift apart from each other or
from what the middleware actually enforces (v1.1 WP-D).
"""
size = MAX_IMPORT_BODY_BYTES if limit is None else limit
megabytes = size / (1024 * 1024)
return f"{megabytes:.0f} MB" if abs(megabytes - round(megabytes)) < 0.05 else f"{megabytes:.1f} MB"
def oversized_export_warning(export_bytes: int, limit: int | None = None) -> str:
"""What to tell a reader whose export is larger than import will accept.
v1.1 WP-D. The file is written and is not damaged: what it exceeds is this
version's import ceiling, so it cannot be brought back in *here*. Saying that
plainly is the whole point — the alternative is a reader who finds out when
they try to restore it.
"""
size = MAX_IMPORT_BODY_BYTES if limit is None else limit
return (
f"This export is larger than this version's {import_limit_label(size)} import "
f"limit ({export_bytes:,} bytes). The file was exported successfully, but this "
f"version cannot import it."
)
class BodySizeLimitMiddleware:
"""Rejects oversized request bodies by their declared `Content-Length`.
+450 -76
View File
@@ -17,9 +17,10 @@ database session. It does three things:
then evicts the bank down to its capacity. Evicted memories are marked as
forgotten and kept so that the UI can still show them.
When the app generates a turn, `retrieve_memories` embeds the recent story text
and ranks the bank by cosine similarity. The highest-ranked memories become the
Memories section of the context.
When the app generates a turn, `retrieve_memories` embeds the player's input
and the current scene, and ranks the bank by a fixed mix of cosine similarity
and rarity-weighted word overlap with the input (v1.1 WP-B.2). The
highest-ranked memories become the Memories section of the context.
Every AI call in this module is best-effort. A failure is logged to the debug
page and retried on a later turn, because the cursors advance only after a call
@@ -28,6 +29,7 @@ succeeds.
import asyncio
import logging
import math
from array import array
from collections import OrderedDict
@@ -36,6 +38,7 @@ from sqlalchemy.orm import Session, defer, object_session
from . import derived, models, summaries, tree, vectors
from .context import (
count_tokens,
cursors,
history,
lineage,
@@ -43,8 +46,11 @@ from .context import (
story_actions,
truncate_to_last_tokens,
)
from .context.builder import _encoding as _token_encoding
from .database import SessionLocal
from .knowledge import embeddings as knowledge_embeddings
from .knowledge import fts
from .narrative import model as narrative_model
from .providers import OpenAICompatibleProvider, ProviderError
from .vectors import cosine # re-exported: the ranking lives here, the maths there
@@ -55,10 +61,14 @@ MEMORY_START = 12 # first memory once the adventure reaches this many actions
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
MAX_MEMORIES_PER_RUN = 5 # cap catch-up work (e.g. imported adventures) per turn
MAX_EMBED_BATCH = 32
RETRIEVAL_WINDOW_TOKENS = 600 # recent story text used as the similarity query
RETRIEVAL_WINDOW_ACTIONS = 4 # ...taken from this many of the newest actions
SUMMARY_MAX_WORDS = 250
MEMORY_EXCERPT_TOKENS = 2000 # of the block, when a block is longer than this
MEMORY_EXCERPT_TOKENS = 2000 # the most of a block the summariser is shown
# v1.1 WP-B.2: what stands between the two parts of a block too long to send
# whole. It says a part is missing, so the summariser does not read the end as
# following straight on from the opening, and `summarize_block` removes it from
# anything the model repeats back.
EXCERPT_OMISSION_MARKER = "[… the middle of this stretch of story is left out here …]"
# How much story has to sit past a block before that block is summarized.
#
@@ -278,9 +288,10 @@ def set_vector(memory: models.Memory, vector: list[float] | None) -> None:
"""
memory.embedding_blob = None if vector is None else vectors.pack(vector)
memory.embedded = vector is not None
cached = _vector_cache.get(memory.adventure_id)
if cached is not None:
cached.pop(memory.id, None)
for cache in (_vector_cache, _terms_cache):
cached = cache.get(memory.adventure_id)
if cached is not None:
cached.pop(memory.id, None)
# ---------- The vector cache ----------
@@ -309,6 +320,37 @@ _vector_cache: OrderedDict[int, dict[int, array]] = OrderedDict()
VECTOR_CACHE_ADVENTURES = 8 # ~600 KB each at a 100-memory bank
# v1.1 WP-B.2: each memory's lexical terms, held the same way and by the same
# two rules as its vector. `set_vector` is also where a memory's text changes
# (an edit clears the vector to re-embed it), so dropping the entry there covers
# a rewritten text as well as a rewritten vector. Text is read only for memories
# not already held, and only on a turn whose input has words to match.
_terms_cache: OrderedDict[int, dict[int, frozenset[str]]] = OrderedDict()
def _terms_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, frozenset[str]]:
"""The lexical terms for `ids`, reading text only for the ones not already held."""
cached = _terms_cache.get(adventure_id)
if cached is None:
cached = _terms_cache[adventure_id] = {}
_terms_cache.move_to_end(adventure_id)
while len(_terms_cache) > VECTOR_CACHE_ADVENTURES:
_terms_cache.popitem(last=False)
wanted = set(ids)
for gone in set(cached) - wanted:
del cached[gone]
missing = [memory_id for memory_id in ids if memory_id not in cached]
if missing:
rows = db.execute(
select(models.Memory.id, models.Memory.text)
.where(models.Memory.id.in_(missing))
).all()
for memory_id, text in rows:
cached[memory_id] = lexical_terms(text or "")
return cached
def forget_cached_vectors(adventure_id: int) -> None:
"""Drops an adventure's cached vectors.
@@ -316,6 +358,7 @@ def forget_cached_vectors(adventure_id: int) -> None:
corrects itself, as described in the comment above.
"""
_vector_cache.pop(adventure_id, None)
_terms_cache.pop(adventure_id, None)
def _vectors_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, array]:
@@ -535,6 +578,199 @@ def cast_brief(adventure: models.Adventure, text: str) -> str:
# ---------- Retrieval (runs inside the turn, before build_context) ----------
# v1.1 WP-B.2: what the retrieval query is made of, and how a memory is scored
# against it (CONTEXT-AND-MEMORY §18, §20).
#
# WP-B.1 measured the v1.0.0 query, the newest four actions cut to 600 tokens,
# against a planted early fact. The player's one-line question arrived after
# three turns of narration, so the embedding mostly described the narration: a
# direct question about the fact fell from cosine 0.708 on its own to 0.241 in
# that query, and a real 100-turn campaign ranked the only memory of the fact
# 10th of 19 against a `memory_top_k` of 4.
#
# The query is now two short texts, embedded in one call:
#
# input the player's own action this turn, when there is one
# context the current scene from the authoritative state (summary, location,
# who is present), then the end of the newest narration
#
# The context is still there because a question often cannot be read without
# it ("I ask her where she hid it"), and §18 says retrieval must not rely on raw
# input alone. It is bounded so it can resolve a reference but cannot outweigh
# the question by sheer length.
#
# A memory's score is
#
# semantic_score = INPUT_WEIGHT * cos(input, memory)
# + (1 - INPUT_WEIGHT) * cos(context, memory)
# lexical_score = rarity-weighted share of the input's words the memory holds
# final_score = semantic_score + LEXICAL_WEIGHT * lexical_score
#
# With no player input (a continue, or a dry run from Insights) the semantic
# score is the context cosine alone and the lexical score is 0. Pins are
# unchanged: a pinned memory is always used and counts toward `memory_top_k`.
INPUT_TYPES = ("do", "say", "story") # player actions that carry words to search for
QUERY_INPUT_TOKENS = 200 # of the player's action; a long `story` entry is cut
QUERY_SCENE_TOKENS = 60 # of the state's scene line
QUERY_NARRATION_TOKENS = 120 # from the end of the newest narration
INPUT_WEIGHT = 0.6
# Chosen by sweep (0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5) over the deterministic
# ranking fixtures, recorded in the WP-B.2 report (§C, §D). The two-part query
# alone already ranks the planting-era memory first; 0.15 is the smallest weight
# at which the lexical term by itself also lifts it into `memory_top_k` against
# the v1.0.0 narration-filled query, and no rare-word negative control put an
# unrelated memory above it. At 0.5 an incidental shared word was enough to
# select it for an unrelated question, which is the failure a larger weight buys.
LEXICAL_WEIGHT = 0.15
# `fts.terms` drops these already; the plural fold below is the only stemming.
_MIN_FOLD_LENGTH = 5
def _fold(word: str) -> str:
"""One term, reduced so "shelves'" and "shelf" do not meet, but "teapots"
and "teapot" do. Possessives lose their `'s`, and a trailing `s` goes from a
word long enough to be a plural and not ending in `ss`. Deliberately no more
than that: a stemmer is a dependency, and a wrong fold merges two words."""
word = word.split("'", 1)[0]
if len(word) >= _MIN_FOLD_LENGTH and word.endswith("s") and not word.endswith("ss"):
word = word[:-1]
return word
def lexical_terms(text: str) -> frozenset[str]:
"""The words of `text` that lexical matching compares, folded.
The tokenizer and stop list are imported knowledge's (`knowledge.fts`), so
the two retrieval paths agree on what a word is.
"""
return frozenset(t for t in (_fold(w) for w in fts.terms(text)) if len(t) >= fts.MIN_TERM_LENGTH)
def lexical_scores(input_terms: frozenset[str], terms_of: dict[int, frozenset[str]]) -> dict[int, float]:
"""Each candidate's share of the input's rarity, in [0, 1].
A term's weight is `ln((N + 1) / (df + 1))`: N candidates, df of them holding
it. A word every candidate holds weighs exactly 0, so a protagonist's name or
a word the whole bank shares moves nothing, and a word no candidate holds
weighs the most. The share is taken over **all** the input's terms, so a
memory that happens to hold one rare word of a longer question gets that
word's part of the question, not the whole of it. The weights live only for
this call, over this candidate set: no index, no stored field.
"""
if not input_terms or not terms_of:
return {memory_id: 0.0 for memory_id in terms_of}
n = len(terms_of)
weight = {
term: math.log((n + 1) / (sum(1 for terms in terms_of.values() if term in terms) + 1))
for term in input_terms
}
total = sum(weight.values())
if total <= 0:
return {memory_id: 0.0 for memory_id in terms_of}
return {
memory_id: min(1.0, sum(w for term, w in weight.items() if term in terms) / total)
for memory_id, terms in terms_of.items()
}
def _scene_text(state) -> str:
"""The scene as the authoritative state has it: summary, location, who is present.
Names only, read straight off the document. The full entity list is left
out on purpose: a campaign with a large cast would turn every query into a
search for everyone.
"""
if not isinstance(state, dict):
return ""
scene = state.get("scene")
if not isinstance(scene, dict):
return ""
pieces: list[str] = []
summary = scene.get("summary")
if isinstance(summary, str) and summary.strip():
pieces.append(summary.strip())
location = scene.get("location")
if isinstance(location, str) and location.strip():
pieces.append(narrative_model.entity_name(state, location.strip()))
present = scene.get("present")
if isinstance(present, list):
names = [narrative_model.entity_name(state, key) for key in present[:8]
if isinstance(key, str) and key.strip()]
if names:
pieces.append(", ".join(names))
return truncate_to_last_tokens(". ".join(pieces), QUERY_SCENE_TOKENS)
def retrieval_query(adventure: models.Adventure, exclude_action_id: int | None = None) -> dict:
"""The two texts a turn's memory retrieval embeds, and the words it matches.
Returns `{"input", "context", "input_terms"}`. `input` is empty when the
newest action is not a player action with text, which is a continue turn or a
dry run. `context` is empty only for a story with no scene and no narration.
"""
recent = history.tail(adventure, 2, exclude_action_id)
newest = recent[-1] if recent else None
player_input = ""
if newest is not None and newest.type in INPUT_TYPES:
player_input = truncate_to_last_tokens(newest.text.strip(), QUERY_INPUT_TOKENS)
narration = recent[0].text if len(recent) > 1 else ""
else:
narration = newest.text if newest is not None else ""
context = "\n".join(part for part in (
_scene_text(adventure.narrative_state),
truncate_to_last_tokens(narration.strip(), QUERY_NARRATION_TOKENS),
) if part.strip())
return {
"input": player_input,
"context": context,
"input_terms": sorted(lexical_terms(player_input)),
}
def score_candidates(
ids: list[int],
held: dict,
terms_of: dict[int, frozenset[str]],
input_vec,
context_vec,
input_terms,
) -> list[tuple[float, int, float, float]]:
"""`(final_score, memory_id, semantic_score, lexical_score)`, best first.
Ties on the final score are broken by id, so the order never depends on the
order the database returned rows in.
"""
lexical = lexical_scores(frozenset(input_terms), {i: terms_of.get(i, frozenset()) for i in ids})
rows = []
for memory_id in ids:
vector = held[memory_id]
if input_vec is not None and context_vec is not None:
semantic = (INPUT_WEIGHT * cosine(input_vec, vector)
+ (1.0 - INPUT_WEIGHT) * cosine(context_vec, vector))
else:
semantic = cosine(input_vec if input_vec is not None else context_vec, vector)
lex = lexical.get(memory_id, 0.0)
rows.append((semantic + LEXICAL_WEIGHT * lex, memory_id, semantic, lex))
rows.sort(key=lambda row: (-row[0], row[1]))
return rows
def select_memories(scored, pinned_of, held, authority_of, top_k):
"""Pins first, then the best-scoring rest, skipping repeats (§22).
Returns `(used, suppressed)`. `used` is `(final_score, memory_id, pinned)`
rows, best first.
"""
rows = [(final, memory_id, pinned_of[memory_id]) for final, memory_id, _, _ in scored]
used = [row for row in rows if row[2]]
remaining = max(0, top_k - len(used))
candidates = [row for row in rows if not row[2]]
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
used += kept
used.sort(key=lambda row: (-row[0], row[1]))
return used, suppressed
async def retrieve_memories(
adventure: models.Adventure,
settings: models.Settings,
@@ -544,16 +780,17 @@ async def retrieve_memories(
"""Returns the memories to inject, or None when the bank is off.
The result is a dict of the form
`{"used": [{id, text, similarity, pinned}], "error": str | None}`. It is
None when the memory bank is disabled for this adventure.
`{"used": [{id, text, similarity, semantic_score, lexical_score,
final_score, pinned, authority, source}], "query": {...}, "error": str | None}`.
`similarity` is the semantic score, under the name the inspector has always
shown. It is None when the memory bank is disabled for this adventure.
This only reads. A turn counts the memories it used with `record_use`, just
before the commit that saves the turn; see that function for why the count
cannot be written here.
`exclude_action_id` removes the action being retried from the similarity
query, so that a discarded attempt cannot influence which memories are
returned.
`exclude_action_id` removes the action being retried from the query, so that
a discarded attempt cannot influence which memories are returned.
"""
if not adventure.memory_bank_enabled:
return None
@@ -585,39 +822,33 @@ async def retrieve_memories(
if not catalogue:
return {"used": [], "error": None}
recent = history.tail(adventure, RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
query = truncate_to_last_tokens(
"\n\n".join(a.text for a in recent), RETRIEVAL_WINDOW_TOKENS
)
if not query.strip():
query = retrieval_query(adventure, exclude_action_id)
texts = [t for t in (query["input"], query["context"]) if t.strip()]
if not texts:
return {"used": [], "error": None}
try:
[query_vec] = await embedding_provider(settings).embed([query])
embedded = await embedding_provider(settings).embed(texts)
except ProviderError as exc:
return {"used": [], "error": str(exc)}
vectors_by_text = dict(zip(texts, embedded))
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
held = _vectors_for(db, adventure.id, [memory_id for memory_id, _, _ in catalogue])
ids = [memory_id for memory_id, _, _ in catalogue]
held = _vectors_for(db, adventure.id, ids)
# Memory text is read only when there are input words to match against, and
# then only for memories not already held (see `_terms_for`).
terms_of = _terms_for(db, adventure.id, ids) if query["input_terms"] else {}
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
scored = sorted(
(
(cosine(query_vec, held[memory_id]), memory_id, pinned)
for memory_id, pinned, _ in catalogue
if memory_id in held
),
key=lambda row: row[0],
reverse=True,
pinned_of = {memory_id: pinned for memory_id, pinned, _ in catalogue}
scored = score_candidates(
[memory_id for memory_id in ids if memory_id in held],
held, terms_of, input_vec, context_vec, query["input_terms"],
)
# Pinned memories are always used, and they count toward `top_k`, so the
# injected set stays within the budget unless the pinned memories alone
# exceed it.
top_k = max(1, settings.memory_top_k)
used = [row for row in scored if row[2]]
remaining = max(0, top_k - len(used))
candidates = [row for row in scored if not row[2]]
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
used += kept
used.sort(key=lambda row: row[0], reverse=True)
components = {memory_id: (semantic, lex) for _, memory_id, semantic, lex in scored}
used, suppressed = select_memories(
scored, pinned_of, held, authority_of, max(1, settings.memory_top_k))
if not used:
return {"used": [], "error": None}
@@ -638,14 +869,19 @@ async def retrieve_memories(
).where(models.Memory.id.in_(used_ids))
).all()
}
texts = {memory_id: row.text for memory_id, row in detail.items()}
texts_of = {memory_id: row.text for memory_id, row in detail.items()}
return {
"used": [
{
"id": memory_id,
"text": texts.get(memory_id, ""),
"similarity": round(score, 4),
"text": texts_of.get(memory_id, ""),
"similarity": round(components[memory_id][0], 4),
# v1.1 WP-B.2: the parts of the score, so an inspector can see
# why this memory beat the ones below it.
"semantic_score": round(components[memory_id][0], 4),
"lexical_score": round(components[memory_id][1], 4),
"final_score": round(final, 4),
"pinned": pinned,
# M6: what weight this carries, and where it came from.
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
@@ -656,7 +892,7 @@ async def retrieve_memories(
"source_end": getattr(detail.get(memory_id), "source_end", None),
},
}
for score, memory_id, pinned in used
for final, memory_id, pinned in used
],
"considered": len(catalogue),
# M6: how many candidates were set aside as repeating one already
@@ -665,6 +901,15 @@ async def retrieve_memories(
{"id": memory_id, "duplicate_of": kept_id}
for memory_id, kept_id in suppressed
],
# v1.1 WP-B.2: what was searched for. Recorded per turn, like the rest.
"query": {
"input": query["input"],
"context": query["context"],
"input_terms": query["input_terms"],
"input_weight": INPUT_WEIGHT if input_vec is not None and context_vec is not None
else (1.0 if input_vec is not None else 0.0),
"lexical_weight": LEXICAL_WEIGHT,
},
"error": None,
}
@@ -822,6 +1067,61 @@ async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
db.commit()
def _excerpt_encoding():
return _token_encoding()
def excerpt_split(budget: int = MEMORY_EXCERPT_TOKENS) -> tuple[int, int]:
"""`(head_tokens, tail_tokens)` for a block longer than `budget`.
The marker and the blank lines around it are paid for first; what is left is
halved, and an odd token goes to the tail, the most recent part. So the two
parts plus the marker come to exactly `budget`.
"""
room = max(0, budget - count_tokens(f"\n\n{EXCERPT_OMISSION_MARKER}\n\n"))
head = room // 2
return head, room - head
def memory_excerpt(raw: str, budget: int = MEMORY_EXCERPT_TOKENS) -> str:
"""What the summariser is shown of one block.
v1.1 WP-B.2. A block that fits in `budget` tokens is sent whole, exactly as
before. A longer block used to be cut to its last `budget` tokens, and B.1
showed that a fact near its start then never reached the summariser at all.
It is now sent as its opening and its end, in order, with
`EXCERPT_OMISSION_MARKER` between them, still inside `budget`.
Rejoining two token runs can tokenise a little differently at the seams, so
the result is measured, and the head gives up tokens until it fits. A fact in
the middle of a very long block is still left out: this bounds the input, it
does not summarise everything.
"""
enc = _excerpt_encoding()
tokens = enc.encode(raw)
if len(tokens) <= budget:
return raw
head_n, tail_n = excerpt_split(budget)
while True:
excerpt = (f"{enc.decode(tokens[:head_n]).rstrip()}\n\n{EXCERPT_OMISSION_MARKER}\n\n"
f"{enc.decode(tokens[-tail_n:]).lstrip()}" if tail_n else
enc.decode(tokens[:head_n]))
over = count_tokens(excerpt) - budget
if over <= 0 or head_n == 0:
return excerpt
head_n = max(0, head_n - over)
def memory_user_prompt(brief: str, excerpt: str) -> str:
"""The user message of a memory call: the cast brief, then the excerpt.
Kept apart from `summarize_block` so an evaluation can send a model exactly
what the application sends (v1.1 WP-B.2, `tools/memory_fidelity.py`).
"""
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
return f"{brief}\n\n{prompt}" if brief else prompt
async def summarize_block(
adventure: models.Adventure,
provider: OpenAICompatibleProvider,
@@ -840,15 +1140,16 @@ async def summarize_block(
old text in place and moves on.
"""
raw = "\n\n".join(a.text for a in block)
excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)
excerpt = memory_excerpt(raw)
# Match the cast against the untruncated block. The excerpt is what the
# model reads, but a character named in the part that was trimmed is still
# model reads, but a character named in the part that was left out is still
# one the memory may have to name.
brief = cast_brief(adventure, raw)
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
return await provider.complete(
MEMORY_SYSTEM_PROMPT, f"{brief}\n\n{prompt}" if brief else prompt
)
text = await provider.complete(MEMORY_SYSTEM_PROMPT, memory_user_prompt(brief, excerpt))
# The marker is an instruction to the summariser, never a fact of the story.
if text and EXCERPT_OMISSION_MARKER in text:
text = " ".join(text.replace(EXCERPT_OMISSION_MARKER, " ").split())
return text
async def _create_due_memories(
@@ -1028,11 +1329,93 @@ async def _embed_pending(
return len(pending)
def eviction_order(rows, limit: int) -> list[int]:
"""The ids eviction would take, first to last, at most `limit` of them.
v1.1 WP-B.2. `rows` are the active memories of one adventure, each with
`id`, `pinned`, `source_start`, `source_end`, `last_used_at`, `created_at`
and `use_count`. Nothing here reads a vector or the database, so the same
function is what the eviction pass runs and what a diagnostic reports.
WP-B.1 showed what pure least-recently-used order does to a long campaign.
Retrieval is steered by the present scene, so a memory of an early stretch
nothing recent resembles stops being used. It then becomes the least
recently used row, and it goes first, while the bank keeps several memories
of the last few scenes that the history window still holds in full. The
rule below keeps the bank spread over the whole story instead.
**Coverage.** Memories with a source range say which stretch of the story
they describe. A memory is judged by the hole its removal would leave: the
number of depths between the end of the nearest memory before it and the
start of the nearest memory after it. The smallest hole goes first, so the
bank thins where it is densest. A memory whose start another memory shares
(a retried or re-played stretch, or a sibling line) leaves no hole, and is
the first kind to go. Pinned memories count as coverage, since they stay.
**Boundaries.** The earliest and the latest memory by position leave a hole
with no memory on one side: removing the first loses the only record of the
opening, and removing the last loses the only record of the most recent
stretch, which is also what keeps a memory written this turn from being
evicted by the pass that wrote it (the frozen bank, below). Boundaries are
not coverage candidates.
**Recency.** Among memories whose removal leaves the same hole, the least
recently used goes first (`coalesce(last_used_at, created_at)`), then the
less used, then the lower id. Ties are therefore never left to the order the
database returned rows in.
**Fallback.** When no memory is a coverage candidate — memories typed by the
player or migrated from before coordinates have no range, and a bank can be
all boundaries — the rest are taken least recently used first, exactly as
v1.0.0 did. The bank stays bounded either way. Pinned memories are never
taken; if every active memory is pinned, capacity yields to the pins.
Recomputed after each pick, because removing one memory widens the holes
of its neighbours.
"""
remaining = {row.id: row for row in rows if not row.pinned}
coverers = {row.id: row for row in rows
if row.source_start is not None and row.source_end is not None}
def recency(row):
return (row.last_used_at or row.created_at, row.use_count or 0, row.id)
order: list[int] = []
while remaining and len(order) < limit:
spans = sorted(coverers.values(), key=lambda r: (r.source_start, r.source_end, r.id))
starts: dict[int, int] = {}
for row in spans:
starts[row.source_start] = starts.get(row.source_start, 0) + 1
best = None
furthest_end = None # the largest source_end before index i
for i, row in enumerate(spans):
if row.id in remaining:
if starts[row.source_start] > 1:
cost = 0
elif i == 0 or i == len(spans) - 1:
cost = None # a boundary
else:
cost = max(0, spans[i + 1].source_start - furthest_end - 1)
if cost is not None:
key = (cost, *recency(row))
if best is None or key < best[0]:
best = (key, row.id)
furthest_end = row.source_end if furthest_end is None else max(furthest_end, row.source_end)
if best is None:
victim = min(remaining.values(), key=recency).id
else:
victim = best[1]
order.append(victim)
del remaining[victim]
coverers.pop(victim, None)
return order
def _evict_over_capacity(
adventure: models.Adventure, settings: models.Settings, db: Session
) -> None:
# The database performs both the count and the ranking, and returns neither
# the rows nor the vectors. Counting by walking `adventure.memories` fetched
# The database performs the count, and the rows read for ordering carry
# neither text nor vectors. Counting by walking `adventure.memories` fetched
# every vector in the bank on every turn, whether or not the bank was over
# capacity.
in_this_bank = (models.Memory.adventure_id == adventure.id,
@@ -1043,32 +1426,23 @@ def _evict_over_capacity(
overflow = active - max(1, settings.memory_bank_capacity)
if overflow <= 0:
return
# Evict the least recently used memory first, and use the use count only to
# break ties.
# v1.1 WP-B.2: the order is `eviction_order`, coverage first and recency
# second. It replaces least recently used alone; see that function.
#
# Ordering by use count first froze the bank. A memory written on this turn
# has never been used, so once every other memory had been retrieved at
# least once, the new memory held the lowest count in the bank. The same
# post-turn run that wrote it then evicted it, one pass after embedding it.
# Use counts only increase, so the bank never recovered. An adventure kept
# whatever memories it held when the bank first filled, and every later
# memory was summarized, marked as forgotten, and never ranked.
#
# Ordering by recency avoids that. A new memory carries the newest
# timestamp, so it is the last row to be evicted rather than the first, and
# it remains until other memories are used. Demoting the use count costs
# little, because retrieving a useful memory also makes it recent. The two
# orderings differ only for memories that were used once and have not been
# retrieved since, which are the rows a full bank should evict.
doomed = db.execute(
select(models.Memory.id)
.where(*in_this_bank, models.Memory.pinned.is_(False))
.order_by(
func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
models.Memory.use_count,
)
.limit(overflow)
).scalars().all()
# What the old ordering fixed still holds. Ordering by use count first froze
# the bank: a memory written on this turn has never been used, so once every
# other memory had been retrieved at least once, the new memory held the
# lowest count in the bank, and the same post-turn run that wrote it evicted
# it. Counts only increase, so the bank never recovered. Under the coverage
# rule the newest memory is the latest boundary, so it is not a coverage
# candidate, and in the fallback it carries the newest timestamp.
rows = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
models.Memory.source_end, models.Memory.last_used_at,
models.Memory.created_at, models.Memory.use_count)
.where(*in_this_bank)
).all()
doomed = eviction_order(rows, overflow)
if not doomed:
return # Every active memory is pinned, so the pins override capacity.
db.execute(
+25 -4
View File
@@ -35,6 +35,8 @@ creating a second, empty Mara.
from __future__ import annotations
import json
# Field types the schema layer enforces. Kept deliberately small: a narrative
# state event carries names, labels and plain values, and nothing here needs a
# nested structure a model could hide something inside.
@@ -184,11 +186,30 @@ def vocabulary_for_prompt() -> str:
Generated from `SPECS` rather than written out beside it, so the model can
never be told about an event the application does not implement — the drift
that would produce proposals rejected for reasons nobody could see.
v1.1 WP-A2: each event is shown as the object the model must put in the
`events` list, with its required fields, not as `name(field, …)`. The call
notation was never the wire format, and a 3B narrator copied it into its
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
is a proposal the extractor already recognises and removes; a call is not.
"""
lines = []
for name, definition in SPECS.items():
fields = list(definition["required"]) + [
f"{field}?" for field in definition["optional"]
]
lines.append(f' {name}({", ".join(fields)}) — {definition["summary"]}')
shape = {"type": name}
for field, kind in definition["required"].items():
shape[field] = _PLACEHOLDER[kind]
body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
line = f" {body} — {definition['summary']}"
if definition["optional"]:
line += f" (optional: {', '.join(definition['optional'])})"
lines.append(line)
return "\n".join(lines)
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
#: never example identifiers, so the vocabulary names nothing a story could copy.
#: A list field is shown as a list, so the model is told its shape; every other
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
#: and every line is still the object the model must send.
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}
+188 -14
View File
@@ -37,17 +37,17 @@ EMIT_RULE = (
"appeared, record it.\n"
"\n"
"Every value is ABSOLUTE — the new state of things, never a change or a "
"difference. Use only these events:\n"
"difference. Use only these events, in exactly this shape:\n"
f"{events.vocabulary_for_prompt()}\n"
"\n"
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must "
"match the ones already in the state you were shown. Introduce a person, place "
"or thing with create_entity before referring to it. If the turn established "
"nothing, send an empty events list.\n"
"Identifiers are short lower-case slugs and must match the ones already in the "
"state you were shown; the example's identifiers are placeholders. Introduce a "
"person, place or thing with create_entity before referring to it. If the turn "
"established nothing, send an empty events list.\n"
"Example:\n"
'```state\n'
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},'
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n'
'{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
'```'
)
@@ -58,6 +58,47 @@ EMIT_REMINDER = (
"nothing changed.]"
)
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
# builds the hint from these, and the extractor recognises an echo of it by
# them, so the two cannot drift apart.
LENGTH_HINT_OPENING = "[Hard limit:"
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
#: The application's wording inside a hint. A 3B narrator reworded the front
#: ("your next turn") and the end ("This story ends here."), and kept one or the
#: other of these every time.
_LENGTH_HINT_PHRASE_RE = re.compile(
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
)
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
#: replay tool and the report can say which removed what.
RULE_EVENT_CALL = "event_call_line"
RULE_LENGTH_HINT = "echoed_length_hint"
RULE_SCENE_LINE = "rendered_scene_line"
RULE_EMPTY_FENCE = "empty_dangling_fence"
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
#: of the hint reworded around it. Kept as a copy rather than an import, so the
#: narrative package does not depend on the provider; a test pins that the
#: hint still contains it.
CONTINUE_HINT_PHRASE = "Output only story text"
# R1. A whole line opening with a call to an event this protocol has. The names
# come from the vocabulary, so a call-shaped line naming anything else — a
# character's `open_door(north)` — is not matched.
_EVENT_CALL_LINE_RE = re.compile(
r"^[ \t]*(?:>[ \t]*)?(?:"
+ "|".join(re.escape(name) for name in events.SPECS)
+ r")[ \t]*\(",
re.IGNORECASE,
)
# R3. The renderer's scene line carries its location this way.
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
# R4. An opener with nothing after it.
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
# Three patterns, and the difference between them is the whole of this module's
# safety. A story is allowed to contain code, and taking a code block out of
# someone's prose is a worse failure than leaving a stray proposal in it.
@@ -128,11 +169,42 @@ def _is_echoed_instruction(inner: str) -> bool:
# opening words only, because the echo is often cut off before it ends.
if low.lstrip().startswith("continue the story directly"):
return True
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
# closeout-era identity re-run stored "[You don't need to continue; … Continue
# the story here, directly. Output only story text.]" as the last line of a
# reply, and because nothing recognised it, nothing above it was trailing.
if CONTINUE_HINT_PHRASE.lower() in low:
return True
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
# "events list", so it passed every check above.
if _is_length_hint(inner):
return True
# The reminder names both; prose about the protocol rarely names either the
# way the instruction does, and effectively never both.
return "state block" in low and "events list" in low
def _opens_like_length_hint(inner: str) -> bool:
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
way. `_clean` takes it only directly above an echoed instruction it has already
removed from the end of the same reply.
"""
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
def _is_length_hint(inner: str) -> bool:
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
It must open the way the hint opens *and* carry the hint's own wording. An
in-world "Hard limit: forty days" has the opening and none of the wording.
"""
opening = LENGTH_HINT_OPENING[1:].lower()
return (inner.lstrip().lower().startswith(opening)
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
# A heading the model writes above a block it did not fence: `State`, sometimes
# as `State:`, `**State**` or `### State`. It is removed only in two places:
# directly above a proposal that is removed, and as the last line of the reply.
@@ -145,7 +217,7 @@ _LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
def _clean(prose: str) -> str:
def _clean(prose: str, *, after_block: bool = False) -> str:
"""Removes protocol the block extraction could not, and nothing else.
Found by the M5 realistic-context run (§12), which is the failure class
@@ -160,9 +232,23 @@ def _clean(prose: str) -> str:
story after it. Stored text is replayed as history, so every leak also
showed the next prompt a second, older account of the state, which is what
M5 review Finding 4 removed from replayed history.
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
v1 corpus, each anchored to something the application owns rather than to
what prose looks like: a line opening with a vocabulary call (R1), the
length hint echoed at the end (R2), the renderer's scene line left last
(R3), and an empty fence opener left last (R4). `after_block` says a
proposal block was already taken out of this reply, which is what lets R3
remove a bare scene line that sat above it.
"""
cleaned, _found = _inline_proposals(prose)
cleaned, calls_removed = _strip_event_call_lines(prose)
cleaned, _found = _inline_proposals(cleaned)
cleaned = _strip_echoed_state(cleaned)
protocol_cut = after_block or calls_removed
# R5: set once an echoed instruction bracket has come off the end. Only then
# may a bracket that merely opens the way the length hint opens be taken as
# part of the same echoed tail.
instruction_cut = False
# The end of the reply is cut until nothing more comes off, because one kind
# of leftover can hide another. In a real reply, a `State` heading sat above
# a block the model never finished, and a parroted reminder sat above an
@@ -171,7 +257,12 @@ def _clean(prose: str) -> str:
before = cleaned
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
bracket = pattern.search(cleaned)
if bracket is not None and _is_echoed_instruction(bracket.group(1)):
if bracket is None:
continue
if _is_echoed_instruction(bracket.group(1)):
cleaned = cleaned[: bracket.start()]
instruction_cut = True
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
cleaned = cleaned[: bracket.start()]
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
dangling = _DANGLING_JSON_RE.search(cleaned)
@@ -182,10 +273,93 @@ def _clean(prose: str) -> str:
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
# A bare quote marker, the start of a quoted block that never came.
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
cleaned = _strip_empty_dangling_fence(cleaned)
if cleaned.rstrip() != before.rstrip():
protocol_cut = True
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
if cleaned == before:
return cleaned.strip()
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
"""R1. Removes whole lines that open with a call to a vocabulary event.
A line inside a fenced code block is the story's own code and is never
examined. Returns the text and whether anything was removed.
"""
kept: list[str] = []
in_fence = False
removed = False
for line in text.split("\n"):
if line.lstrip().startswith("```"):
in_fence = not in_fence
kept.append(line)
continue
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
removed = True
continue
kept.append(line)
if not removed:
return text, False
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
def _strip_empty_dangling_fence(text: str) -> str:
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
Only an *opener*: the fence lines are counted, and an even count means the
last one closes a story's own code block, which stays.
"""
lines = text.rstrip().split("\n")
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
return text
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
if fences % 2 == 0:
return text
return "\n".join(lines[:-1]).rstrip()
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
"""R3. The renderer's scene line, left as the last line of the reply.
Taken when it carries the renderer's own `(at <location>)`, or when protocol
was already cut from this reply, which makes a bare scene line part of the
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
no protocol in it stays, and so does any scene line with story after it.
"""
lines = text.rstrip().split("\n")
if len(lines) < 2:
return text
last = lines[-1].strip()
if not last.startswith(render.HEADING_SCENE + " "):
return text
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
return text
return "\n".join(lines[:-1]).rstrip()
def explain_removed_line(line: str) -> str | None:
"""Which v1.1 rule removes a line of this shape, for the replay report.
None means no v1.1 rule explains it, which the replay treats as a failure.
"""
stripped = line.strip()
if _EVENT_CALL_LINE_RE.match(line):
return RULE_EVENT_CALL
if stripped.startswith("["):
inner = stripped[1:]
inner = inner[:-1] if inner.endswith("]") else inner
if _is_length_hint(inner):
return RULE_LENGTH_HINT
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
return RULE_INSTRUCTION_TAIL
if stripped.startswith(render.HEADING_SCENE + " "):
return RULE_SCENE_LINE
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
return RULE_EMPTY_FENCE
return None
def _is_state_heading(line: str) -> bool:
return bool(_STATE_HEADING_RE.match(line))
@@ -451,7 +625,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
if matches:
match = matches[-1]
raw = match.group(1).strip()
prose = _clean(text[: match.start()] + text[match.end():])
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
return prose, _tolerant_load(raw), raw
# A `json` or unlabelled fence is ours only when its contents are this
@@ -465,7 +639,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
raw = match.group(1).strip()
parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
prose = _clean(text[: match.start()] + text[match.end():])
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
return prose, parsed, raw
match = _TRAILING_RE.search(text)
@@ -473,7 +647,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
raw = match.group(1)
parsed = _tolerant_load(raw)
if _looks_like_proposal(parsed):
return _clean(text[: match.start()]), parsed, raw
return _clean(text[: match.start()], after_block=True), parsed, raw
# An unfenced proposal on its own lines but not at the end: quoted, or
# followed by more story. The last one is the turn's proposal, as with
@@ -481,7 +655,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
without, found = _inline_proposals(text)
if found:
parsed, raw = found[-1]
return _clean(without), parsed, raw
return _clean(without, after_block=True), parsed, raw
# No block at all — but the reply may still carry protocol the model wrote
# as prose, or a fence it never closed.
+16 -4
View File
@@ -26,6 +26,10 @@ EMBED_READ_TIMEOUT = 60.0
#: v1.1 WP-A1: ask a stream to report its token usage. Without it Ollama sends
#: none, and a prompt the server cut cannot be told from one it read whole.
STREAM_OPTIONS = {"include_usage": True}
# Completion endpoints have no roles, so a chat has to be flattened into one
# labeled transcript that ends on "Assistant:" for the model to continue.
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
@@ -84,10 +88,14 @@ class OpenAICompatibleProvider(Provider):
def _record_usage(self, payload: dict) -> None:
"""Records the endpoint's own token accounting, if it reported any.
OpenRouter now always reports usage, and `usage: {include: true}` and
`stream_options` are deprecated and do nothing. In a stream the usage
arrives on a final chunk that carries no choices, which is why this is
read separately from the text extraction.
In a stream the usage arrives on a final chunk that carries no choices,
which is why this is read separately from the text extraction.
v1.1 WP-A1: Ollama sends that chunk only when asked. Measured on Ollama
0.33: a stream with no `stream_options` carried no usage at all, and not
one of the 514 AI turns in the v1 evidence had a count stored. Every
streaming body therefore sets `stream_options.include_usage`
(`STREAM_OPTIONS`), and the turn compares the count with what it sent.
"""
usage = payload.get("usage")
if isinstance(usage, dict) and usage:
@@ -102,6 +110,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
else:
url = f"{self.base_url}/chat/completions"
@@ -114,6 +123,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
return url, body
@@ -183,6 +193,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
else:
url = f"{self.base_url}/chat/completions"
@@ -192,6 +203,7 @@ class OpenAICompatibleProvider(Provider):
"temperature": temperature,
"max_tokens": max_tokens,
"stream": True,
"stream_options": dict(STREAM_OPTIONS),
}
async for event in self._stream(url, body):
yield event
+34 -2
View File
@@ -28,7 +28,9 @@ repair. Refusing a whole campaign because a search index would not build would
trade the valuable thing for the cheap one.
"""
from fastapi import Body, Depends, Request
import json
from fastapi import Body, Depends, Request, Response
from sqlalchemy.orm import Session
from ... import bundle, head, limits, models, schemas
@@ -46,8 +48,38 @@ def export_adventure(
`app/bundle.py` owns the format, in all three of its versions. A backup
outlives the schema, so no call site decides anything about its shape.
**v1.1 WP-D: the export also says whether this version could import it back.**
A campaign large enough to pass `limits.MAX_IMPORT_BODY_BYTES` still exports —
the file is complete and not damaged, and refusing to write it would destroy
the only copy the reader was trying to make. What it cannot do is come back
in here, and the reader is told that at the moment they take it rather than
at the moment they need it.
It travels in headers, not in the body. The body is the bundle, the browser
saves exactly those bytes as the file, and a warning inside it would become
part of a portable story file and of every checksum taken over one.
The size measured is the compact serialisation, because that is both what
this response sends and what the browser POSTs back on import, which is what
`BodySizeLimitMiddleware` weighs. The pretty-printed file the reader
downloads is larger, and is not what import reads.
"""
return bundle.export(db, adv)
payload = bundle.export(db, adv)
# Serialised exactly as Starlette's JSONResponse would, so the bytes counted
# are the bytes sent.
body = json.dumps(payload, ensure_ascii=False, allow_nan=False,
separators=(",", ":")).encode("utf-8")
limit = limits.MAX_IMPORT_BODY_BYTES
importable = len(body) <= limit
headers = {
"X-Export-Bytes": str(len(body)),
"X-Import-Limit-Bytes": str(limit),
"X-Importable-By-This-Version": "true" if importable else "false",
}
if not importable:
headers["X-Export-Warning"] = limits.oversized_export_warning(len(body), limit)
return Response(content=body, media_type="application/json", headers=headers)
@router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201)
+38 -3
View File
@@ -6,6 +6,7 @@ lock guards one set only while one module owns it. And a test that replaces
`OpenAICompatibleProvider` or `generate_turn` patches this module, which every
caller reads through.
"""
import logging
import threading
from fastapi import Depends, HTTPException, Request
@@ -28,6 +29,8 @@ from .deps import CurrentUser, current_adventure, router
from .nodes import _move_to_after, next_depth
from .paging import annotate_takes
log = logging.getLogger(__name__)
def world_delta_of(snapshot: dict | None) -> dict | None:
"""Returns the bulk-read slice of a context snapshot, for `Action.world_delta`.
@@ -212,8 +215,18 @@ async def _generate_turn(
# network calls — and cached per endpoint and model, so it costs one short
# request per session rather than one per turn. An unverified window does
# not block the turn; it is recorded as unverified in the snapshot below.
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
#
# v1.1 WP-A1 corrective: a model that is not resident cannot report its window,
# and a turn built to the configured budget against it was silently cut in the
# A1 evidence (13,875 tokens sent, 2,050 read). So an unverified window gets
# one bounded attempt to load the model, and one more probe, before the
# prompt is assembled. No story text is generated by it and nothing is
# written. A window still unverified afterwards changes nothing below.
window, preflight = await contextwindow.ensure_window(
settings.endpoint_url, settings.model,
declared=settings.context_window_override,
warm_timeout=float(settings.model_timeout_seconds or 300),
)
try:
system_text, story_text, snapshot = build_context(
adventure,
@@ -232,6 +245,9 @@ async def _generate_turn(
yield turn_error(str(exc))
return
if isinstance(snapshot.get("window"), dict):
snapshot["window"]["preflight"] = preflight
parts = PromptParts(system=system_text, story=story_text)
provider = OpenAICompatibleProvider(
@@ -318,6 +334,24 @@ async def _generate_turn(
# prompt came from cache rather than being billed in full. This is recorded
# per attempt, next to the prompt it priced.
snapshot["usage"] = provider.last_usage
# v1.1 WP-A1: what the server says it read, against what was sent. Recorded
# and shown, never acted on: the narration has already streamed to the
# reader, and discarding an accepted turn over an accounting discrepancy
# would lose story to hide a problem. A server that cut the prompt answers
# 200 either way, so this record is the only place the cut is visible.
tokens = snapshot.get("tokens") or {}
accounting = contextwindow.classify_usage(
provider.last_usage,
estimate=tokens.get("estimate") or tokens.get("total") or 0,
budget=tokens.get("budget") or settings.context_token_budget,
max_output_tokens=settings.max_output_tokens,
window_verified=bool((snapshot.get("window") or {}).get("verified")),
)
snapshot["accounting"] = accounting
if accounting["status"] in (contextwindow.EXCEEDED,
contextwindow.TRUNCATION_SUSPECTED):
log.warning("turn accounting for adventure %s: %s — %s",
adventure.id, accounting["status"], accounting["detail"])
reasoning = "".join(reasoning_chunks).strip() or None
ai_action = models.Action(
@@ -381,7 +415,8 @@ async def _generate_turn(
db.commit()
db.refresh(ai_action)
yield _SAVED
yield sse({"type": "done", "action": action_json(ai_action, db)})
yield sse({"type": "done", "action": action_json(ai_action, db),
"accounting": accounting})
# Phase 6: schedule summarization and embedding without waiting for them.
# The task opens its own database session.
memorybank.schedule_post_turn(adventure)
+21 -4
View File
@@ -246,11 +246,28 @@ def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
assert report_before["history"]["floor_depth"] is not None, (
"this fixture is meant to be over budget; trimming never engaged")
_play_one_more(db, adventure, 60)
_, after, report_after = _builder.build_context(adventure, settings)
# v1.1 WP-A1: the fixture used to be positioned so that the very next turn
# held the floor. The safety reserve takes 256 tokens of this 2,048 budget,
# the block is now the minimum of two, and the next turn is a step. So walk
# forward until a turn holds, requiring every move on the way to be exactly
# one block: a window that slides by one action every turn fails either way.
held = None
depth = 60
for _ in range(4):
_play_one_more(db, adventure, depth)
depth += 1
_, after, report_after = _builder.build_context(adventure, settings)
floor_before = report_before["history"]["floor_depth"]
floor_after = report_after["history"]["floor_depth"]
block = report_after["history"]["trim_block"]
assert floor_after - floor_before in (0, block), (floor_before, floor_after, block)
if floor_after == floor_before:
held = (before, after)
break
before, report_before = after, report_after
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
assert _shared_prefix(before, after) > 0.85
assert held is not None, "the floor never held across a turn"
assert _shared_prefix(*held) > 0.85
def test_without_a_stable_floor_the_prefix_collapses(saturated):
@@ -0,0 +1,86 @@
"""v1.1 WP-B.1: the long run's `recovered_through_memory_independent` verdict.
The new verdict must never be reported when anything other than memory could
have carried the fact. Each precondition is named when it fails. The existing M04
verdicts keep their meaning exactly.
python -m pytest tests/test_v11_b1_long_run_verdict.py -v
"""
import pytest
from tools import m11_long_run as lr
GOOD = {
"independent_planted_depth": 3,
"planted_turn_outside_history": True,
"absent_from_state": True,
"absent_from_summary": True,
"absent_from_knowledge": True,
"absent_from_later_narration": True,
"memory_covering_planting_carries_fact": True,
"memory_forgotten": False,
"memory_injected": True,
}
def test_every_precondition_and_an_injected_memory_is_the_new_verdict():
assert lr._independent_memory_verdict(GOOD) == "recovered_through_memory_independent"
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
def test_a_failed_precondition_is_named_and_never_a_recovery(name):
assert lr._independent_memory_verdict({**GOOD, name: False}) == f"precondition_failed:{name}"
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
def test_an_unmeasured_precondition_is_unknown_not_a_pass(name):
assert lr._independent_memory_verdict({**GOOD, name: None}) == f"precondition_unknown:{name}"
def test_no_planted_depth_is_unknown():
assert lr._independent_memory_verdict({**GOOD, "independent_planted_depth": None}) == \
"precondition_unknown:planted_depth"
@pytest.mark.parametrize("change, verdict", [
({"memory_covering_planting_carries_fact": False}, "not_recovered:not_created"),
({"memory_forgotten": True}, "not_recovered:evicted"),
({"memory_injected": False}, "not_recovered:not_injected"),
])
def test_the_failing_memory_stage_is_named(change, verdict):
assert lr._independent_memory_verdict({**GOOD, **change}) == verdict
def test_preconditions_are_judged_before_memory():
"""A carried fact disqualifies the run even when memory also failed."""
both = {**GOOD, "absent_from_state": False, "memory_covering_planting_carries_fact": False}
assert lr._independent_memory_verdict(both) == "precondition_failed:absent_from_state"
def test_the_fact_is_matched_as_whole_words():
assert lr._mentions_fact("She hid the amber Sundial.")
assert lr._mentions_fact("a cracked TEAPOT on the shelf")
assert not lr._mentions_fact("teapots") # a different word, not the fact's
assert not lr._mentions_fact("the sun dialled down")
def test_the_m04_verdicts_are_unchanged():
base = {"planted_turn_in_history_window": False, "in_memories_section": False,
"in_summary_section": False, "in_state_section": False}
assert lr._m04_verdict(base) == "not_recovered"
assert lr._m04_verdict({**base, "in_state_section": True}) == "recovered_through_state_only"
assert lr._m04_verdict({**base, "in_memories_section": True}) == \
"recovered_through_memory_or_summary"
assert lr._m04_verdict({**base, "planted_turn_in_history_window": True}) == \
"precondition_not_met"
def test_the_independent_fact_is_not_in_any_imported_knowledge_file():
for text in (lr.CANON_MD, lr.REFERENCE_MD, lr.INSPIRATION_MD, *lr.BEATS):
assert not lr._mentions_fact(text)
def test_the_planting_text_and_recall_carry_the_fact():
assert lr._mentions_fact(lr.INDEPENDENT_FACT_TEXT)
assert lr._mentions_fact(lr.INDEPENDENT_RECALL_TEXT)
@@ -0,0 +1,434 @@
"""v1.1 WP-B.1: the memory-retention diagnostic, deterministically.
B.1 changes no memory behaviour. These tests prove two things about the
diagnostic in `tools/memory_diagnostic.py`:
1. **It measures what it claims.**
- The fixture keeps the planted fact out of every layer except memory.
- Each stage (created, retained, ranked, injected) is reported from the rows
and the recall turn's own stored context.
- Its ranking agrees with the selection production stored.
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
one that keeps a fact if and only if the fact reached it. Any failure is
therefore the application's mechanism, not a model's writing.
- WP-B.1 marked the criteria v1.0.0 did not meet `xfail(strict=True)`. WP-B.2
fixed ranking, eviction and the creation excerpt, and those tests are now
ordinary passes; the v1.0.0 results are recorded in the WP-B.2 report.
- The same file is run unchanged against v1.0.0 for the baseline.
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
"""
import asyncio
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, limits, memorybank, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures
from tools import memory_diagnostic as md
_results: dict = {}
def scenario(name: str) -> dict:
"""Runs a named scenario once per session and keeps the result."""
if name not in _results:
_results[name] = md.run_scenario(md.SCENARIOS[name])
return _results[name]
# ------------------------------------------------------- fixture preconditions
def test_the_fact_is_planted_early_and_recalled_past_depth_one_hundred():
result = scenario("independent_default")
assert result["plant_depth"] is not None and result["plant_depth"] <= 3
assert result["recall_depth"] >= 100
@pytest.mark.parametrize("check", ["state_document", "state_snapshots", "later_narration",
"summary", "knowledge", "recent_history", "state_section"])
def test_no_layer_but_memory_carries_the_fact(check):
"""A test where another layer carries F is not evidence about memory."""
isolation = scenario("independent_default")["isolation"]
assert isolation["checks"][check]["ok"], isolation["checks"][check]
assert isolation["ok"]
def test_the_isolation_check_fails_when_another_layer_carries_the_fact():
"""The negative control for the precondition itself: a state fact naming F."""
fact = md.FACT_F
with SessionLocal() as db:
Base.metadata.create_all(bind=engine)
try:
user = models.User(is_guest=False, email="b1-iso@example.com")
db.add(user)
db.flush()
adventure = models.Adventure(user_id=user.id, title="iso")
adventure.narrative_state = {"facts": [{"id": "x", "predicate": "hidden",
"value": "the amber sundial is in the teapot"}]}
db.add(adventure)
db.commit()
result = md.isolation(db, adventure, fact, 1)
assert result["ok"] is False
assert result["checks"]["state_document"]["ok"] is False
finally:
db.close()
Base.metadata.drop_all(bind=engine)
# ------------------------------------------------------------------- stages
def test_creation_is_reported_with_the_covering_memory_and_what_the_summariser_saw():
created = scenario("independent_default")["diagnosis"]["created"]
assert created["yes"] is True
assert created["source_start"] <= scenario("independent_default")["plant_depth"] <= created["source_end"]
assert md.FACT_F.carried_by(created["memory_text"])
covering = [c for c in created["covering_memories"] if c["memory_id"] == created["memory_id"]]
assert covering and covering[0]["fact_in_block"] and covering[0]["fact_in_summariser_excerpt"]
def test_retention_is_reported_with_the_bank_and_its_eviction_order():
retained = scenario("independent_default")["diagnosis"]["retained"]
assert retained["yes"] is True and retained["forgotten"] is False
assert retained["on_active_lineage"] is True
assert retained["active_memories"] <= retained["memory_bank_capacity"]
assert retained["eviction_position"] is not None
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
ranked = scenario("independent_default")["diagnosis"]["ranked"]
assert ranked["replica_matches_stored_selection"] is True
assert ranked["top_k_cutoff"] == 5
assert ranked["yes"] is True and ranked["selected"] is True
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
# v1.1 WP-B.2: every part of the score is reported, and they add up.
assert 0.0 <= ranked["lexical_score"] <= 1.0
assert ranked["final_score"] == pytest.approx(
ranked["semantic_score"] + memorybank.LEXICAL_WEIGHT * ranked["lexical_score"], abs=2e-4)
assert ranked["query"]["input"].endswith(md.SCENARIOS["independent_default"].recall_text)
def test_injection_is_read_from_the_recall_turns_own_context():
diagnosis = scenario("independent_default")["diagnosis"]
assert diagnosis["injected"]["yes"] is True
assert diagnosis["injected"]["context_component"] == md.MEMORIES_LABEL
assert diagnosis["injected"]["token_count"] > 0
assert diagnosis["verdict"] == "injected"
def test_ranking_variants_direct_paraphrase_and_unrelated():
variants = scenario("independent_default")["ranking_variants"]
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
assert (variants["direct"]["final_score"] > variants["paraphrase"]["final_score"]
> variants["unrelated"]["final_score"])
def test_retrieval_still_fills_top_k_whatever_the_similarity():
"""There is still no relevance floor: an unrelated question selects a full
`memory_top_k`. B.2 changed which memories those are, not how many — the
early fact is no longer carried along by an unrelated question."""
variants = scenario("independent_default")["ranking_variants"]
assert variants["unrelated"]["selected_count"] == 5
assert variants["unrelated"]["selected"] is False
assert variants["unrelated"]["rank"] > 5
# ------------------------------------------------- WP-B.2 ranking acceptance
@pytest.mark.parametrize("name", ["ranking_crowded", "ranking_context_dependent"])
def test_acceptance_the_early_memory_is_ranked_and_injected_below_capacity(name):
"""The B.1 ranking failure, made deterministic. On v1.0.0 both fixtures are
`retained_but_not_ranked` (ranks 7 and 6 of 17 against a top-k of 4)."""
result = scenario(name)
assert result["isolation"]["ok"], result["isolation"]
diagnosis = result["diagnosis"]
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
assert diagnosis["retained"]["active_memories"] <= diagnosis["retained"]["memory_bank_capacity"]
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["rank"] <= 4
assert diagnosis["ranked"]["replica_matches_stored_selection"] is True
assert diagnosis["injected"]["yes"] is True
assert diagnosis["verdict"] == "injected"
def test_the_crowded_fixture_is_won_by_the_players_question():
ranked = scenario("ranking_crowded")["diagnosis"]["ranked"]
assert ranked["rank"] == 1
assert ranked["lexical_score"] > 0 # "sundial" and "amber" are in the question
def test_a_paraphrase_is_found_by_meaning_not_by_shared_words():
"""Lexical matching must not replace semantic retrieval. The paraphrase
shares none of F's distinctive words, yet ranks first."""
for name in ("ranking_crowded", "independent_default"):
paraphrase = scenario(name)["ranking_variants"]["paraphrase"]
assert paraphrase["rank"] == 1 and paraphrase["selected"]
# Only "Mara" is shared, which is far less than the direct question holds.
direct = scenario(name)["ranking_variants"]["direct"]
assert paraphrase["lexical_score"] < direct["lexical_score"] / 2
def test_a_context_dependent_question_needs_the_scene():
""""I ask her what she keeps up there" names nothing F's memory holds. The
scene the last narration set up (Mara, the top shelf, a kettle) is what
finds it; without that context it ranks last."""
result = scenario("ranking_context_dependent")
ranked = result["diagnosis"]["ranked"]
assert ranked["lexical_score"] == 0.0
assert ranked["rank"] <= 4
assert "top shelf" in ranked["query"]["context"]
assert result["ranking_variants"]["input_only"]["rank"] > 4
def test_an_unrelated_rare_word_does_not_outrank_the_relevant_memory():
"""Negative control: the paraphrase plus a place only one other memory
holds. The decoy gains lexical score, and still ranks below F."""
for name in ("ranking_crowded", "independent_default"):
control = scenario(name)["ranking_variants"]["rare_word_with_paraphrase"]
assert control["decoy_lexical_score"] > control["lexical_score"]
assert control["rank"] == 1
assert control["decoy_rank"] > control["rank"]
def test_common_words_contribute_nothing():
common = scenario("independent_default")["ranking_variants"]["common_words"]
assert common["lexical_score"] == 0.0
# ---------------------------------------------------------- capacity/eviction
@pytest.mark.parametrize("name", ["past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"])
def test_past_capacity_the_early_memory_is_retained(name):
"""v1.1 WP-B.2. On v1.0.0 all three are `created_but_evicted`: F was the
least recently used row once recent narration stopped retrieving it, and
went first (turns 21, 21 and 36). Coverage-first eviction keeps the only
memory of the opening, so it stays active and is recalled at depth 106."""
result = scenario(name)
assert result["isolation"]["ok"], result["isolation"]
diagnosis = result["diagnosis"]
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
assert result["eviction"]["f_evicted_at_turn"] is None
# The bank really was past capacity, and stayed bounded.
assert result["eviction"]["first_eviction_turn"] is not None
assert all(t["active"] <= result["scenario"]["capacity"] + (1 if result["scenario"]["pin_first_memory"] else 0)
for t in result["trace"])
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
def test_past_capacity_the_bank_still_describes_the_whole_story():
"""What the rule buys in general, not only for F: the active bank reaches
from the opening to the newest block, and no stretch between them goes
undescribed for more than twice the average spacing a bank of this capacity
can afford (story span / capacity). On v1.0.0 these banks began at depths 36
and 18: the opening was simply gone."""
for name in ("past_capacity", "past_capacity_low_top_k"):
result = scenario(name)
cover = result["trace"][-1]["coverage"]
assert cover["first_start"] == 0
assert cover["last_end"] >= result["recall_depth"] - 2 * memorybank.MEMORY_INTERVAL
assert cover["largest_gap"] <= 2 * (cover["last_end"] + 1) / result["scenario"]["capacity"]
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
"""The frozen-bank regression the v1.0.0 rule fixed, still holding."""
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
def test_a_pinned_memory_survives_capacity():
eviction = scenario("past_capacity_pinned")["eviction"]
assert eviction["pinned_memory_id"] is not None
assert eviction["pinned_memory_forgotten"] is False
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
"""Was `xfail(strict=True)` in WP-B.1; B.2 fixed the eviction rule."""
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
assert scenario(name)["diagnosis"]["verdict"] == "injected"
# ---------------------------------------------------------- creation window
def test_a_fact_early_in_a_long_block_now_reaches_the_summariser():
"""v1.1 WP-B.2. On v1.0.0 this block (2,079 tokens) was cut to its last
2,000, the fact at its start was never seen, and the stage was
`not_created`. The excerpt is now the block's opening and end."""
result = scenario("long_block_fact_early")
created = result["diagnosis"]["created"]
covering = created["covering_memories"]
assert covering, "the long block must have been summarised"
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert covering[0]["fact_in_block"] is True
assert covering[0]["fact_in_summariser_excerpt"] is True
assert created["yes"] is True
assert created["source_start"] <= result["plant_depth"] <= created["source_end"]
assert memorybank.EXCERPT_OMISSION_MARKER not in created["memory_text"]
def test_the_same_fact_late_in_the_same_sized_block_does():
result = scenario("long_block_fact_late")
covering = result["diagnosis"]["created"]["covering_memories"]
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert covering[0]["fact_in_summariser_excerpt"] is True
assert result["diagnosis"]["created"]["yes"] is True
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
"""Was `xfail(strict=True)` in WP-B.1; B.2 changed the excerpt."""
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
assert scenario("long_block_fact_early")["diagnosis"]["verdict"] == "injected"
# ------------------------------------------ WP-B.2 full deterministic acceptance
def test_acceptance_full_isolation_holds_on_every_turn():
"""`independent_full`: long blocks, a crowded query, a bank past capacity.
F must be carried by memory alone for the whole run, not only at recall."""
result = scenario("independent_full")
assert result["plant_depth"] <= 3 and result["recall_depth"] >= 100
assert not any(t["f_in_state"] for t in result["trace"])
assert not any(t["f_in_summary"] for t in result["trace"])
assert result["isolation"]["ok"], result["isolation"]
for check in ("state_document", "state_snapshots", "later_narration", "summary",
"knowledge", "recent_history", "state_section"):
assert result["isolation"]["checks"][check]["ok"], check
def test_acceptance_full_every_stage_passes_past_capacity_with_long_blocks():
"""On v1.0.0 this fixture fails at creation: every block is over 2,000
tokens, and the fact at the start of the first one is never summarised."""
result = scenario("independent_full")
diagnosis = result["diagnosis"]
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
assert diagnosis["created"]["covering_memories"][0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["replica_matches_stored_selection"]
assert diagnosis["injected"]["yes"]
assert diagnosis["verdict"] == "injected"
def test_acceptance_full_provenance_resolves_to_the_planting_turn():
provenance = scenario("independent_full")["provenance"]
assert provenance["recorded"] is not None
assert provenance["range_covers_plant"] and provenance["matches_row"]
assert provenance["source_block_holds_planting"] is True
assert provenance["recorded"]["authority"] == memorybank.ACCEPTED_STORY
def test_acceptance_full_is_the_long_run_independent_memory_verdict():
"""The same measurements, judged by the long-run tool's own verdict."""
from tools import m11_long_run as lr
result = scenario("independent_full")
checks = result["isolation"]["checks"]
diagnosis = result["diagnosis"]
verdict = lr._independent_memory_verdict({
"independent_planted_depth": result["plant_depth"],
"planted_turn_outside_history": checks["recent_history"]["ok"],
"absent_from_state": checks["state_document"]["ok"] and checks["state_snapshots"]["ok"]
and not any(t["f_in_state"] for t in result["trace"]),
"absent_from_summary": checks["summary"]["ok"]
and not any(t["f_in_summary"] for t in result["trace"]),
"absent_from_knowledge": checks["knowledge"]["ok"],
"absent_from_later_narration": checks["later_narration"]["ok"],
"memory_covering_planting_carries_fact": diagnosis["created"]["yes"],
"memory_forgotten": not diagnosis["retained"]["yes"],
"memory_injected": diagnosis["injected"]["yes"],
})
assert verdict == "recovered_through_memory_independent"
# ------------------------------------------------------- lineage control (G)
def test_an_abandoned_lines_memory_is_stored_but_never_eligible_or_injected():
g = scenario("lineage_control")["lineage_control"]
assert g["memory_ids"], "G's memory must exist on line A before it is abandoned"
assert sorted(g["stored"]) == sorted(g["memory_ids"])
assert g["eligible_on_active_line"] == []
assert g["g_text_ever_in_used_memories"] is False
# Any turn that did name G's memory was on line A, before the divergence.
assert g["eligible_after_returning_to_line_a"] == g["memory_ids"]
def test_the_lineage_scenario_still_diagnoses_f_on_the_active_line():
result = scenario("lineage_control")
assert result["isolation"]["ok"], result["isolation"]
assert result["diagnosis"]["verdict"] == "injected"
# ----------------------------------------------------- authority control
@pytest.fixture()
def authority_client(monkeypatch):
embedder = md.ConceptEmbedder()
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
with SessionLocal() as db:
user = models.User(is_guest=False, email="b1-auth@example.com")
db.add(user)
db.flush()
db.add(models.Settings(user_id=user.id, model="script",
endpoint_url="http://127.0.0.1:9/v1",
embedding_model="concept-embed", memory_top_k=5))
adventure = models.Adventure(user_id=user.id, title="auth", memory_bank_enabled=True,
auto_summarize=True)
db.add(adventure)
db.flush()
db.add(models.Action(adventure_id=adventure.id, type="start", text="The tavern at dusk."))
db.commit()
adv, user_id = adventure.id, user.id
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", md.ScriptNarrator)
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: embedder)
monkeypatch.setattr(memorybank, "summary_provider", lambda s: md.BestCaseSummariser())
monkeypatch.setattr(memorybank, "schedule_post_turn", lambda a: None)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id))
client = TestClient(app)
client.adv = adv
try:
yield client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
def test_a_memory_that_contradicts_state_loses_and_changes_nothing(authority_client):
client, adv = authority_client, authority_client.adv
corrected = client.post(f"/api/adventures/{adv}/state/corrections", json={"events": [
{"type": "add_fact", "predicate": "the tavern lamp is lit", "fact_id": "lamp-lit"}]})
assert corrected.status_code in (200, 201), corrected.text[:300]
made = client.post(f"/api/adventures/{adv}/memories",
json={"text": "The tavern lamp was never lit that night."})
assert made.status_code == 201, made.text[:300]
client.patch(f"/api/adventures/{adv}/memories/{made.json()['id']}", json={"pinned": True})
asyncio.run(memorybank.run_post_turn(adv)) # embed it
before = client.get(f"/api/adventures/{adv}/state").json()["document"]
md.ScriptNarrator.next_reply = 'The fire crackles.\n```state\n{"events": []}\n```'
played = client.post(f"/api/adventures/{adv}/actions",
json={"type": "do", "text": "I look at the lamp."})
assert played.status_code == 200 and '"type": "error"' not in played.text
after = client.get(f"/api/adventures/{adv}/state").json()["document"]
assert after == before # retrieval mutated no state
with SessionLocal() as db:
action = (db.query(models.Action).filter_by(adventure_id=adv, type="ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
snapshot = action.context_snapshot
state_text = md._section(snapshot, md.STATE_LABEL)
memory_text = md._section(snapshot, md.MEMORIES_LABEL)
assert "the tavern lamp is lit" in state_text
assert "never lit" in memory_text
assert memory_text.startswith("Memories from earlier in the story")
labels = [s["label"] for s in snapshot["sections"]]
# State is read last of the live sections: it settles the conflict.
assert labels.index(md.STATE_LABEL) > labels.index(md.MEMORIES_LABEL)
@@ -0,0 +1,277 @@
"""v1.1 WP-B.2 (B2.2): which memory a full bank lets go of.
v1.0.0 evicted the least recently used memory. B.1 showed that this discards
the only memory of an early stretch first, because retrieval follows the present
scene and nothing recent resembles it. `memorybank.eviction_order` now thins the
bank where it is densest and keeps the opening and the newest stretch, with
recency as the tie-break and least-recently-used as the fallback.
The scenario-level tests (the planted fact kept past capacity) are in
`test_v11_b1_memory_diagnostic.py`; the v1.0.0 eviction tests in
`test_memory_retrieval.py` still pass unchanged, because their memories carry
no source range and take the fallback.
python -m pytest tests/test_v11_b2_memory_eviction.py -v
"""
import random
from collections import namedtuple
from datetime import datetime, timedelta
import pytest
from sqlalchemy import select
from app import memorybank, models, tree
from app.context import lineage
from app.database import Base, SessionLocal, engine
T0 = datetime(2026, 1, 1, 12, 0, 0)
Row = namedtuple("Row", "id pinned source_start source_end last_used_at created_at use_count")
def row(id, start, end=None, *, pinned=False, used=None, created=None, uses=0):
"""A memory as eviction sees it. Times are minutes after T0."""
return Row(id, pinned, start, (start + 5) if end is None and start is not None else end,
None if used is None else T0 + timedelta(minutes=used),
T0 + timedelta(minutes=id if created is None else created), uses)
def blocks(n, *, first_id=1):
return [row(first_id + i, 6 * i) for i in range(n)]
def largest_gap_from_opening(rows):
"""The longest uncovered run of depths from depth 0 to the last memory."""
ordered = sorted((r.source_start, r.source_end) for r in rows)
gaps = [ordered[0][0]]
reach = ordered[0][1]
for start, end in ordered[1:]:
gaps.append(max(0, start - reach - 1))
reach = max(reach, end)
return max(gaps)
# ------------------------------------------------------------ the pure order
def test_the_opening_and_the_newest_memory_are_kept():
bank = blocks(7)
doomed = memorybank.eviction_order(bank, 5)
assert bank[0].id not in doomed and bank[-1].id not in doomed
assert len(doomed) == 5
def test_the_densest_stretch_is_thinned_first():
# Memories every 6 depths to 30, then a sparse stretch. Removing one of the
# dense ones leaves a 6-depth hole; removing a sparse one leaves far more.
bank = [row(1, 0), row(2, 6), row(3, 12), row(4, 18), row(5, 60), row(6, 120), row(7, 180)]
assert memorybank.eviction_order(bank, 1)[0] in {2, 3, 4}
assert set(memorybank.eviction_order(bank, 2)) <= {2, 3, 4}
def test_a_stretch_two_memories_describe_loses_one_of_them_first():
"""A shared start (a re-played stretch, or a sibling line) leaves no hole.
Of the two, the less recently used goes, even though a unique memory
elsewhere is older and less used than both."""
bank = [row(1, 0), row(2, 6, used=5), row(3, 12, used=50), row(4, 12, used=40), row(5, 18)]
assert memorybank.eviction_order(bank, 1) == [4]
def test_equal_holes_fall_to_the_least_recently_used():
bank = [row(1, 0), row(2, 6, used=30), row(3, 12, used=10), row(4, 18, used=20), row(5, 24)]
assert memorybank.eviction_order(bank, 1) == [3]
def test_a_newborn_can_be_the_legitimate_first_to_go():
"""The frozen bank is about a newborn losing to a count it cannot have yet.
A newborn that only repeats a stretch another memory describes, one used
after it was written, is legitimately the first to go."""
bank = [row(1, 0), row(2, 6, used=100), row(3, 12), row(4, 6, created=90)]
assert memorybank.eviction_order(bank, 1) == [4]
def test_the_newest_memory_is_not_evicted_by_the_bank_it_joins():
"""The frozen-bank regression under the new rule: every older memory has
been used, the newborn never has, and it still stays."""
bank = [row(i, 6 * (i - 1), used=200 + i, uses=3) for i in range(1, 6)]
newborn = row(6, 30, created=300)
assert newborn.id not in memorybank.eviction_order(bank + [newborn], 1)
def test_pins_are_never_taken_but_still_count_as_coverage():
bank = [row(1, 0), row(2, 6, pinned=True), row(3, 12), row(4, 18, pinned=True), row(5, 24)]
doomed = memorybank.eviction_order(bank, 10)
assert not {2, 4} & set(doomed)
# With the pins covering 6 and 18, memory 3's hole is only its own block.
assert doomed[0] == 3
def test_memories_without_a_range_take_the_least_recently_used_fallback():
hand_written = [row(1, None, None, used=30), row(2, None, None, used=10),
row(3, None, None, used=20)]
assert memorybank.eviction_order(hand_written, 3) == [2, 3, 1]
def test_the_fallback_is_used_only_once_no_interior_memory_remains():
bank = [row(1, 0, used=1), row(2, 6, used=90), row(3, 12, used=2),
row(10, None, None, used=0)]
order = memorybank.eviction_order(bank, 4)
assert order[0] == 2 # the interior memory, although recently used
assert order[1:] == [10, 1, 3] # then least recently used
def test_the_order_does_not_depend_on_row_order():
bank = [row(i, 6 * (i - 1), used=(i * 37) % 11, uses=i % 3) for i in range(1, 30)]
bank += [row(40, 12), row(41, 12)] # a shared start with identical timestamps
expected = memorybank.eviction_order(bank, 20)
for seed in range(5):
shuffled = bank[:]
random.Random(seed).shuffle(shuffled)
assert memorybank.eviction_order(shuffled, 20) == expected
def test_a_tie_on_every_signal_is_broken_by_id():
bank = [row(1, 0), row(9, 6, created=0), row(4, 12, created=0), row(20, 18)]
assert memorybank.eviction_order(bank, 1) == [4]
@pytest.mark.parametrize("seed", range(8))
def test_capacity_holds_and_pins_survive_for_any_bank(seed):
rng = random.Random(seed)
bank = []
for i in range(1, rng.randint(2, 60)):
start = None if rng.random() < 0.15 else rng.randrange(0, 400)
bank.append(row(i, start, None if start is None else start + rng.choice([3, 5, 8]),
pinned=rng.random() < 0.1, used=rng.choice([None, rng.randrange(500)]),
uses=rng.randrange(4)))
capacity = rng.randint(1, 30)
overflow = len(bank) - capacity
doomed = memorybank.eviction_order(bank, max(0, overflow))
pinned = {r.id for r in bank if r.pinned}
assert not pinned & set(doomed)
assert len(set(doomed)) == len(doomed)
remaining = len(bank) - len(doomed)
assert remaining == max(capacity, len(pinned)) if overflow > 0 else remaining == len(bank)
@pytest.mark.parametrize("n, capacity, irregular", [(60, 10, False), (500, 80, False), (500, 80, True)])
def test_a_long_bank_keeps_describing_the_whole_story(n, capacity, irregular):
"""The general property, with nothing ever retrieved: memories arrive one
block at a time and the bank is kept at capacity. The opening stays, and no
stretch goes undescribed for more than twice the average spacing. Least
recently used order, on the same arrivals, keeps only the newest stretch."""
rng = random.Random(n)
kept, lru = [], []
depth = 0
for i in range(1, n + 1):
size = rng.choice([4, 6, 6, 9]) if irregular else 6
memory = row(i, depth, depth + size - 1)
depth += size
kept.append(memory)
lru.append(memory)
if len(kept) > capacity:
doomed = set(memorybank.eviction_order(kept, len(kept) - capacity))
kept = [m for m in kept if m.id not in doomed]
lru = sorted(lru, key=lambda m: (m.created_at, m.id))[len(lru) - capacity:]
assert len(kept) == capacity
assert min(m.source_start for m in kept) == 0
assert max(m.id for m in kept) == n
assert largest_gap_from_opening(kept) <= 2 * depth / capacity
assert largest_gap_from_opening(lru) > depth / 2 # v1.0.0 order: the opening is gone
# --------------------------------------------------------- on real rows
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
session = SessionLocal()
try:
yield session
finally:
session.close()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture()
def adventure(db):
user = models.User(is_guest=False, email="b2-evict@example.com")
db.add(user)
db.flush()
settings = models.Settings(user_id=user.id, model="m", embedding_model="e",
memory_bank_capacity=3)
adv = models.Adventure(user_id=user.id, title="Evict", script_state={}, memory_bank_enabled=True)
db.add_all([settings, adv])
db.commit()
adv.settings_row = settings
return adv
def test_the_pass_changes_nothing_but_forgotten(db, adventure):
for i in range(6):
memory = models.Memory(adventure_id=adventure.id, text=f"block {i}",
source_start=6 * i, source_end=6 * i + 5, branch_id=None, depth=6 * i + 5)
db.add(memory)
db.commit()
columns = (models.Memory.id, models.Memory.text, models.Memory.source_start,
models.Memory.source_end, models.Memory.branch_id, models.Memory.depth,
models.Memory.pinned, models.Memory.use_count, models.Memory.last_used_at)
before = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
after = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
assert before == after
active = db.execute(select(models.Memory.id).where(models.Memory.forgotten.is_(False))).scalars().all()
assert len(active) == 3
assert min(active) == min(before) and max(active) == max(before) # the boundaries
def test_eviction_does_not_make_an_abandoned_lines_memory_eligible(db, adventure):
"""Eviction and lineage are separate: the pass decides only `forgotten`, so
a memory on a line the story left is exactly as ineligible afterwards."""
trunk = []
for i in range(4):
action = models.Action(adventure_id=adventure.id, type="ai", text=f"trunk {i}")
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
trunk.append(action)
abandoned_node = models.Action(adventure_id=adventure.id, type="ai", text="the abandoned line")
tree.place_action(db, adventure, abandoned_node)
db.add(abandoned_node)
db.flush()
abandoned = models.Memory(adventure_id=adventure.id, text="on the abandoned line",
source_start=4, source_end=4)
tree.attach_memory(abandoned, abandoned_node)
db.add(abandoned)
db.commit()
# Move the head back and diverge, so the abandoned node is off the path.
adventure.head_depth = trunk[-1].depth
db.commit()
from app import head
head.fork_if_behind_head(db, adventure)
divergent = models.Action(adventure_id=adventure.id, type="ai", text="the new line")
tree.place_action(db, adventure, divergent)
db.add(divergent)
db.flush()
for i, node in enumerate(trunk + [divergent]):
memory = models.Memory(adventure_id=adventure.id, text=f"active {i}",
source_start=node.depth, source_end=node.depth)
tree.attach_memory(memory, node)
db.add(memory)
db.commit()
def eligible():
return set(db.execute(select(models.Memory.id).where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.forgotten.is_(False))).scalars().all())
assert abandoned.id not in eligible()
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
db.expire_all()
assert abandoned.id not in eligible()
assert len(db.execute(select(models.Memory.id).where(
models.Memory.forgotten.is_(False))).scalars().all()) == 3
+189
View File
@@ -0,0 +1,189 @@
"""v1.1 WP-B.2 (B2.3): what the memory summariser is shown of a long block.
v1.0.0 sent the last 2,000 tokens of a block, so a fact early in a longer block
never reached the summariser (B.1 §E). A block that fits is still sent whole. A
longer one is now sent as its opening and its end, with a marker between them,
inside the same 2,000-token budget.
The scenario-level test (the planted fact early in a long block, remembered) is
in `test_v11_b1_memory_diagnostic.py`.
python -m pytest tests/test_v11_b2_memory_excerpt.py -v
"""
import asyncio
import random
import pytest
from app import memorybank, models, tree
from app.context import builder
from app.database import Base, SessionLocal, engine
BUDGET = memorybank.MEMORY_EXCERPT_TOKENS
MARKER = memorybank.EXCERPT_OMISSION_MARKER
FILLER = "The travellers walked the long grey road north past the salt market and the reed beds. "
def words_to_tokens(tokens: int) -> str:
"""Filler at least `tokens` long."""
text = FILLER
while builder.count_tokens(text) < tokens:
text += FILLER
return text
# --------------------------------------------------------------- the excerpt
def test_a_block_that_fits_is_sent_whole_and_unchanged():
raw = words_to_tokens(BUDGET - 200)
assert builder.count_tokens(raw) <= BUDGET
assert memorybank.memory_excerpt(raw) == raw
def test_a_block_of_exactly_the_budget_is_unchanged():
raw = words_to_tokens(BUDGET)
tokens = memorybank._excerpt_encoding().encode(raw)[:BUDGET]
exact = memorybank._excerpt_encoding().decode(tokens)
if builder.count_tokens(exact) == BUDGET:
assert memorybank.memory_excerpt(exact) == exact
def test_a_long_block_keeps_its_opening_and_its_end_in_order():
opening = "Mara slipped the amber sundial inside the cracked teapot. "
ending = "Aldric finally reached the north gate at dawn."
raw = opening + words_to_tokens(3 * BUDGET) + ending
excerpt = memorybank.memory_excerpt(raw)
assert excerpt.startswith(opening)
assert excerpt.endswith(ending)
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
assert tail, "the marker must sit between the two parts"
assert excerpt.index(opening) < excerpt.index(MARKER) < excerpt.index(ending)
def test_the_split_is_even_and_documented():
raw = words_to_tokens(4 * BUDGET)
head_budget, tail_budget = memorybank.excerpt_split(BUDGET)
marker_tokens = builder.count_tokens(f"\n\n{MARKER}\n\n")
assert head_budget + tail_budget + marker_tokens == BUDGET
assert abs(head_budget - tail_budget) <= 1
excerpt = memorybank.memory_excerpt(raw)
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
# Each part is cut as a run of `head_budget` / `tail_budget` tokens. Measured
# on its own, a cut run can come to one token more, because the text either
# side of the cut tokenises differently once it is separated; the hard limit
# is the whole excerpt, tested below.
assert builder.count_tokens(head) <= head_budget + 1
assert builder.count_tokens(tail) <= tail_budget + 1
assert builder.count_tokens(excerpt) <= BUDGET
@pytest.mark.parametrize("extra", [1, 7, 500, BUDGET, 9 * BUDGET])
def test_the_excerpt_never_exceeds_the_budget(extra):
raw = words_to_tokens(BUDGET + extra)
excerpt = memorybank.memory_excerpt(raw)
assert builder.count_tokens(excerpt) <= BUDGET
assert memorybank.memory_excerpt(raw) == excerpt # deterministic
@pytest.mark.parametrize("seed", range(4))
def test_the_budget_holds_for_awkward_text(seed):
"""Token boundaries can merge differently once the parts are rejoined, and
text that is not plain English tokenises unevenly. The budget still holds."""
rng = random.Random(seed)
alphabet = "abcdefghij ÄÖÜ ßé漢字かな 🙂🐉 \n\t.,;:—'\""
raw = "".join(rng.choice(alphabet) for _ in range(12000))
assert builder.count_tokens(memorybank.memory_excerpt(raw)) <= BUDGET
def test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted():
"""The documented limit of a bounded excerpt: head and tail, not everything."""
half = words_to_tokens(3 * BUDGET)
raw = half + "Mara slipped the amber sundial inside the cracked teapot. " + half
assert "sundial" not in memorybank.memory_excerpt(raw)
# ------------------------------------------------------------ memory creation
class EchoSummariser:
"""Returns the whole excerpt it was given as the memory: the worst case for
a marker leaking into stored text."""
def __init__(self):
self.users: list[str] = []
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
self.users.append(user)
return user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
session = SessionLocal()
try:
yield session
finally:
session.close()
Base.metadata.drop_all(bind=engine)
def campaign(db, texts):
user = models.User(is_guest=False, email="b2-excerpt@example.com")
db.add(user)
db.flush()
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
adventure = models.Adventure(user_id=user.id, title="Excerpt", script_state={},
auto_summarize=True, memory_bank_enabled=True)
db.add(adventure)
db.flush()
nodes = []
for i, text in enumerate(texts):
action = models.Action(adventure_id=adventure.id, type="ai" if i % 2 else "do", text=text)
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
nodes.append(action)
db.commit()
return adventure, nodes
def write_memory(db, adventure, monkeypatch, summariser):
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
monkeypatch.setattr(memorybank, "summary_provider", lambda s: summariser)
settings = db.query(models.Settings).first()
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
def test_the_marker_is_never_stored_as_part_of_a_memory(db, monkeypatch):
long = words_to_tokens(800)
adventure, _ = campaign(db, [long] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK))
summariser = EchoSummariser()
[memory] = write_memory(db, adventure, monkeypatch, summariser)
assert MARKER in summariser.users[0] # the summariser was told
assert MARKER not in memory.text # and the memory does not repeat it
assert "[…" not in memory.text and "omitted" not in memory.text
def test_a_short_block_is_prompted_exactly_as_before(db, monkeypatch):
texts = [f"Short action {i}." for i in range(memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK)]
adventure, _ = campaign(db, texts)
summariser = EchoSummariser()
write_memory(db, adventure, monkeypatch, summariser)
block = "\n\n".join(texts[:memorybank.MEMORY_INTERVAL])
assert summariser.users[0] == f"Story excerpt:\n\n{block}\n\nMemory:"
def test_a_long_blocks_memory_keeps_its_source_provenance(db, monkeypatch):
early = "Mara slipped the amber sundial inside the cracked teapot. " + words_to_tokens(900)
texts = [early] + [words_to_tokens(900)] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK - 1)
adventure, nodes = campaign(db, texts)
[memory] = write_memory(db, adventure, monkeypatch, EchoSummariser())
block = nodes[:memorybank.MEMORY_INTERVAL]
assert "sundial" in memory.text # the early fact reached the summariser
assert (memory.source_start, memory.source_end) == (block[0].depth, block[-1].depth)
assert (memory.branch_id, memory.depth) == (block[-1].branch_id, block[-1].depth)
+326
View File
@@ -0,0 +1,326 @@
"""v1.1 WP-B.2 (B2.1): what memory retrieval searches for, and how it scores.
B.1 found the planting-era memory created and retained but ranked out of
`memory_top_k`, because the query was three turns of narration with the player's
question at the end. The query is now the player's input plus a short scene
context, and the score adds one transparent lexical term over the input.
These tests pin the pieces. The end-to-end fixture tests (crowded bank,
context-dependent question, negative controls) are in
`test_v11_b1_memory_diagnostic.py`, beside the diagnostic they use.
python -m pytest tests/test_v11_b2_memory_ranking.py -v
"""
import asyncio
import math
import pytest
from sqlalchemy import event
from app import memorybank, models, tree
from app.context import builder
from app.database import Base, SessionLocal, engine
# ------------------------------------------------------------------ lexical
def test_terms_are_folded_but_not_stemmed():
terms = memorybank.lexical_terms("The tavern's teapots, the glass and the SUNDIAL")
assert {"tavern", "teapot", "glass", "sundial"} <= terms
assert "the" not in terms # the knowledge path's stop list
assert "glas" not in terms # a word ending in "ss" is not a plural
def test_a_word_every_candidate_holds_weighs_nothing():
scores = memorybank.lexical_scores(
frozenset({"travellers"}), {1: frozenset({"travellers", "road"}), 2: frozenset({"travellers"})})
assert scores == {1: 0.0, 2: 0.0}
def test_a_rarer_word_weighs_more_than_a_common_one():
scores = memorybank.lexical_scores(
frozenset({"sundial", "road"}),
{1: frozenset({"sundial"}), 2: frozenset({"road"}), 3: frozenset({"road"}),
4: frozenset({"gate"})})
assert scores[1] > scores[2] == scores[3] > scores[4] == 0.0
def test_a_single_rare_word_of_a_longer_question_is_only_its_share():
"""The share is over the whole question, so one incidental word match
cannot score like a memory that answers it."""
question = frozenset({"where", "amber", "sundial", "fish"})
scores = memorybank.lexical_scores(
question, {1: frozenset({"fish"}), 2: frozenset({"amber", "sundial"}), 3: frozenset({"road"})})
assert 0.0 < scores[1] < scores[2] <= 1.0
assert scores[1] < 0.5
def test_scores_are_bounded_and_empty_inputs_score_zero():
candidates = {1: frozenset({"a1", "b2"}), 2: frozenset({"a1"})}
assert all(0.0 <= v <= 1.0 for v in memorybank.lexical_scores(frozenset({"a1", "b2"}), candidates).values())
assert memorybank.lexical_scores(frozenset(), candidates) == {1: 0.0, 2: 0.0}
assert memorybank.lexical_scores(frozenset({"a1"}), {}) == {}
# ------------------------------------------------------------------ scoring
def _unit(angle):
return [math.cos(angle), math.sin(angle)]
def test_ties_are_broken_by_id_not_by_row_order():
held = {7: [1.0, 0.0], 3: [1.0, 0.0], 5: [1.0, 0.0]}
rows = memorybank.score_candidates([7, 3, 5], held, {}, [1.0, 0.0], None, [])
assert [row[1] for row in rows] == [3, 5, 7]
def test_the_semantic_score_mixes_input_and_context_by_the_fixed_weight():
held = {1: [1.0, 0.0]}
[(final, _, semantic, lexical)] = memorybank.score_candidates(
[1], held, {}, [1.0, 0.0], [0.0, 1.0], [])
assert semantic == pytest.approx(memorybank.INPUT_WEIGHT)
assert final == semantic and lexical == 0.0
# Either part alone is used as it is.
[(_, _, only_context, _)] = memorybank.score_candidates([1], held, {}, None, [0.0, 1.0], [])
assert only_context == pytest.approx(0.0)
@pytest.mark.parametrize("margin, relevant_first", [(0.01, True), (-0.01, False)])
def test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight(margin, relevant_first):
"""The bound that keeps rarity from overruling meaning: a memory more than
`LEXICAL_WEIGHT` behind semantically cannot pass one ahead of it, however
rare the word it shares."""
decoy_cos = 1.0 - memorybank.LEXICAL_WEIGHT - margin
held = {1: [1.0, 0.0], 2: _unit(math.acos(decoy_cos))}
terms = {1: frozenset(), 2: frozenset({"zeppelin"})}
rows = memorybank.score_candidates([1, 2], held, terms, [1.0, 0.0], None, ["zeppelin"])
order = [row[1] for row in rows]
assert (order[0] == 1) is relevant_first
# ---------------------------------------------------- the query, on real rows
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
memorybank._terms_cache.clear()
session = SessionLocal()
try:
yield session
finally:
session.close()
memorybank._vector_cache.clear()
memorybank._terms_cache.clear()
Base.metadata.drop_all(bind=engine)
@pytest.fixture(autouse=True)
def restore_embedding_provider():
real = memorybank.embedding_provider
try:
yield
finally:
memorybank.embedding_provider = real
@pytest.fixture()
def settings(db):
user = models.User(is_guest=False, email="b2-rank@example.com")
db.add(user)
db.flush()
row = models.Settings(user_id=user.id, model="m", embedding_model="stub-embed",
memory_top_k=2, memory_bank_capacity=80)
db.add(row)
db.commit()
return row
SCENE_STATE = {
"entities": {"mara": {"type": "character", "name": "Mara"},
"tavern": {"type": "location", "name": "The Crooked Lantern"}},
"scene": {"summary": "Closing time", "location": "tavern", "present": ["mara"]},
}
def make_adventure(db, settings, texts, state=None):
"""`texts` is `[(type, text)]`, oldest first, each placed on the tree."""
adventure = models.Adventure(user_id=settings.user_id, title="Rank", script_state={},
memory_bank_enabled=True, narrative_state=state or {})
db.add(adventure)
db.flush()
placed = []
for kind, text in texts:
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
placed.append(action)
db.commit()
return adventure, placed
def test_the_query_is_the_players_input_and_the_scene(db, settings):
adventure, _ = make_adventure(db, settings, [
("start", "Rain over the harbour."),
("ai", "Mara wipes down the counter and glances up at the shelf."),
("do", "> You ask Mara about the brass dial."),
], state=SCENE_STATE)
query = memorybank.retrieval_query(adventure)
assert query["input"] == "> You ask Mara about the brass dial."
assert "The Crooked Lantern" in query["context"] and "Mara" in query["context"]
assert "glances up at the shelf" in query["context"]
assert "brass" in query["input_terms"] and "dial" in query["input_terms"]
def test_a_continue_turn_has_no_input_and_searches_by_the_scene(db, settings):
adventure, _ = make_adventure(db, settings, [
("do", "> You sit down."),
("ai", "The fire burns low in the grate."),
])
query = memorybank.retrieval_query(adventure)
assert query["input"] == "" and query["input_terms"] == []
assert "fire burns low" in query["context"]
def test_a_retry_searches_with_the_input_it_is_retrying(db, settings):
adventure, placed = make_adventure(db, settings, [
("ai", "The market is quiet."),
("do", "> You ask about the sundial."),
("ai", "A discarded attempt about lanterns."),
])
query = memorybank.retrieval_query(adventure, exclude_action_id=placed[-1].id)
assert query["input"] == "> You ask about the sundial."
assert "lanterns" not in query["context"]
assert "market is quiet" in query["context"]
def test_the_query_is_bounded_however_long_the_story(db, settings):
long = "The travellers walked the long grey road north past the salt market. " * 400
adventure, _ = make_adventure(db, settings, [
("ai", long), ("story", long)], state=SCENE_STATE)
query = memorybank.retrieval_query(adventure)
assert builder.count_tokens(query["input"]) <= memorybank.QUERY_INPUT_TOKENS
assert builder.count_tokens(query["context"]) <= (
memorybank.QUERY_SCENE_TOKENS + memorybank.QUERY_NARRATION_TOKENS + 2)
# ------------------------------------------------ retrieval, end to end
class SameVector:
"""Every text embeds the same, so only the lexical term separates memories."""
async def embed(self, texts):
return [[1.0, 0.0, 0.0] for _ in texts]
def add_memory(db, adventure, text, **kwargs):
memory = models.Memory(adventure_id=adventure.id, text=text, **kwargs)
db.add(memory)
db.flush()
memorybank.set_vector(memory, [1.0, 0.0, 0.0])
db.commit()
return memory
def retrieve(adventure, settings, **kwargs):
memorybank.embedding_provider = lambda s: SameVector()
return asyncio.run(memorybank.retrieve_memories(adventure, settings, **kwargs))
@pytest.fixture()
def played(db, settings):
adventure, _ = make_adventure(db, settings, [
("ai", "The tavern is warm."),
("do", "> You ask Mara where the amber sundial went."),
], state=SCENE_STATE)
bank = {
"road": add_memory(db, adventure, "Aldric walked the north road."),
"sundial": add_memory(db, adventure, "Mara hid the amber sundial in the teapot."),
"gate": add_memory(db, adventure, "The gate guard asked for a toll."),
}
return adventure, bank
def test_every_used_memory_reports_the_parts_of_its_score(db, settings, played):
adventure, bank = played
result = retrieve(adventure, settings)
first = result["used"][0]
assert first["id"] == bank["sundial"].id
assert first["similarity"] == first["semantic_score"]
assert first["lexical_score"] > 0
assert first["final_score"] == pytest.approx(
first["semantic_score"] + memorybank.LEXICAL_WEIGHT * first["lexical_score"], abs=2e-4)
assert result["query"]["input"] == "> You ask Mara where the amber sundial went."
assert result["query"]["lexical_weight"] == memorybank.LEXICAL_WEIGHT
assert result["query"]["input_weight"] == memorybank.INPUT_WEIGHT
def test_a_pin_is_still_always_used_and_counts_toward_top_k(db, settings, played):
adventure, bank = played
settings.memory_top_k = 1
bank["gate"].pinned = True
db.commit()
used = retrieve(adventure, settings)["used"]
assert [m["id"] for m in used] == [bank["gate"].id]
assert used[0]["pinned"] is True
def memory_text_reads(statements):
return [s for s in statements
if s.lstrip().upper().startswith("SELECT") and "FROM memories" in s
and "memories.text" in s]
@pytest.fixture()
def sql_log():
statements: list[str] = []
def record(conn, cursor, statement, parameters, context, executemany):
statements.append(statement)
event.listen(engine, "before_cursor_execute", record)
try:
yield statements
finally:
event.remove(engine, "before_cursor_execute", record)
def test_memory_text_is_read_once_and_then_held(db, settings, played, sql_log):
adventure, _ = played
retrieve(adventure, settings)
sql_log.clear()
result = retrieve(adventure, settings)
reads = memory_text_reads(sql_log)
# Only the detail read of the memories chosen remains. (Every memory here
# embeds identically, so redundancy suppression keeps just one of them.)
assert len(reads) == 1 and reads[0].count("?") == len(result["used"])
def test_a_continue_turn_reads_no_memory_text_to_rank(db, settings, sql_log):
adventure, _ = make_adventure(db, settings, [("do", "> You wait."), ("ai", "Night falls.")])
for text in ("one", "two", "three"):
add_memory(db, adventure, f"memory {text}")
sql_log.clear()
result = retrieve(adventure, settings)
assert all(m["lexical_score"] == 0.0 for m in result["used"])
assert len(memory_text_reads(sql_log)) == 1 # the top-k detail read only
def test_an_edited_memory_is_matched_on_its_new_text(db, settings, played):
adventure, bank = played
assert retrieve(adventure, settings)["used"][0]["id"] == bank["sundial"].id
# An edit clears the vector (the route calls set_vector(None)); re-embedding
# sets it again. Both go through set_vector, which drops the held terms.
bank["road"].text = "The amber sundial was traded for the road toll."
memorybank.set_vector(bank["road"], None)
memorybank.set_vector(bank["road"], [1.0, 0.0, 0.0])
bank["sundial"].text = "Mara hid a bottle in the cellar."
memorybank.set_vector(bank["sundial"], None)
memorybank.set_vector(bank["sundial"], [1.0, 0.0, 0.0])
db.commit()
assert retrieve(adventure, settings)["used"][0]["id"] == bank["road"].id
@@ -0,0 +1,225 @@
"""v1.1 WP-B.2: the memory summariser, after the rejected B2.4 prompt experiment.
B2.4 tried a memory prompt instructing the model to keep named facts and objects.
Measured against the reference model, it did not correct the creation failure it
was for, and it was not shipped (`V1.1-WP-B2-REPORT.md` §T). The shipped prompt
is v1.0.0's.
This file keeps two kinds of test apart.
**Acceptance tests** gate the tree:
- the shipped memory prompt is exactly v1.0.0's, so the experiment is gone;
- every fidelity fixture reaches the summariser whole, through the application's
own prompt assembly;
- a long memory is stored as the model wrote it, never cut;
- the memory the attempt-2 block should have produced ranks first under B2.1.
**Diagnostic-measurement tests** check only that `tools/memory_fidelity.py`
measures correctly: fact retention, attribution, invention, word count, a leading
"Memory:", second person and promise retention, on hand-written memories whose
answers are known. What a real model scores on those measurements is
nondeterministic, is taken with inference, and is reported. It is never a gate
here.
python -m pytest tests/test_v11_b2_summarizer_fidelity.py -v
"""
import asyncio
import re
import subprocess
import pytest
from app import memorybank, models, tree
from app.database import Base, SessionLocal, engine
from tools import memory_diagnostic as md
from tools import memory_fidelity as mf
# ==================================================================== acceptance
def test_the_shipped_memory_prompt_is_v1_0_0s():
"""The B2.4 experiment is reverted: production sends the prompt v1.0.0 and
WP-B.1 shipped, unchanged."""
try:
source = subprocess.run(["git", "show", "beb17ad:backend/app/memorybank.py"],
capture_output=True, text=True, check=True).stdout
except (OSError, subprocess.CalledProcessError):
pytest.skip("git history not available")
block = re.search(r"^MEMORY_SYSTEM_PROMPT = \((.*?)^\)$", source, re.S | re.M).group(1)
shipped = eval(f"({block})", {"MEMORY_MAX_WORDS": 50}) # noqa: S307 - our own source
assert memorybank.MEMORY_SYSTEM_PROMPT == shipped
assert memorybank.MEMORY_MAX_WORDS == 50
def test_the_rejected_experiment_is_not_what_ships():
assert mf.B24_EXPERIMENT_PROMPT != memorybank.MEMORY_SYSTEM_PROMPT
assert "Keep each fact with the person it belongs to" not in memorybank.MEMORY_SYSTEM_PROMPT
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
def test_every_fixture_reaches_the_summariser_whole(fixture):
"""Creation can only fail at the model if the fact was sent. Each fixture
fits the excerpt budget, so the whole block is the excerpt."""
user = mf.user_prompt_for(fixture)
assert memorybank.count_tokens(fixture.raw) <= memorybank.MEMORY_EXCERPT_TOKENS
assert f"Story excerpt:\n\n{fixture.raw}\n\nMemory:" in user
assert user.startswith("Cast:\n- " + fixture.protagonist + " — the protagonist.")
class Scripted:
def __init__(self, reply):
self.reply = reply
self.calls: list[tuple[str, str]] = []
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
self.calls.append((system, user))
return self.reply
@pytest.fixture()
def db():
Base.metadata.create_all(bind=engine)
session = SessionLocal()
try:
yield session
finally:
session.close()
Base.metadata.drop_all(bind=engine)
def campaign(db, fixture):
user = models.User(is_guest=False, email="b2-fidelity@example.com")
db.add(user)
db.flush()
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
adventure = models.Adventure(user_id=user.id, title="Fidelity", script_state={}, auto_summarize=True,
persona_name=fixture.protagonist)
db.add(adventure)
db.flush()
nodes = []
for kind, text in fixture.actions + (("ai", "The story moves on."),):
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
tree.place_action(db, adventure, action)
db.add(action)
db.flush()
nodes.append(action)
db.commit()
return adventure, nodes
def write_memory(db, adventure, monkeypatch, provider):
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
monkeypatch.setattr(memorybank, "summary_provider", lambda s: provider)
settings = db.query(models.Settings).first()
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
def test_the_application_sends_the_shipped_prompt_and_the_whole_planting_block(db, monkeypatch):
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
adventure, nodes = campaign(db, fixture)
provider = Scripted(fixture.faithful)
[memory] = write_memory(db, adventure, monkeypatch, provider)
system, user = provider.calls[0]
assert system == memorybank.MEMORY_SYSTEM_PROMPT
assert "> I watch Mara slip the amber sundial inside the cracked teapot" in user
assert (memory.source_start, memory.source_end) == (nodes[0].depth, nodes[5].depth)
def test_an_over_long_memory_is_stored_as_written_never_cut(db, monkeypatch):
"""The word target is an instruction, not a truncation: cutting a memory
after the fact can split or drop exactly the fact it was written to keep."""
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
adventure, _ = campaign(db, fixture)
long_reply = fixture.faithful + " " + " ".join(["They advanced cautiously through the dark."] * 12)
[memory] = write_memory(db, adventure, monkeypatch, Scripted(long_reply))
assert memory.text == long_reply
assert len(memory.text.split()) > 2 * memorybank.MEMORY_MAX_WORDS
def test_a_faithful_regression_memory_ranks_first_for_its_question():
"""If the summariser keeps the fact, B2.1 finds it: the memory the attempt-2
block should have produced, among the memories its bank really held for that
stretch, under production scoring."""
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
stored, _ = fixture.unfaithful[0]
bank = {
1: fixture.faithful,
2: stored,
3: "Aldric, Mara and Edrin advanced through the cold crypt, the silver key heavy in Aldric's hands.",
4: "Aldric told Mara the silver key opens the crypt beneath the Old Abbey.",
5: "Rain kept falling on Westhaven as the travellers walked toward the abbey grounds.",
}
embed = md.ConceptEmbedder.vector
query = {"input": "> I ask Mara quietly where she hid the amber sundial.",
"context": "Aldric and Mara in the Crooked Lantern, rain outside."}
held = {i: embed(t) for i, t in bank.items()}
terms = {i: memorybank.lexical_terms(t) for i, t in bank.items()}
rows = memorybank.score_candidates(list(bank), held, terms, embed(query["input"]),
embed(query["context"]),
sorted(memorybank.lexical_terms(query["input"])))
assert rows[0][1] == 1
assert rows[0][3] > 0
# ======================================================= diagnostic measurements
# These prove the measuring instrument. They say nothing about any model.
def test_the_fixtures_cover_each_measurement_in_more_than_one_genre():
requirements = {f.requirement for f in mf.FIXTURES}
assert {"distinctive object and place", "player-established concrete fact", "promise / commitment",
"attribution", "clutter pressure", "no invention", "multiple concrete facts",
"the actual failed-run block"} <= requirements
assert {"office", "contemporary", "science-fiction-neutral"} <= {f.genre for f in mf.FIXTURES}
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
def test_the_checker_passes_a_faithful_memory(fixture):
result = mf.evaluate(fixture, fixture.faithful)
assert result["passed"], result
assert not result["over_target"] and not result["memory_prefix"] and not result["second_person"]
@pytest.mark.parametrize("fixture, memory, reason", [
(f, memory, reason) for f in mf.FIXTURES for memory, reason in f.unfaithful
], ids=lambda v: v.fixture_id if isinstance(v, mf.Fixture) else None)
def test_the_checker_fails_each_failure_shape(fixture, memory, reason):
result = mf.evaluate(fixture, memory)
assert not result["passed"], result
if reason == "not retained":
assert not result["retained"]
elif reason == "misattributed":
assert result["misattributed"]
elif reason == "invented":
assert result["inventions"]
def test_the_checker_reads_the_stored_attempt_2_memory_as_the_real_failure():
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
stored, _ = fixture.unfaithful[0]
result = mf.evaluate(fixture, stored)
assert result["retained"] is False and result["words"] == 102 and result["over_target"]
@pytest.mark.parametrize("memory, prefix, you", [
("Memory: Dana promised Marcus the lease by Friday.", True, False),
(" memory: Dana promised the lease.", True, False),
("You thanked Marcus and left.", False, True),
("Dana thanked Marcus; your lease is due.", False, True),
("Dana promised Marcus she would bring the signed lease by Friday.", False, False),
])
def test_the_checker_measures_framing(memory, prefix, you):
result = mf.evaluate(mf.FIXTURES_BY_ID["promise_contemporary"], memory)
assert result["memory_prefix"] is prefix
assert result["second_person"] is you
def test_the_checker_measures_promise_retention():
fixture = mf.FIXTURES_BY_ID["promise_contemporary"]
kept = mf.evaluate(fixture, "Dana promised to bring Marcus the signed lease by Friday.")
scenery = mf.evaluate(fixture, "Memory: Dana looked around the empty living room while a dog barked.")
assert kept["facts"]["lease by Friday"]["kept"] and kept["passed"]
assert not scenery["facts"]["lease by Friday"]["kept"] and scenery["memory_prefix"]
@@ -0,0 +1,91 @@
"""v1.1 WP-C: the browser harness's download helpers, without a browser.
`tools/m11_webdriver.wait_for_download` is what decides that an export actually
left the browser as a file. It must never call a download finished because a
file appeared, because it is still being written, or because it is empty — each
of those would make "the export works" a claim the evidence does not support.
python -m pytest tests/test_v11_c_browser_helpers.py -v
"""
import threading
import time
from pathlib import Path
import pytest
from tools import m11_webdriver as wd
def later(seconds, action):
timer = threading.Timer(seconds, action)
timer.start()
return timer
def test_a_finished_file_is_returned(tmp_path):
later(0.2, lambda: (tmp_path / "campaign.json").write_text('{"format": "x"}'))
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05)
assert found == tmp_path / "campaign.json"
def test_a_file_that_was_already_there_is_not_the_download(tmp_path):
(tmp_path / "old.json").write_text("{}")
with pytest.raises(wd.WebDriverError):
wd.wait_for_download(tmp_path, {"old.json"}, timeout=0.6, poll=0.05)
def test_an_empty_file_never_counts(tmp_path):
(tmp_path / "empty.json").write_bytes(b"")
with pytest.raises(wd.WebDriverError):
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
def test_nothing_counts_while_firefox_is_still_writing(tmp_path):
"""Firefox writes `<name>.part` beside the final name until it is done."""
(tmp_path / "campaign.json").write_text('{"format": "x"}')
(tmp_path / "campaign.json.part").write_text("")
with pytest.raises(wd.WebDriverError):
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
later(0.1, lambda: (tmp_path / "campaign.json.part").unlink())
assert wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05).name == "campaign.json"
def test_a_file_that_is_still_growing_is_not_finished(tmp_path):
target = tmp_path / "big.json"
target.write_text("{")
stop = threading.Event()
def grow():
for _ in range(8):
if stop.is_set():
return
with target.open("a") as fh:
fh.write("x" * 100)
time.sleep(0.05)
writer = threading.Thread(target=grow)
started = time.monotonic()
writer.start()
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05, stable_polls=3)
writer.join()
# Returned only once the size held still, so after the last write.
assert found == target
assert target.stat().st_size == 1 + 8 * 100
assert time.monotonic() - started >= 0.4
def test_the_prefs_save_downloads_unasked_to_the_folder_given(tmp_path):
prefs = wd.firefox_download_prefs(tmp_path)
assert prefs["browser.download.folderList"] == 2
assert prefs["browser.download.dir"] == str(tmp_path)
assert prefs["browser.download.useDownloadDir"] is True
assert prefs["browser.download.always_ask_before_handling_new_types"] is False
assert "application/json" in prefs["browser.helperApps.neverAsk.saveToDisk"]
def test_a_download_folder_must_be_under_home():
with pytest.raises(wd.WebDriverError):
wd.require_under_home(Path("/tmp/wp-c-downloads"))
inside = Path.home() / "v11-evidence" / "wp-c" / "downloads"
assert wd.require_under_home(inside) == inside.resolve()
+348
View File
@@ -0,0 +1,348 @@
"""v1.1 WP-A1 corrective: a cold model is loaded, not guessed about.
A1's accounting caught a real cold-model turn: `/api/ps` knew nothing because the
model was not resident, `/api/show` found no `num_ctx`, the window was therefore
unverified, and the prompt was built to the configured 16,384. Ollama loaded the
model at its own 4,096 default, read 2,050 of the 13,875 tokens and answered 200.
Detection was right. The case is also preventable: once the model is loaded its
window is readable. So before an unverified turn is assembled, the application
asks the configured server, once, to load the model (`POST /api/generate` with a
model and no prompt, which Ollama answers with `"done_reason": "load"` and no
text), probes again, and builds the turn to whatever that probe says. A window
still unverified afterwards changes nothing: the configured budget stands and
the post-response accounting still watches for a cut prompt.
python -m pytest tests/test_v11_cold_window.py -v
"""
import asyncio
import httpx
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy import text
from sqlalchemy.orm import undefer
from app import auth, contextwindow, limits, models
from app.context import builder
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers.base import ProviderError
from app.routers import adventures
from fakes import ScriptedProvider
ENDPOINT = "http://127.0.0.1:11434/v1"
MODEL = "qwen2.5:3b-instruct"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
class ColdOllama:
"""The shapes a real Ollama 0.33 returned, with a model that starts cold.
`/api/ps` lists only loaded models. `/api/show` carries no `num_ctx`.
`/api/generate` with no prompt loads the model at `load_window`, exactly as
the real server answered: HTTP 200, `"response": ""`, `"done_reason": "load"`.
"""
def __init__(self, *, loaded=None, load_window=4096, generate_status=200,
report_after_load=True):
self.loaded = dict(loaded or {})
self.load_window = load_window
self.generate_status = generate_status
self.report_after_load = report_after_load
self.requests: list[tuple[str, str, dict | None]] = []
def handler(self, request: httpx.Request) -> httpx.Response:
body = None
if request.content:
import json
body = json.loads(request.content)
self.requests.append((request.method, str(request.url), body))
path = request.url.path
if path == "/api/ps":
return httpx.Response(200, json={"models": [
{"name": name, "model": name, "context_length": tokens}
for name, tokens in self.loaded.items()
]})
if path == "/api/show":
return httpx.Response(200, json={
"model_info": {"qwen2.context_length": 32768}, "parameters": ""})
if path == "/api/generate":
if self.generate_status != 200:
return httpx.Response(self.generate_status, json={"error": "model not found"})
if self.report_after_load:
self.loaded[body["model"]] = self.load_window
return httpx.Response(200, json={
"model": body["model"], "response": "", "done": True, "done_reason": "load"})
return httpx.Response(404)
def paths(self):
return [httpx.URL(url).path for _method, url, _body in self.requests]
@pytest.fixture()
def server(monkeypatch):
def install(fake: ColdOllama):
original = httpx.AsyncClient
def build(*args, **kwargs):
kwargs.pop("verify", None)
return original(*args, transport=httpx.MockTransport(fake.handler), **kwargs)
monkeypatch.setattr(contextwindow.httpx, "AsyncClient", build)
return fake
return install
# ------------------------------------------------------------ ensure_window
def test_a_cold_model_is_loaded_once_and_its_window_verified(server):
fake = server(ColdOllama(loaded={}, load_window=4096))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert (window.tokens, window.source, window.verified) == (4096, contextwindow.LOADED, True)
assert preflight == {"attempted": True, "loaded": True, "verified_before": False,
"verified_after": True,
"detail": "the server loaded the model (load)"}
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate", "/api/ps"]
# One load request, naming the model and nothing else: no prompt, so no text.
warms = [body for _m, url, body in fake.requests if url.endswith("/api/generate")]
assert warms == [{"model": MODEL}]
def test_an_already_loaded_model_is_not_warmed(server):
fake = server(ColdOllama(loaded={MODEL: 16384}))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert window.verified and window.tokens == 16384
assert preflight["attempted"] is False
assert "/api/generate" not in fake.paths()
def test_a_model_that_loads_but_still_cannot_be_read_stays_unverified(server):
"""A server that loads the model but whose `/api/ps` still cannot say. The
existing unknown path stands: no guessed window, the configured budget kept."""
fake = server(ColdOllama(loaded={}, report_after_load=False))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert not window.verified and window.tokens is None
assert preflight["attempted"] is True and preflight["loaded"] is True
assert preflight["verified_after"] is False
assert fake.paths().count("/api/generate") == 1
assert contextwindow.effective_budget(16384, window) == 16384
@pytest.mark.parametrize("status", [404, 500])
def test_a_failed_load_is_recorded_and_leaves_the_window_unverified(server, status):
fake = server(ColdOllama(loaded={}, generate_status=status))
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert not window.verified
assert preflight["attempted"] is True and preflight["loaded"] is False
assert f"HTTP {status}" in preflight["detail"]
# Bounded: one attempt, and no second probe after a failed load.
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate"]
def test_a_declared_window_does_not_stop_the_server_being_asked(server):
"""A declaration fills a hole the server leaves. Loading the model can close
the hole, and a verified answer always wins over a declaration."""
server(ColdOllama(loaded={}, load_window=4096))
window, _preflight = asyncio.run(
contextwindow.ensure_window(ENDPOINT, MODEL, declared=8192))
assert (window.tokens, window.source) == (4096, contextwindow.LOADED)
def test_an_unreachable_server_is_not_asked_to_load_anything():
"""Nothing listens here. No load is attempted against a server that did not
answer the probe, so an offline turn costs no second timeout."""
window, preflight = asyncio.run(
contextwindow.ensure_window("http://127.0.0.1:1/v1", MODEL))
assert not window.verified
assert preflight["attempted"] is False
assert "did not answer" in preflight["detail"]
def test_the_load_request_obeys_the_endpoint_policy():
"""ADR 011. No transport is installed: a load that ignored the policy would
try to reach a public address for real."""
for url in ("https://api.openai.com/v1", "http://8.8.8.8:11434/v1"):
loaded, detail = asyncio.run(contextwindow.warm(url, MODEL, timeout=2))
assert loaded is False
assert "not allowed" in detail
def test_the_load_request_goes_only_to_the_configured_host(server):
fake = server(ColdOllama(loaded={}))
asyncio.run(contextwindow.ensure_window("http://192.168.0.50:11434/v1", MODEL))
hosts = {httpx.URL(url).host for _m, url, _b in fake.requests}
ports = {httpx.URL(url).port for _m, url, _b in fake.requests}
assert hosts == {"192.168.0.50"} and ports == {11434}
# ------------------------------------------------------------- end to end
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="v11cold@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
context_token_budget=16384, max_output_tokens=500,
))
adventure = models.Adventure(
user_id=user.id, title="Cold",
campaign_canon={"rules": ["The sealed crypt is named CANON-SENTINEL-COLD-2050."]},
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(adventure_id=adventure.id, type="start", text="Rain."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def _long_story(adv_id, turns=120):
from app import tree
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv_id)
for i in range(turns):
for kind, body in (
("do", f"I search the {i}th chamber of the undercroft."),
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
):
action = models.Action(adventure_id=adv_id, type=kind, text=body)
db.add(action)
db.flush()
tree.place_action(db, adventure, action)
db.commit()
def _counts():
with SessionLocal() as db:
return {table: db.execute(text(f"SELECT COUNT(*) FROM {table}")).scalar()
for table in ("actions", "state_events", "state_proposals", "memories",
"summaries")}
def _latest_ai(adv_id):
with SessionLocal() as db:
return (db.query(models.Action)
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
def test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports(client, server):
"""The observed failure, prevented. Without the load this turn would be built
to the configured 16,384 against a 4,096 server."""
_long_story(client.adv_id)
fake = server(ColdOllama(loaded={}, load_window=4096))
ScriptedProvider.replies = ["The seal holds."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is True
assert snapshot["tokens"]["budget"] == 4096
assert snapshot["window"]["preflight"]["attempted"] is True
assert snapshot["window"]["preflight"]["verified_after"] is True
system, story = ScriptedProvider.prompts[-1]
sent = builder.count_tokens(system) + builder.count_tokens(story)
assert sent + snapshot["tokens"]["transport"] + 500 + 256 <= 4096
assert "CANON-SENTINEL-COLD-2050" in system
assert fake.paths().count("/api/generate") == 1
def test_without_the_load_the_same_cold_turn_would_have_been_built_too_large(client, server,
monkeypatch):
"""The negative control: v1.0.0 and the first A1 tree probed only."""
_long_story(client.adv_id)
server(ColdOllama(loaded={}, load_window=4096))
async def probe_only(endpoint_url, model, *, declared=None, warm_timeout=300.0):
window = await contextwindow.probe(endpoint_url, model, declared=declared)
return window, {"attempted": False}
monkeypatch.setattr(adventures.turns.contextwindow, "ensure_window", probe_only)
ScriptedProvider.replies = ["The seal holds."]
client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is False
assert snapshot["tokens"]["budget"] == 16384
system, story = ScriptedProvider.prompts[-1]
assert builder.count_tokens(system) + builder.count_tokens(story) > 4096 * 2
def test_the_load_itself_writes_nothing(client, server):
"""No action, narration, state event, proposal, memory or summary comes from
the preflight: it is a request to the server and nothing else."""
fake = server(ColdOllama(loaded={}))
before = _counts()
asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
assert _counts() == before
assert fake.paths().count("/api/generate") == 1
def test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe(client, server):
"""The ordinary failure semantics: the error is reported, no narration is
accepted, and nothing about the state changes."""
server(ColdOllama(loaded={}, generate_status=404))
before = _counts()
ScriptedProvider.replies = [ProviderError("Endpoint or model not found (HTTP 404).")]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the door"})
assert response.status_code == 200
assert '"type": "error"' in response.text or '"error"' in response.text
after = _counts()
assert after["state_events"] == before["state_events"]
assert after["state_proposals"] == before["state_proposals"]
with SessionLocal() as db:
assert db.query(models.Action).filter_by(adventure_id=client.adv_id,
type="ai").count() == 0
def test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer(client, server):
server(ColdOllama(loaded={}, generate_status=500))
ScriptedProvider.replies = ["The door opens."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the door"})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
assert snapshot["window"]["verified"] is False
assert snapshot["window"]["preflight"]["loaded"] is False
assert snapshot["tokens"]["budget"] == 16384
assert snapshot["accounting"]["status"] == contextwindow.UNKNOWN
def test_the_context_dry_run_never_loads_a_model(client, server):
fake = server(ColdOllama(loaded={}))
response = client.get(f"/api/adventures/{client.adv_id}/context")
assert response.status_code == 200
assert "/api/generate" not in fake.paths()
+495
View File
@@ -0,0 +1,495 @@
"""v1.1 WP-A1: a deliberate safety reserve, and a turn the server cut is not silent.
M11 made the verified window a ceiling. It did not make the application's count
the server's count. The application counts with `cl100k_base`, the narrator with
its own tokenizer, and the v1 evidence left 23-42 real tokens between the largest
prompt and the edge of a 16,384 window. Past that edge Ollama does not refuse.
Measured against the reference CPU host (Ollama 0.33, a 4,096 window), a
6,316-token prompt came back 200 with `prompt_tokens` 2,050: the front of the
prompt, which in this design is the narrator's rules and the canon, was gone.
So the tests below are in three halves.
**The reserve.** `max(256, ceil(5% of the effective window))`, taken from the
budget before any history is chosen, on top of an exact reply allocation.
**The arithmetic.** The assembled prompt, plus the application text the provider
adds to every request, plus the reply allocation, plus the reserve, fits the
effective window. Protected context that cannot fit that way fails before the
model is called.
**The accounting.** Where the server reports how many prompt tokens it read, the
turn records `fits`, `exceeded` or `truncation_suspected`. Where it reports
nothing, the turn says `unknown`, never `fits`. A discrepancy found after the
reply is recorded and shown; it never costs the reader an accepted turn.
python -m pytest tests/test_v11_context_reserve.py -v
"""
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, contextwindow, limits, models
from app.context import builder
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.providers.openai_compatible import CHAT_CONTINUE_HINT, OpenAICompatibleProvider
from app.routers import adventures
from fakes import ScriptedProvider
ENDPOINT = "http://127.0.0.1:11434/v1"
@pytest.fixture(autouse=True)
def _clear_window_cache():
contextwindow.cache_clear()
yield
contextwindow.cache_clear()
# ------------------------------------------------------------- the reserve
@pytest.mark.parametrize("window, reserve", [
(1024, 256),
(4096, 256), # 5% is 204.8, so the floor holds
(5120, 256), # exactly 5% is the floor
(5121, 257), # 256.05 rounds up
(8192, 410), # 409.6 rounds up
(16384, 820), # 819.2 rounds up
(32768, 1639), # 1638.4 rounds up
])
def test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up(window, reserve):
assert contextwindow.safety_reserve(window) == reserve
def test_the_reserve_is_far_larger_than_the_v1_margin_at_the_evidence_window():
"""The v1 evidence left 23-42 tokens at 16,384. 64 tokens of slack was all
the arithmetic kept for drift and separators together."""
assert contextwindow.safety_reserve(16384) >= 10 * 64
# ---------------------------------------------------------- the arithmetic
@pytest.fixture()
def client(monkeypatch):
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="v11reserve@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(
user_id=user.id, model="qwen2.5:3b-instruct", endpoint_url=ENDPOINT,
embedding_model="", context_token_budget=16384, max_output_tokens=500,
))
adventure = models.Adventure(
user_id=user.id, title="Reserved",
campaign_canon={"rules": [
"The abbey seal has never been broken.",
"The sealed crypt is named CANON-SENTINEL-RESERVE-5120.",
]},
)
setup.add(adventure)
setup.flush()
setup.add(models.Action(
adventure_id=adventure.id, type="start", text="Rain over Westhaven."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
test_client = TestClient(app)
test_client.adv_id = adv_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
adventures.turns._active_turns.clear()
Base.metadata.drop_all(bind=engine)
def _long_story(adv_id, turns=120):
from app import tree
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv_id)
for i in range(turns):
for kind, text in (
("do", f"I search the {i}th chamber of the undercroft."),
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
):
action = models.Action(adventure_id=adv_id, type=kind, text=text)
db.add(action)
db.flush()
tree.place_action(db, adventure, action)
db.commit()
def _settings(**changes):
with SessionLocal() as db:
settings = db.query(models.Settings).first()
for key, value in changes.items():
setattr(settings, key, value)
db.commit()
def _build(client, window):
with SessionLocal() as db:
adventure = db.get(models.Adventure, client.adv_id)
settings = db.query(models.Settings).first()
return builder.build_context(adventure, settings, window=window)
def _sent(system, story) -> int:
"""What the provider actually sends in chat mode, by the application's count."""
return (builder.count_tokens(system) + builder.count_tokens(story)
+ builder.count_tokens(CHAT_CONTINUE_HINT))
CONFIGURATIONS = {
"verified 4,096": (dict(context_token_budget=16384),
contextwindow.Window(4096, contextwindow.LOADED)),
"verified 8,192": (dict(context_token_budget=16384),
contextwindow.Window(8192, contextwindow.PARAMETERS)),
"verified 16,384": (dict(context_token_budget=16384),
contextwindow.Window(16384, contextwindow.LOADED)),
"declared 6,000": (dict(context_token_budget=16384),
contextwindow.Window(6000, contextwindow.DECLARED)),
"unverified, configured 12,000": (dict(context_token_budget=12000),
contextwindow.UNVERIFIED),
}
@pytest.mark.parametrize("name", list(CONFIGURATIONS))
def test_the_prompt_leaves_the_reply_and_the_reserve_free(client, name):
"""A1-2, on the assembled text rather than the builder's own arithmetic."""
changes, window = CONFIGURATIONS[name]
_settings(**changes)
_long_story(client.adv_id, turns=120)
system, story, report = _build(client, window)
tokens = report["tokens"]
budget = tokens["budget"]
assert tokens["safety_reserve"] == contextwindow.safety_reserve(budget)
assert tokens["output_reserve"] == 500
sent = _sent(system, story)
assert sent + tokens["output_reserve"] + tokens["safety_reserve"] <= budget, (
name, sent, tokens)
# The history is what gave way, not the canon.
assert "CANON-SENTINEL-RESERVE-5120" in system
assert report["history"]["included"] < report["history"]["total"]
def test_the_report_prices_the_text_the_provider_adds(client):
"""The chat hint rides on every request and was never counted."""
_, _, report = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
tokens = report["tokens"]
assert tokens["transport"] >= builder.count_tokens(CHAT_CONTINUE_HINT)
assert tokens["estimate"] == tokens["total"] + builder.count_tokens(CHAT_CONTINUE_HINT)
def test_the_reserve_follows_the_effective_window_not_the_setting(client):
"""5% of a 4,096 server, not 5% of a 16,384 setting it will never read."""
_, _, capped = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
_, _, full = _build(client, contextwindow.Window(16384, contextwindow.LOADED))
assert capped["tokens"]["safety_reserve"] == 256
assert full["tokens"]["safety_reserve"] == 820
def test_protected_context_that_only_fits_without_the_reserve_fails_explicitly(client):
"""A1-3 in the builder. Before v1.1 this prompt would have been built.
The canon is sized so that protected text plus the reply fits a 4,096 window
with room to spare, and does not fit once the 256-token reserve is taken.
"""
small = contextwindow.Window(4096, contextwindow.LOADED)
# Measured with a window large enough never to overflow, because repeated
# text merges tokens at its seams and cannot be priced by multiplication.
roomy = contextwindow.Window(32768, contextwindow.LOADED)
rules = None
with SessionLocal() as db:
adventure = db.get(models.Adventure, client.adv_id)
settings = db.query(models.Settings).first()
base_rules = list(adventure.campaign_canon["rules"])
filler = "The bell tolls once for every name in the ledger."
copies = 1
while True:
candidate = base_rules + [" ".join([filler] * copies)]
adventure.campaign_canon = {"rules": candidate}
_, _, measured = builder.build_context(adventure, settings, window=roomy)
t = measured["tokens"]
# What protected context costs at 4,096, without the reserve.
without_reserve = t["protected"] + t["transport"] + t["output_reserve"]
if without_reserve + 64 >= 4096 - 60:
break
copies += 1
rules = candidate
adventure.campaign_canon = {"rules": rules}
db.commit()
# The case this test is about: v1's arithmetic, with its 64-token margin,
# would have built this prompt. v1.1's reserve does not fit.
assert without_reserve + 64 < 4096
assert without_reserve + contextwindow.safety_reserve(4096) >= 4096
with pytest.raises(builder.ContextOverflow) as caught:
builder.build_context(adventure, settings, window=small)
message = str(caught.value)
assert "safety" in message
assert "load the model with a larger window" in message
def test_an_overflowing_turn_never_reaches_the_model(client, monkeypatch):
"""A1-3 end to end: the refusal happens before the provider is called."""
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(1024, contextwindow.LOADED)
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
ScriptedProvider.replies = ["This must never be generated."]
response = client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "open the crypt"})
assert response.status_code == 200
assert "safety" in response.text
assert ScriptedProvider.calls == 0
with SessionLocal() as db:
assert db.query(models.Action).filter_by(
adventure_id=client.adv_id, type="ai").count() == 0
# ---------------------------------------------------------- the accounting
def _classify(prompt_tokens=None, *, usage=None, estimate=3500, budget=4096,
output=500, verified=True):
if usage is None and prompt_tokens is not None:
usage = {"prompt_tokens": prompt_tokens, "completion_tokens": 40}
return contextwindow.classify_usage(
usage, estimate=estimate, budget=budget, max_output_tokens=output,
window_verified=verified,
)
def test_a_prompt_the_server_read_in_full_fits():
# The 13-token chat-template overhead measured against the real server.
result = _classify(3513)
assert result["status"] == contextwindow.FITS
assert result["server_prompt_tokens"] == 3513
assert result["difference"] == 13
assert result["safety_reserve"] == 256
assert result["observed_margin"] == 4096 - 500 - 3513
def test_a_server_that_counts_more_than_the_reserve_allows_is_exceeded():
"""The prompt plus the reply allocation no longer fits the window."""
result = _classify(3700)
assert result["status"] == contextwindow.EXCEEDED
assert result["observed_margin"] < 0
def test_a_server_that_read_far_less_than_was_sent_is_suspected_of_truncating():
"""The real shape: 6,316 sent, 2,050 read, HTTP 200, no error."""
result = _classify(2050, estimate=6316)
assert result["status"] == contextwindow.TRUNCATION_SUSPECTED
assert result["difference"] == 2050 - 6316
def test_a_small_undercount_is_tokenizer_drift_not_truncation():
"""A tokenizer thriftier than `cl100k_base` reads fewer tokens honestly. Only
a shortfall larger than the reserve is called truncation."""
assert _classify(3500 - 255)["status"] == contextwindow.FITS
assert _classify(3500 - 257)["status"] == contextwindow.TRUNCATION_SUSPECTED
@pytest.mark.parametrize("usage", [
None,
{},
{"completion_tokens": 40},
{"prompt_tokens": 0},
{"prompt_tokens": "3500"},
{"prompt_tokens": -1},
])
def test_no_usable_count_is_unknown_never_fits(usage):
result = _classify(usage=usage)
assert result["status"] == contextwindow.UNKNOWN
assert result["server_prompt_tokens"] is None
assert result["observed_margin"] is None
def test_the_accounting_says_when_the_window_itself_was_not_verified():
result = _classify(3513, verified=False)
assert result["status"] == contextwindow.FITS
assert result["window_verified"] is False
assert "not verified" in result["detail"]
def test_the_stream_asks_the_server_to_report_its_usage():
"""Measured: Ollama 0.33 sends no usage in a stream unless asked."""
provider = OpenAICompatibleProvider(ENDPOINT, "m")
from app.providers.base import PromptParts
for mode in ("chat", "completion"):
provider.api_mode = mode
_url, body = provider._request(PromptParts(system="s", story="t"), 0.7, 50)
assert body["stream"] is True
assert body["stream_options"] == {"include_usage": True}
def _latest_ai(adv_id):
with SessionLocal() as db:
return (
db.query(models.Action)
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first()
)
def _play(client, monkeypatch, usage, window=4096, reply="The crypt is still sealed."):
async def verified(endpoint, model, declared=None, use_cache=True):
return contextwindow.Window(window, contextwindow.LOADED, 32768, "fake")
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
monkeypatch.setattr(ScriptedProvider, "last_usage", usage)
ScriptedProvider.replies = [reply]
return client.post(f"/api/adventures/{client.adv_id}/actions",
json={"type": "do", "text": "look at the seal"})
def test_a_turn_records_what_the_server_read(client, monkeypatch):
response = _play(client, monkeypatch, None)
assert response.status_code == 200, response.text[:300]
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
response = _play(client, monkeypatch,
{"prompt_tokens": estimate + 13, "completion_tokens": 9})
assert response.status_code == 200, response.text[:300]
snapshot = _latest_ai(client.adv_id).context_snapshot
accounting = snapshot["accounting"]
assert accounting["status"] == contextwindow.FITS
assert accounting["server_prompt_tokens"] == estimate + 13
assert accounting["estimate"] == snapshot["tokens"]["estimate"]
assert '"accounting"' in response.text
assert contextwindow.FITS in response.text
def test_a_turn_with_no_reported_usage_is_unknown(client, monkeypatch):
response = _play(client, monkeypatch, None)
assert response.status_code == 200
assert _latest_ai(client.adv_id).context_snapshot["accounting"]["status"] == (
contextwindow.UNKNOWN)
def test_a_suspected_truncation_keeps_the_turn_and_says_so(client, monkeypatch, caplog):
"""A1-7 and A1-8. The reader watched the narration arrive; it stays."""
response = _play(client, monkeypatch, {"prompt_tokens": 12, "completion_tokens": 9},
reply="The seal holds, and the rain goes on.")
assert response.status_code == 200, response.text[:300]
action = _latest_ai(client.adv_id)
assert action is not None
assert action.text == "The seal holds, and the rain goes on."
accounting = action.context_snapshot["accounting"]
assert accounting["status"] == contextwindow.TRUNCATION_SUSPECTED
assert contextwindow.TRUNCATION_SUSPECTED in response.text
assert any(contextwindow.TRUNCATION_SUSPECTED in r.getMessage() for r in caplog.records)
# Inspectable afterwards through the same route the context panel reads.
context = client.get(
f"/api/adventures/{client.adv_id}/actions/{action.id}/context")
assert context.status_code == 200
assert context.json()["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
def test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves():
"""Found by the A2 long run. Accounting belongs to one API call, not to the
turn's shared prompt. A retry demotes the old attempt, and a take selection
hands the prompt from one attempt to another. Neither may drop an attempt's
accounting or give it another attempt's."""
from app import attempts
class Node:
def __init__(self, snapshot):
self.context_snapshot = snapshot
shared = {"tokens": {"estimate": 3000}, "sections": [], "window": {"verified": True}}
first = Node(shared | {"raw_output": "one", "usage": {"prompt_tokens": 3015},
"accounting": {"status": contextwindow.FITS, "server_prompt_tokens": 3015}})
second = Node({"raw_output": "two", "usage": {"prompt_tokens": 12},
"accounting": {"status": contextwindow.TRUNCATION_SUSPECTED,
"server_prompt_tokens": 12}})
# Superseded by a retry: the old attempt keeps only its own slices.
attempts.keep_own_slices(Node(dict(first.context_snapshot)))
demoted = Node(dict(first.context_snapshot))
attempts.keep_own_slices(demoted)
assert demoted.context_snapshot["accounting"]["server_prompt_tokens"] == 3015
assert "tokens" not in demoted.context_snapshot
# The prompt moves to the second attempt; each keeps its own accounting.
attempts.hand_over_the_prompt(first, second)
assert second.context_snapshot["tokens"] == {"estimate": 3000}
assert second.context_snapshot["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
assert second.context_snapshot["accounting"]["server_prompt_tokens"] == 12
assert first.context_snapshot["accounting"]["status"] == contextwindow.FITS
assert "tokens" not in first.context_snapshot
def test_a_retry_leaves_each_take_with_its_own_accounting(client, monkeypatch):
"""End to end, through the real retry route. Before the fix the live take
inherited the superseded take's accounting, so the inspector could show one
call's server count as another's."""
response = _play(client, monkeypatch, None, reply="The first take.")
assert response.status_code == 200, response.text[:300]
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
# Replay the first take with a real count, so it has accounting of its own.
with SessionLocal() as db:
first = (db.query(models.Action)
.filter(models.Action.adventure_id == client.adv_id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot)).first())
snapshot = dict(first.context_snapshot)
snapshot["accounting"] = contextwindow.classify_usage(
{"prompt_tokens": estimate + 15}, estimate=estimate, budget=4096,
max_output_tokens=500, window_verified=True)
first.context_snapshot = snapshot
db.commit()
first_id = first.id
monkeypatch.setattr(ScriptedProvider, "last_usage",
{"prompt_tokens": 12, "completion_tokens": 9})
ScriptedProvider.replies = ["The second take."]
retried = client.post(f"/api/adventures/{client.adv_id}/retry")
assert retried.status_code == 200, retried.text[:300]
with SessionLocal() as db:
rows = (db.query(models.Action)
.filter(models.Action.adventure_id == client.adv_id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id).all())
by_id = {row.id: row for row in rows}
old = by_id[first_id]
new = [row for row in rows if row.id != first_id][-1]
assert new.text == "The second take."
assert new.live and not old.live
# The superseded take keeps its own accounting and gives up the prompt.
assert old.context_snapshot["accounting"]["status"] == contextwindow.FITS
assert old.context_snapshot["accounting"]["server_prompt_tokens"] == estimate + 15
assert "tokens" not in old.context_snapshot
# The live take carries the prompt and its own accounting, not the old one's.
assert "tokens" in new.context_snapshot
assert new.context_snapshot["accounting"]["status"] == (
contextwindow.TRUNCATION_SUSPECTED)
assert new.context_snapshot["accounting"]["server_prompt_tokens"] == 12
def test_an_exceeded_turn_is_also_kept(client, monkeypatch):
response = _play(client, monkeypatch, {"prompt_tokens": 3900, "completion_tokens": 9})
assert response.status_code == 200
action = _latest_ai(client.adv_id)
assert action.text == "The crypt is still sealed."
assert action.context_snapshot["accounting"]["status"] == contextwindow.EXCEEDED
+294
View File
@@ -0,0 +1,294 @@
"""v1.1 WP-D: a backup that was really checked, and an export that says what it is.
Two recovery-path claims, each of which was true only in the small before this
package:
- **A backup is verified.** M9 ran `PRAGMA quick_check` on the finished copy.
That reads every page and every record, and skips the cross-check between a
table and its indexes — so a copy whose index disagrees with its table passed.
`test_the_fixture_is_the_difference_between_the_two_checks` builds exactly that
damage and shows the two pragmas disagreeing about it, before anything here
uses it as evidence.
- **An export is importable.** Nothing compared the bundle with
`limits.MAX_IMPORT_BODY_BYTES`, so a campaign could be exported and then
refused by its own importer, with the reader finding out at the moment they
needed it. The export still succeeds — the file is complete, and a version
that refused to write it would destroy the copy someone was trying to make —
and now it says so.
python -m pytest tests/test_v11_d_recovery.py -v
"""
import json
import sqlite3
from pathlib import Path
import pytest
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, backup, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
@pytest.fixture()
def client():
Base.metadata.create_all(bind=engine)
setup = SessionLocal()
user = models.User(is_guest=False, email="wp-d@example.com")
setup.add(user)
setup.flush()
setup.add(models.Settings(user_id=user.id, model="test-model"))
adventure = models.Adventure(user_id=user.id, title="Recovery")
setup.add(adventure)
setup.flush()
setup.add(models.Action(adventure_id=adventure.id, type="start", text="The story opens."))
setup.commit()
adv_id, user_id = adventure.id, user.id
setup.close()
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id))
test_client = TestClient(app)
test_client.adv_id = adv_id
test_client.user_id = user_id
try:
yield test_client
finally:
app.dependency_overrides.clear()
Base.metadata.drop_all(bind=engine)
# ------------------------------------------------------------ the fixture
def build_corrupt_copy(path: Path) -> None:
"""A database whose index disagrees with its table, and nothing else.
One digit inside one index leaf page is changed, so that entry names a key
no row holds and one row's key is in no index entry. Every page is still
structurally sound and every record still parses, which is the whole point:
this is the damage `quick_check` is not looking for.
"""
path.unlink(missing_ok=True)
connection = sqlite3.connect(path)
connection.execute("PRAGMA page_size=4096")
connection.execute("CREATE TABLE t (id INTEGER PRIMARY KEY, k TEXT NOT NULL, filler TEXT)")
connection.execute("CREATE INDEX i_t_k ON t(k)")
connection.executemany("INSERT INTO t (k, filler) VALUES (?, ?)",
[(f"k{n:06d}", "x" * 40) for n in range(400)])
connection.commit()
page_size = connection.execute("PRAGMA page_size").fetchone()[0]
leaves = [row[0] for row in connection.execute(
"SELECT pageno FROM dbstat WHERE name='i_t_k' AND pagetype='leaf' ORDER BY pageno")]
connection.close()
assert leaves, "the index must have a leaf page to damage"
raw = bytearray(path.read_bytes())
start = (leaves[0] - 1) * page_size
page = raw[start:start + page_size]
at = page.find(b"k000")
assert at != -1, "expected an indexed key on the index's first leaf page"
page[at + 4] = ord("9") # k000144 -> k000944: a key no row has
raw[start:start + page_size] = page
path.write_bytes(bytes(raw))
def check(path: Path, pragma: str) -> str:
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
return ", ".join(str(row[0]) for row in connection.execute(f"PRAGMA {pragma}").fetchall())
finally:
connection.close()
def test_the_fixture_is_the_difference_between_the_two_checks(tmp_path):
"""Before using it as evidence: quick_check calls this database fine."""
damaged = tmp_path / "damaged.db"
build_corrupt_copy(damaged)
assert check(damaged, "quick_check") == "ok"
integrity = check(damaged, "integrity_check")
assert integrity != "ok"
assert "i_t_k" in integrity # it names the index that disagrees
# ---------------------------------------------------------------- backups
def test_a_healthy_backup_passes_the_full_check_and_is_kept(client, tmp_path):
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
result = backup.create(source)
assert result.integrity == "ok"
assert result.path.exists() and result.bytes > 0
assert check(result.path, "integrity_check") == "ok"
assert result.path.parent == backup.directory(source)
def test_the_backup_runs_the_full_check_not_the_quick_one(client, tmp_path, monkeypatch):
"""The pragma itself, named. SQLite traces every statement it executes, so
this reads what the backup actually asked the copy rather than inferring it."""
asked: list[str] = []
real_connect = sqlite3.connect
def tracing(*args, **kwargs):
connection = real_connect(*args, **kwargs)
connection.set_trace_callback(asked.append)
return connection
monkeypatch.setattr(backup.sqlite3, "connect", tracing)
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
backup.create(source).path.unlink()
assert any("integrity_check" in sql for sql in asked), asked
assert not any("quick_check" in sql for sql in asked), asked
def test_a_copy_the_full_check_rejects_is_not_kept(client, tmp_path, monkeypatch):
"""The copy is damaged after it is written and before it is verified, which
is where a real page-level fault would appear: between the copy and the
rename. Nothing wearing a backup's name may be left behind."""
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
real_copy = backup._copy
def damage(source_path, working):
pages = real_copy(source_path, working)
build_corrupt_copy(working)
return pages
monkeypatch.setattr(backup, "_copy", damage)
with pytest.raises(backup.BackupError) as refused:
backup.create(source)
assert "did not verify" in str(refused.value)
assert "i_t_k" in str(refused.value) # it says what was wrong
kept = list(backup.directory(source).glob("*"))
assert kept == [], f"a rejected backup was left behind: {kept}"
def test_a_rejected_backup_leaves_an_earlier_good_one_alone(client, tmp_path, monkeypatch):
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
good = backup.create(source)
before = good.path.read_bytes()
real_copy = backup._copy
def damage(source_path, working):
pages = real_copy(source_path, working)
build_corrupt_copy(working)
return pages
monkeypatch.setattr(backup, "_copy", damage)
with pytest.raises(backup.BackupError):
backup.create(source)
assert good.path.exists()
assert good.path.read_bytes() == before
assert check(good.path, "integrity_check") == "ok"
assert [p.name for p in backup.directory(source).glob("*")] == [good.path.name]
def test_the_backup_file_semantics_are_unchanged(client, tmp_path):
"""Same directory, same stamped name, same reported fields: WP-D changed the
check, not the file."""
source = tmp_path / "campaign.db"
source.write_bytes(Path(str(engine.url.database)).read_bytes())
first = backup.create(source)
second = backup.create(source)
assert first.path.name.startswith(backup.PREFIX) and first.path.suffix == ".db"
assert first.path != second.path, "an existing backup is never overwritten"
assert set(first.as_dict()) == {"filename", "bytes", "pages", "seconds", "integrity"}
listed = [row["filename"] for row in backup.existing(source)]
assert sorted(listed) == sorted([first.path.name, second.path.name])
# ----------------------------------------------------------------- exports
def export(client, adv_id):
response = client.get(f"/api/adventures/{adv_id}/export")
assert response.status_code == 200, response.text[:200]
return response
def test_a_normal_export_carries_no_warning(client):
response = export(client, client.adv_id)
assert "X-Export-Warning" not in response.headers
assert response.headers["X-Importable-By-This-Version"] == "true"
assert int(response.headers["X-Import-Limit-Bytes"]) == limits.MAX_IMPORT_BODY_BYTES
assert int(response.headers["X-Export-Bytes"]) == len(response.content)
assert response.json()["format"] == "ai-dnd-adventure-v3"
def fill_past_the_limit(adv_id: int) -> int:
"""Real rows, until the campaign's bundle is genuinely over the ceiling.
Not a mocked size: the export below serialises all of it.
"""
chunk = "The rain kept on over the harbour road, and nobody came. " * 900 # ~50 kB
written = 0
with SessionLocal() as db:
while written < limits.MAX_IMPORT_BODY_BYTES + 2 * 1024 * 1024:
db.add_all([models.Action(adventure_id=adv_id, type="ai", text=chunk)
for _ in range(40)])
db.commit()
written += 40 * len(chunk)
return written
def test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back(client):
fill_past_the_limit(client.adv_id)
response = export(client, client.adv_id)
# Delivered, whole, and still the same format.
body = response.content
assert len(body) > limits.MAX_IMPORT_BODY_BYTES
parsed = json.loads(body)
assert parsed["format"] == "ai-dnd-adventure-v3"
assert len(parsed["actions"]) > 40
# And honest about what this version can do with it.
assert response.headers["X-Importable-By-This-Version"] == "false"
warning = response.headers["X-Export-Warning"]
assert limits.import_limit_label() in warning
assert "exported successfully" in warning
assert "cannot import" in warning
assert int(response.headers["X-Export-Bytes"]) == len(body)
def test_the_warning_follows_the_configured_limit(monkeypatch):
"""The text is generated from the constant, so changing the constant changes
the sentence rather than leaving a stale number in it."""
assert "20 MB" in limits.oversized_export_warning(21_000_000)
monkeypatch.setattr(limits, "MAX_IMPORT_BODY_BYTES", 50 * 1024 * 1024)
assert limits.import_limit_label() == "50 MB"
assert "50 MB" in limits.oversized_export_warning(60_000_000)
assert "20 MB" not in limits.oversized_export_warning(60_000_000)
def test_the_bundle_itself_never_carries_the_warning(client):
"""The warning is about the export, not part of the portable story file."""
fill_past_the_limit(client.adv_id)
parsed = json.loads(export(client, client.adv_id).content)
flat = json.dumps(parsed).lower()
assert "import limit" not in flat
assert "cannot import" not in flat
for key in parsed:
assert "warning" not in key.lower()
def test_that_same_bundle_is_refused_by_import_naming_the_limit(client):
fill_past_the_limit(client.adv_id)
body = export(client, client.adv_id).content
response = client.post("/api/adventures/import", content=body,
headers={"Content-Type": "application/json"})
assert response.status_code == 413
detail = response.json()["detail"]
assert "too large" in detail.lower()
assert limits.import_limit_label() in detail
def test_a_bundle_under_the_limit_still_imports(client):
"""The refusal is about size alone: the ordinary path is untouched."""
body = export(client, client.adv_id).content
assert len(body) < limits.MAX_IMPORT_BODY_BYTES
response = client.post("/api/adventures/import", content=body,
headers={"Content-Type": "application/json"})
assert response.status_code == 201, response.text[:200]
+152
View File
@@ -0,0 +1,152 @@
"""v1.1 WP-E: the contrast audit is a gate, not a report.
Before this package `tools/contrast_audit.py` measured control boundaries,
printed that two of them were below 3:1, and exited 0 anyway — on the argument
that a control is identified by its label rather than its edge. A check that
cannot fail is not a check, and these tests are what make it one: the threshold
is exercised from both sides, on a real tokens file, so a future palette change
that dims a control's edge stops the run instead of adding a line to it.
python -m pytest tests/test_v11_e_contrast.py -v
"""
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tools"))
import contrast_audit as audit # noqa: E402
# --------------------------------------------------------------- the maths
def test_the_ratio_is_the_wcag_ratio():
"""Anchored on values with known answers, so a broken formula is visible."""
assert audit.ratio("#ffffff", "#000000") == pytest.approx(21.0, abs=0.01)
assert audit.ratio("#ffffff", "#ffffff") == pytest.approx(1.0, abs=0.001)
# Order does not matter: contrast is symmetric.
assert audit.ratio("#131320", "#676792") == pytest.approx(
audit.ratio("#676792", "#131320"), abs=1e-9)
# ------------------------------------------------------- the gate itself
def tokens_file(tmp_path: Path, **overrides: str) -> Path:
"""A real tokens file with named colours replaced."""
source = audit.TOKENS.read_text()
for name, value in overrides.items():
token = "--" + name.replace("_", "-")
start = source.index(f"{token}: ")
end = source.index(";", start)
source = source[:start] + f"{token}: {value}" + source[end:]
written = tmp_path / "tokens.css"
written.write_text(source)
return written
def run_against(path: Path, monkeypatch) -> int:
monkeypatch.setattr(audit, "TOKENS", path)
return audit.main()
def test_a_boundary_below_three_to_one_fails_the_run(tmp_path, monkeypatch, capsys):
"""2.99:1 against --bg-input — the wrong side of the line by one hundredth.
This value clears 3:1 against --bg-panel (3.21:1), so it would have passed
the audit as M11 wrote it. It fails now because the floor is taken against
the background the control is actually drawn on.
"""
below = tokens_file(tmp_path, border="#58639a")
assert run_against(below, monkeypatch) == 1
printed = capsys.readouterr().out
assert "2.99:1" in printed
assert "FAIL" in printed
assert "below 3:1 (WCAG 1.4.11)" in printed
def test_a_boundary_at_three_to_one_passes(tmp_path, monkeypatch, capsys):
"""3.00:1 — the right side of the same line, one hundredth from the value
above, and not failed for arithmetic the reader cannot see."""
at = tokens_file(tmp_path, border="#58639b")
assert run_against(at, monkeypatch) == 0
printed = capsys.readouterr().out
assert "3.00:1" in printed
assert "every control boundary clears 3:1" in printed
def test_the_v1_0_0_boundary_would_now_fail(tmp_path, monkeypatch, capsys):
"""The value v1.0.0 shipped. This is the defect WP-E closes, and the gate
has to be the thing that would have caught it."""
shipped = tokens_file(tmp_path, border="#2b2b3d")
assert run_against(shipped, monkeypatch) == 1
assert "1.24:1" in capsys.readouterr().out # against --bg-input, the worst case
def test_a_text_pair_below_four_point_five_still_fails(tmp_path, monkeypatch, capsys):
"""WP-E raised the boundary floor and must not have lowered the text one."""
dimmed = tokens_file(tmp_path, text_dim="#5a5750")
assert run_against(dimmed, monkeypatch) == 1
assert "below WCAG AA (1.4.3)" in capsys.readouterr().out
def test_a_missing_token_is_a_failure_not_a_skip(tmp_path, monkeypatch, capsys):
source = audit.TOKENS.read_text().replace("--border-bright:", "--border-was-renamed:")
written = tmp_path / "tokens.css"
written.write_text(source)
assert run_against(written, monkeypatch) == 1
assert "MISSING TOKEN" in capsys.readouterr().out
# ------------------------------------------------- the shipped palette
def test_the_real_tokens_pass_both_criteria(capsys):
"""The palette as it stands, through the same gate CI would run."""
assert audit.main() == 0
printed = capsys.readouterr().out
assert "every text pair clears WCAG AA (1.4.3)" in printed
assert "every control boundary clears 3:1 (1.4.11)" in printed
assert "FAIL" not in printed
def test_every_boundary_pair_is_measured_against_the_background_it_is_drawn_on():
"""The audit used to check borders only against --bg-panel, which is not
where the bordered controls are: inputs and buttons sit on --bg-input
(styles/forms.css), which is lighter and therefore harder. Checking only the
easier background would let a token pass while the real control failed."""
boundary = [(fg, bg) for kind, fg, bg, _, _ in audit.PAIRS if kind == "boundary"]
for background in ("--bg-input", "--bg-panel", "--bg"):
assert ("--border", background) in boundary, background
assert ("--border-bright", "--bg-input") in boundary
def test_the_text_baselines_are_unchanged_by_wp_e():
"""WP-E changed only boundary tokens. These are the M11 text measurements,
and they have to still be exactly what the earlier reports recorded."""
tokens = audit.read_tokens(audit.TOKENS)
measured = {
"body text on the page": audit.ratio(tokens["--text"], tokens["--bg"]),
"body text in a panel": audit.ratio(tokens["--text"], tokens["--bg-panel"]),
"secondary text in a panel": audit.ratio(tokens["--text-dim"], tokens["--bg-panel"]),
"secondary text on the page": audit.ratio(tokens["--text-dim"], tokens["--bg"]),
}
assert measured["body text on the page"] == pytest.approx(14.57, abs=0.01)
assert measured["body text in a panel"] == pytest.approx(13.57, abs=0.01)
assert measured["secondary text in a panel"] == pytest.approx(5.48, abs=0.01)
assert measured["secondary text on the page"] == pytest.approx(5.88, abs=0.01)
def test_the_hover_edge_stays_brighter_than_the_resting_edge():
"""Rest and hover have to remain distinguishable from each other, not merely
each clear the floor against the background."""
tokens = audit.read_tokens(audit.TOKENS)
panel = tokens["--bg-panel"]
assert audit.ratio(tokens["--border-bright"], panel) > audit.ratio(tokens["--border"], panel)
def test_a_boundary_never_becomes_as_loud_as_body_text():
"""A control's edge that outshines the words inside it is its own defect."""
tokens = audit.read_tokens(audit.TOKENS)
panel = tokens["--bg-panel"]
assert audit.ratio(tokens["--border-bright"], panel) < audit.ratio(tokens["--text"], panel)
+324
View File
@@ -0,0 +1,324 @@
"""v1.1 WP-A2: the protocol a narrator copies stays out of the story, and nothing else does.
The M11 closeout's identity run (an office meeting, a 3B narrator, a 4,096
window) stored four turns carrying text the application wrote, not the story:
- `> Create_entity(new_person, "john", …` — the event vocabulary as the prompt
printed it, `name(field, …)`, copied as if it were a call;
- `> set_possession(silver-key, "alice") Adds the silver key to Alice's
possession.` — the call again, naming the fantasy example slug from the fixed
state rule, in a meeting room;
- `Scene: Bill, Alice, … (at The meeting room)` — the renderer's own scene line;
- `[Hard limit: your next turn must not exceed 180 words, … append the state
block well inside the limit.]` — the length hint, reworded at the front and
verbatim at the end.
v1.0.0 removed none of them. The rule this module is held to is unchanged from
M5: **removing story is worse than leaving protocol.** Every removal below is
anchored to a string or a vocabulary the application owns, and every one has
story beside it that must survive.
python -m pytest tests/test_v11_protocol_echo.py -v
"""
import json
import re
import pytest
from app.context import builder
from app.narrative import events, extract, render
# ---------------------------------------------------------------- the prompt
#: The identifiers the v1 state rule taught every campaign, from the fantasy
#: acceptance fixture. None may come back into a fixed instruction.
FANTASY_IDENTIFIERS = ("mara", "silver-key", "silver key", "old-abbey", "abbey",
"aldric", "westhaven", "crypt", "edrin")
#: And nothing from the science-fiction fixture either: neutral means neutral,
#: not "the other genre".
SCIFI_IDENTIFIERS = ("persephone", "imani", "data-crystal", "data crystal", "airlock")
def _fixed_instructions() -> str:
return "\n".join([
extract.EMIT_RULE,
extract.EMIT_REMINDER,
events.vocabulary_for_prompt(),
builder.length_hint(500),
builder.length_hint(500, "brief"),
builder.length_hint(500, "long"),
builder.length_hint(120),
]).lower()
@pytest.mark.parametrize("identifier", FANTASY_IDENTIFIERS + SCIFI_IDENTIFIERS)
def test_the_fixed_state_instructions_name_no_fixture_identifier(identifier):
"""A2-1. The example slug the office run copied cannot come back."""
assert not re.search(rf"\b{re.escape(identifier)}\b", _fixed_instructions())
def test_the_example_uses_neutral_identifiers():
for neutral in ("character-1", "item-1", "location-1"):
assert neutral in extract.EMIT_RULE
def test_the_worked_example_is_a_block_this_extractor_accepts():
"""The example is the wire format, byte for byte, not an illustration of it."""
example = extract.EMIT_RULE[extract.EMIT_RULE.index("```state"):]
prose, parsed, _raw = extract.split("The door opens.\n\n" + example)
assert prose == "The door opens."
assert isinstance(parsed, dict)
assert [e["type"] for e in parsed["events"]] == [
"set_possession", "set_current_location"]
for event in parsed["events"]:
assert events.is_allowed(event["type"])
def test_the_vocabulary_is_not_written_as_function_calls():
"""A2-2. `set_possession(item, owner)` is the notation the narrator copied."""
vocabulary = events.vocabulary_for_prompt()
for name in events.SPECS:
assert not re.search(rf"\b{name}\s*\(", vocabulary), name
def test_every_event_is_described_in_the_shape_the_model_must_send():
lines = events.vocabulary_for_prompt().splitlines()
assert len(lines) == len(events.SPECS)
for name, line in zip(events.SPECS, lines):
shape = line.strip().split(" — ", 1)[0]
obj = json.loads(shape)
assert obj["type"] == name
assert set(obj) - {"type"} == set(events.SPECS[name]["required"])
for optional in events.SPECS[name]["optional"]:
assert optional in line
def test_the_length_hint_carries_the_phrases_the_extractor_recognises():
"""One source for the words, so the builder and the extractor cannot drift."""
for narration_length in ("", "brief", "medium", "long"):
hint = builder.length_hint(500, narration_length)
assert hint.startswith(extract.LENGTH_HINT_OPENING)
assert extract.LENGTH_HINT_TAIL in hint
# ---------------------------------------------------------- observed shapes
STORY = (
"Alice looks at John, the tension in the room palpable.\n\n"
"John nods. \"I'm ready to contribute.\""
)
@pytest.mark.parametrize("leak", [
# Depth 20: the call, the fantasy slug, and a gloss on the same line.
'> set_possession(silver-key, "alice") Adds the silver key to Alice\'s possession.',
# Depth 10: cut off by the output limit mid-call.
'> Create_entity(new_person, "john", "character", "A determined team member", ["john',
# Unquoted, and a vocabulary name in any case.
'SET_CURRENT_LOCATION(bill, office)',
'add_fact(predicate="knows the plan", subject="alice")',
])
def test_an_event_call_line_at_the_end_leaves_the_story(leak):
prose, parsed, _raw = extract.split(f"{STORY}\n\n{leak}")
assert prose == STORY
assert parsed is None
def test_an_event_call_line_in_the_middle_leaves_and_the_story_after_it_stays():
"""Depth 12: the call, then more narration."""
reply = (
f"{STORY}\n\n"
'> Create_entity(new_person, "mike", "character", "A new team member.", ["mike"])\n\n'
"Mike takes the empty chair by the window."
)
prose, _parsed, _raw = extract.split(reply)
assert prose == f"{STORY}\n\nMike takes the empty chair by the window."
def test_the_depth_fourteen_tail_leaves_entirely():
"""A call, a rendered scene line, and a reworded length hint, in that order."""
reply = (
f"{STORY}\n\n"
'> Create_entity(mike, "character", "A new team member.", ["mike"])\n\n'
"Scene: Bill, Alice, Roger, John, and Mike at the table. (at The meeting room)\n\n"
"[Hard limit: your next turn must not exceed 180 words, and it should not stop "
"short of about 70. Prefer the lower end of that range unless the scene genuinely "
"needs more. Finish the narration and append the state block well inside the limit.]"
)
prose, parsed, _raw = extract.split(reply)
assert prose == STORY
assert parsed is None
@pytest.mark.parametrize("hint", [
builder.length_hint(500),
builder.length_hint(500, "brief"),
# Cut off by the output limit before the tail.
"[Hard limit: this turn must not exceed 180 words, and it should not stop short",
# Reworded at the front, as the 3B narrator did.
"[Hard limit: your next turn must not exceed 506 words. Write only as much as the "
"moment needs — a typical turn is much shorter. Finish the narration and append "
"the state block well inside the limit.]",
])
def test_a_parroted_length_hint_at_the_end_leaves_the_story(hint):
prose, _parsed, _raw = extract.split(f"{STORY}\n\n{hint}")
assert prose == STORY
def test_a_rendered_scene_line_at_the_end_leaves_the_story():
prose, _parsed, _raw = extract.split(
f"{STORY}\n\nScene: A tense budget meeting. (at The meeting room)")
assert prose == STORY
def test_a_fenced_block_with_a_call_line_above_it_still_parses_and_applies():
"""A2-6. The proposal is still read when protocol litter surrounds it."""
reply = (
f"{STORY}\n\n"
'> set_current_location(john, office)\n\n'
'```state\n{"events": [{"type": "set_current_location", '
'"entity": "john", "location": "office"}]}\n```'
)
prose, parsed, raw = extract.split(reply)
assert prose == STORY
assert parsed["events"][0]["entity"] == "john"
assert raw.startswith("{")
# ------------------------------------------------ adversarial story that stays
@pytest.mark.parametrize("reply", [
# The owner's cases.
'The engineer writes "set_power(core, 80)" on the whiteboard.',
'She says, "Create_entity is a terrible name for a company."',
'The old manual contains a heading labeled "Scene:"',
'He reads aloud: "[Hard limit: 500 words]" and laughs.',
# A vocabulary name, written into a story, not at the start of a line.
'Nadia squints at the log: the last command was set_possession(badge, guard).',
# Call-shaped, at the start of a line, but not an event this protocol has.
"The terminal scrolls.\n\n> open_door(north)\n\nNothing happens.",
# A vocabulary call inside the story's own code block is the story's code.
"She types:\n\n```python\ncreate_entity(ship)\nset_possession(key, captain)\n```\n\n"
"The console beeps twice.",
# A bracket at the very end, in-world, that is not the application's hint.
"The warning light blinks.\n\n[Hard limit of the reactor: three hours]",
"The contract ends with a clause.\n\n[Hard limit: forty days, no extensions]",
# A scene heading in a screenplay the characters are writing, mid-story.
"Scene: a kitchen, late.\n\nShe crosses it out and starts again.",
# A last line that starts like the renderer's but is not its shape.
"The director calls it.\n\nScene: take two, and nobody moves.",
# A fact restated inside a sentence.
"Alice knew the badge opened the server room, and said nothing.",
"Memory: she remembered the bells.",
])
def test_story_that_resembles_the_new_rules_is_kept(reply):
"""A2-5."""
prose, parsed, _raw = extract.split(reply)
assert prose == reply
assert parsed is None
# ------------------------------------------------------ the replay attribution
@pytest.mark.parametrize("line, rule", [
('> set_possession(silver-key, "alice") Adds the key.', extract.RULE_EVENT_CALL),
("Create_entity(new_person", extract.RULE_EVENT_CALL),
("[Hard limit: this turn must not exceed 90 words. Finish the narration and append "
"the state block well inside the limit.]", extract.RULE_LENGTH_HINT),
("Scene: A meeting. (at The meeting room)", extract.RULE_SCENE_LINE),
("John nods.", None),
('He reads aloud: "[Hard limit: 500 words]" and laughs.', None),
])
def test_a_removed_line_is_attributed_to_the_rule_that_removes_it(line, rule):
assert extract.explain_removed_line(line) == rule
# ------------------------------------- corrective: the depth-16 instruction tail
#: Cut down from the v1.1 identity diagnostic's depth-16 turn, whose stored text
#: was exactly the extractor's output. The two story paragraphs are shortened;
#: the four trailing lines are verbatim.
DEPTH_16_STORY = (
"John's initial ideas are thoughtful and insightful, and the room fills with a "
"sense of optimism.\n\n"
"John's enthusiasm is contagious, and the meeting room is electric with the "
"excitement of a fruitful collaboration ahead."
)
DEPTH_16_TAIL = (
"Scene: Bill, Alice and Roger at the table; John not yet arrived.\n\n"
"[Hard limit: this is now 180 words.]\n\n"
"[Reminder: end your reply with a `state` block listing the events your narration "
"made true, with absolute values.]\n\n"
"[You don't need to continue; your turn must now be about John entering the room. "
"Continue the story here, directly. Output only story text.]"
)
def test_the_depth_sixteen_instruction_tail_leaves_entirely():
"""The corrective's positive regression. v1.1's first A2 left all four lines:
the last bracket was a reworded continue hint nothing recognised, so nothing
above it was ever at the end."""
prose, parsed, _raw = extract.split(f"{DEPTH_16_STORY}\n\n{DEPTH_16_TAIL}")
assert prose == DEPTH_16_STORY
assert parsed is None
def test_the_continue_hint_phrase_is_the_providers_own_sentence():
from app.providers.openai_compatible import CHAT_CONTINUE_HINT
assert extract.CONTINUE_HINT_PHRASE in CHAT_CONTINUE_HINT
def test_an_echoed_continue_hint_alone_at_the_end_leaves():
prose, _p, _r = extract.split(
f"{STORY}\n\n[Keep going. Continue the story here, directly. Output only story text.]")
assert prose == STORY
@pytest.mark.parametrize("reply", [
# A hint-opened bracket with no echoed instruction below it is in-world.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# Nor does a state block below it make it an instruction.
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
# The phrase in the middle of a story is prose, not a trailing echo.
'She wrote "output only story text" on the card, then crossed it out.\n\nThe rain went on.',
# A trailing in-world bracket that only resembles a continuation.
f"{STORY}\n\n[To be continued]",
])
def test_story_brackets_near_the_corrective_rule_are_kept(reply):
prose, _p, _r = extract.split(reply)
assert prose == reply
def test_a_hint_opened_bracket_above_a_state_block_is_kept():
reply = (f"{STORY}\n\n[Hard limit: forty days, no extensions]\n\n"
'```state\n{"events": []}\n```')
prose, parsed, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[Hard limit: forty days, no extensions]"
assert parsed == {"events": []}
def test_an_in_world_bracket_above_an_echoed_hint_is_kept():
"""Only a bracket opening the way the application's hint opens is taken with
the echo. Any other bracket above it is the story's."""
reply = (f"{STORY}\n\n[The sign on the door reads: Closed]\n\n"
"[Continue the story here, directly. Output only story text.]")
prose, _p, _r = extract.split(reply)
assert prose == f"{STORY}\n\n[The sign on the door reads: Closed]"
@pytest.mark.parametrize("line, rule", [
("[Hard limit: this is now 180 words.]", extract.RULE_INSTRUCTION_TAIL),
("[Reminder: end your reply with a `state` block listing the events.]",
extract.RULE_INSTRUCTION_TAIL),
("[You don't need to continue. Output only story text.]", extract.RULE_INSTRUCTION_TAIL),
])
def test_the_corrective_rule_is_attributed(line, rule):
assert extract.explain_removed_line(line) == rule
def test_the_new_rules_do_not_disturb_the_section_headings_they_share_a_module_with():
"""The renderer's headings are the M11 rules' anchor. A2 adds none."""
assert render.HEADING_SCENE == "Scene:"
assert "Scene:" not in render.SECTION_HEADINGS
+52 -31
View File
@@ -26,16 +26,24 @@ TOKENS = Path(__file__).resolve().parent.parent.parent / "frontend/src/styles/to
#: rather than combinatorial, because "every colour against every other" reports
#: pairs that never meet on screen.
#:
#: The `kind` matters and is not a way of grading on a curve. **text** pairs are
#: WCAG 1.4.3 Contrast (Minimum) and are what §21 of the M11 brief asks about;
#: they are pass/fail. **boundary** pairs are WCAG 1.4.11 Non-text Contrast,
#: which applies to "visual information required to identify user interface
#: components" — and in this design a control is identified by its *label*,
#: which is measured above and passes, not by its edge. So a boundary below 3:1
#: is reported with its number and does not fail the run; what it would take to
#: turn it into a real failure is a control with no visible label, and there is
#: no such control (`tools/m11_browser.py` asserts every visible control has an
#: accessible name, and the story controls are text buttons).
#: The `kind` records which success criterion a pair is measured against —
#: **text** is WCAG 1.4.3 Contrast (Minimum), **boundary** is WCAG 1.4.11
#: Non-text Contrast — and **both are pass/fail**.
#:
#: v1.1 WP-E overturned the earlier position here, which was that a boundary
#: below 3:1 could be recorded rather than failed because "a control is
#: identified by its label, not by its edge". That argument understates what
#: 1.4.11 asks: the criterion covers the visual information needed to identify
#: a component *and its boundary*, and a reader who cannot see where a text box
#: ends cannot see that there is a text box to type into, label or no label.
#: The edges were at 1.33:1 and 1.75:1 — the tokens were raised instead.
#:
#: Borders are measured against **every background they are drawn on**, and the
#: floor is the worst of them. Inputs and buttons sit on --bg-input, which is
#: lighter than --bg-panel and so the harder case; checking only --bg-panel
#: would have let a token pass the audit while the real control failed.
#: `--bg-panel-glass` is translucent and cannot be resolved from tokens alone;
#: that edge is measured on the rendered page by `tools/m11_browser.py`.
PAIRS = [
("text", "--text", "--bg", 4.5, "body text on the page"),
("text", "--text", "--bg-panel", 4.5, "body text in a panel"),
@@ -48,8 +56,14 @@ PAIRS = [
("text", "--danger", "--bg-panel", 4.5, "an error message"),
("text", "--warning", "--bg-panel", 4.5, "a caution message"),
("text", "--player", "--bg", 4.5, "the player's own words"),
("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge"),
("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge"),
("boundary", "--border", "--bg-input", 3.0, "a field or button's resting edge"),
("boundary", "--border-bright", "--bg-input", 3.0, "a field or button's hover edge"),
("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge in a panel"),
("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge in a panel"),
("boundary", "--border", "--bg", 3.0, "a divider on the page"),
("boundary", "--border-bright", "--bg", 3.0, "the scrollbar thumb"),
("boundary", "--accent-dim", "--bg-input", 3.0, "a focused field's edge"),
("boundary", "--accent-dim", "--bg-panel", 3.0, "a focused control's edge in a panel"),
("boundary", "--chart-1", "--bg-panel", 3.0, "a chart bar"),
("boundary", "--chart-2", "--bg-panel", 3.0, "a chart bar"),
("boundary", "--chart-3", "--bg-panel", 3.0, "a chart bar"),
@@ -81,37 +95,44 @@ def ratio(a: str, b: str) -> float:
def main() -> int:
tokens = read_tokens(TOKENS)
print(f"{TOKENS.relative_to(TOKENS.parents[3])}: {len(tokens)} colour tokens\n")
# Named defensively: the tests run this against a temporary tokens file,
# which need not sit four directories deep the way the real one does.
label = TOKENS.name
if len(TOKENS.parents) > 3:
label = TOKENS.relative_to(TOKENS.parents[3])
print(f"{label}: {len(tokens)} colour tokens\n")
print(f"{'pair':44} {'kind':9} {'ratio':>7} {'floor':>6} verdict")
print("-" * 82)
failures, advisories = 0, 0
text_failures, boundary_failures = 0, 0
for kind, foreground, background, floor, description in PAIRS:
if foreground not in tokens or background not in tokens:
print(f"{description:44} {kind:9} {'—':>7} {floor:>6.1f} MISSING TOKEN")
failures += 1
text_failures += 1
continue
measured = ratio(tokens[foreground], tokens[background])
ok = measured >= floor
if not ok:
if kind == "text":
failures += 1
verdict = "FAIL"
else:
advisories += 1
verdict = "below 1.4.11 (label carries it)"
else:
# Rounded to the two decimals printed, so the verdict matches what the
# reader is shown: a pair reported as 3.00:1 is not failed for arithmetic
# the output does not display.
if round(measured, 2) >= floor:
verdict = "pass"
else:
verdict = "FAIL"
if kind == "text":
text_failures += 1
else:
boundary_failures += 1
print(f"{description:44} {kind:9} {measured:>6.2f}:1 {floor:>6.1f} {verdict}")
print()
if failures:
print(f"{failures} text pair(s) below WCAG AA — this is a defect")
if text_failures:
print(f"{text_failures} text pair(s) below WCAG AA (1.4.3) — this is a defect")
else:
print("every text pair clears WCAG AA (1.4.3)")
if advisories:
print(f"{advisories} boundary pair(s) below 3:1 (1.4.11). Recorded rather "
"than failed: every control in this design carries a visible text "
"label, which is measured above and passes.")
return 1 if failures else 0
if boundary_failures:
print(f"{boundary_failures} boundary pair(s) below 3:1 (WCAG 1.4.11) — "
"this is a defect")
else:
print("every control boundary clears 3:1 (1.4.11)")
return 1 if (text_failures or boundary_failures) else 0
if __name__ == "__main__":
+1111 -137
View File
File diff suppressed because it is too large Load Diff
+12 -1
View File
@@ -90,6 +90,11 @@ from app.routers import adventures # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
#: v1.1 release Gate 7 asks for this diagnostic with **memory on**. It shipped
#: with no embedding model and the bank switched off, so a release run of it
#: would have reported a clean identity result without memory ever taking part.
#: Empty keeps the old behaviour, which is what `--scripted` wants.
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
#: The cast the finding describes: a protagonist and three others, all on stage.
CAST = [
@@ -140,7 +145,7 @@ def _setup(scripted: bool):
db.add(models.Settings(
user_id=user.id,
model=MODEL or "scripted", endpoint_url=ENDPOINT or "http://127.0.0.1:11434/v1",
embedding_model="", context_token_budget=16384, max_output_tokens=500,
embedding_model=EMBED_MODEL, context_token_budget=16384, max_output_tokens=500,
model_timeout_seconds=300,
))
db.commit()
@@ -167,6 +172,12 @@ def _campaign(client) -> int:
})
created.raise_for_status()
adv = created.json()["id"]
# Memory and summaries are per-campaign switches defaulting to off. Gate 7
# asks for this diagnostic with memory on, and the ten beats below write
# twenty actions — past `MEMORY_START` — so the bank has something to do.
client.patch(f"/api/adventures/{adv}",
json={"memory_bank_enabled": True, "auto_summarize": True}
).raise_for_status()
answer = client.post(f"/api/adventures/{adv}/state/corrections", json={
"events": [
{"type": "create_entity", "entity": key, "entity_type": kind, "name": name}
+234 -1
View File
@@ -190,6 +190,31 @@ CLUE_FACT = {
"fact_id": "silver-key-opens-crypt",
}
#: v1.1 WP-B.1: a second planted fact, established in the **story only**.
#:
#: The M04 clue above is planted as accepted state, and memories are written
#: from story text, so no memory could ever carry it on its own. That is why
#: every M04 recovery so far ran through state. This fact is told to the reader
#: in narration and never corrected into state, so memory is the only layer that
#: is meant to carry it. `--independent-fact` plants it and reports
#: `recovered_through_memory_independent` only when every other layer is proven
#: not to carry it. The words are copied from `tools/memory_diagnostic.FACT_F`,
#: for the reason `HISTORY_LABELS` is copied.
INDEPENDENT_FACT_TEXT = ("I watch Mara slip the amber sundial inside the cracked teapot on "
"the tavern's top shelf, and she makes me promise to tell no one.")
INDEPENDENT_FACT_TERMS = ("sundial", "teapot")
INDEPENDENT_RECALL_TEXT = "I ask Mara quietly where she hid the amber sundial."
#: How far past the planting turn its memory block can reach. Narration inside
#: that block may repeat the fact; narration after it may not.
INDEPENDENT_BLOCK_SLACK = 6
INDEPENDENT_PRECONDITIONS = (
"planted_turn_outside_history",
"absent_from_state",
"absent_from_summary",
"absent_from_knowledge",
"absent_from_later_narration",
)
CANON = [
"The dead do not return. No rite, relic or bargain has ever returned anyone.",
"The abbey crypt has been sealed since the founding.",
@@ -365,6 +390,14 @@ class Run:
#: The depth of the player turn that planted the clue. M04's
#: precondition is that this turn has left the history window.
self.planted_depth: int | None = None
#: v1.1 WP-B.1, with --independent-fact: where the story-only fact was
#: planted, and the accepted-turn count at which each isolation
#: precondition first failed.
self.independent_fact = False
self.independent_depth: int | None = None
self.independent_violations: dict[str, int] = {}
self.last_done: dict = {}
self.last_report: dict = {}
# ------------------------------------------------------------ recording
@@ -393,6 +426,8 @@ class Run:
"turns_target": self.turns_target,
"log_offset": self.log_offset,
"planted_depth": self.planted_depth,
"independent_depth": self.independent_depth,
"independent_violations": self.independent_violations,
"written": datetime.now().isoformat(timespec="seconds"),
}
tmp = self.out / (RESUME_FILE + ".tmp")
@@ -414,6 +449,8 @@ class Run:
self.elapsed_before = prior.get("elapsed_seconds", 0)
self.log_offset = prior.get("log_offset", 0)
self.planted_depth = prior.get("planted_depth")
self.independent_depth = prior.get("independent_depth")
self.independent_violations = dict(prior.get("independent_violations") or {})
self.resumed = True
def reattach(self) -> None:
@@ -559,15 +596,55 @@ class Run:
self.accepted += 1
seconds = time.monotonic() - started
sample = self.measure()
# v1.1 WP-A1: what the server said it read for the turn just played, from
# the `done` event. `.get` because a build before v1.1 sends none.
done = next((e for e in events if e.get("type") == "done"), {})
accounting = done.get("accounting") or {}
sample.update({
"accounting_status": accounting.get("status"),
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
"app_prompt_estimate": accounting.get("estimate"),
"observed_margin": accounting.get("observed_margin"),
"safety_reserve": accounting.get("safety_reserve"),
})
self.last_done = done
if self.independent_fact and self.independent_depth is not None:
self._check_independent_isolation(done, sample)
self.note("turn", text=text, seconds=round(seconds, 1), **sample)
return {"accepted": True, "seconds": seconds, **sample}
def _check_independent_isolation(self, done: dict, sample: dict) -> None:
"""v1.1 WP-B.1: does anything but memory carry the story-only fact yet?
Checked on every accepted turn, so a run knows the first turn at which
the experiment stopped being about memory, instead of finding out at
recall. Each precondition records only its first failure.
"""
depth = sample.get("total_actions", 0) - 1
text = (done.get("action") or {}).get("text") or ""
found = {}
if depth > self.independent_depth + INDEPENDENT_BLOCK_SLACK and _mentions_fact(text):
found["absent_from_later_narration"] = f"narration at depth {depth}"
document = self.state().get("document") or {}
if _mentions_fact(json.dumps(document)):
found["absent_from_state"] = "the narrative state names the fact"
summary = next((sec.get("text", "") for sec in (self.last_report.get("sections") or [])
if sec.get("label") == SUMMARY_LABEL), "")
if _mentions_fact(summary):
found["absent_from_summary"] = "the active summary names the fact"
for name, detail in found.items():
if name not in self.independent_violations:
self.independent_violations[name] = self.accepted
self.note("independent_precondition_failed", precondition=name, detail=detail)
sample["independent_violations"] = dict(self.independent_violations)
def count_actions(self) -> int:
return self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")["total"]
def measure(self) -> dict:
"""M03's numbers, read from the prompt the app would send right now."""
report = self.server.call("GET", f"/adventures/{self.adv}/context")
self.last_report = report
tokens = report["tokens"]
sections = {s["label"]: s["tokens"] for s in report["sections"]}
window = report.get("window") or {}
@@ -703,6 +780,10 @@ def main() -> int:
"--max-consecutive-failures", type=int,
default=DEFAULT_MAX_CONSECUTIVE_FAILURES,
help="stop and write the evidence after this many unaccepted turns")
parser.add_argument(
"--independent-fact", action="store_true",
help=("v1.1 WP-B.1: also plant a story-only fact at depth 3 and report "
"whether memory alone recovers it"))
args = parser.parse_args()
if not (ENDPOINT and MODEL and EMBED_MODEL):
@@ -735,6 +816,7 @@ def main() -> int:
server.start()
run = Run(server, out, turns_target=args.turns,
turn_timeout=args.turn_timeout)
run.independent_fact = args.independent_fact
if prior:
run.adopt(prior)
@@ -771,6 +853,18 @@ def main() -> int:
"the planted clue is not in accepted state, so M04 cannot "
"be measured from this run. Stopping before the campaign "
"starts rather than reporting a recall failure later.")
if args.independent_fact:
# v1.1 WP-B.1: the story-only fact, told in the next turn and
# never corrected into state. Depth 3: the opening, the clue turn
# and its reply come first.
if any(_mentions_fact(md) for md in (CANON_MD, REFERENCE_MD, INSPIRATION_MD)):
raise SystemExit("the imported knowledge names the independent fact")
planting_f = run.turn(INDEPENDENT_FACT_TEXT)
if not planting_f.get("accepted"):
raise SystemExit("the turn that plants the independent fact was not accepted")
run.independent_depth = planting_f["total_actions"] - 2
run.note("independent_fact_planted", depth=run.independent_depth,
terms=list(INDEPENDENT_FACT_TERMS))
# The first checkpoint, and the point from which --resume works: the
# campaign exists and its clue is planted.
run.save_resume()
@@ -835,10 +929,15 @@ def main() -> int:
# Skipped on an aborted run: it asks the narrator a question, and the
# reason the run stopped is that the narrator does not answer.
recall = None
independent = None
if aborted is None:
run.note("recall_begin")
recall = _recall(run)
(out / "recall.json").write_text(json.dumps(recall, indent=2))
if args.independent_fact and run.independent_depth is not None:
independent = _independent_recall(run, out / "campaign.db")
(out / "recall-independent.json").write_text(json.dumps(independent, indent=2))
run.note("independent_recall", verdict=independent["verdict"])
# ---- Export whatever exists, for the recovery evidence. ----
# Attempted even for an aborted run: the recovery check and the storage
@@ -886,6 +985,7 @@ def main() -> int:
"elapsed_seconds": run.elapsed(),
"turn_timeout_seconds": args.turn_timeout,
"recall": recall,
"independent_recall": independent,
"final_state": _or_none(lambda: run.state()["document"]),
"final_measurement": _or_none(run.measure),
"db_bytes": db_path.stat().st_size,
@@ -1213,6 +1313,123 @@ def _m04_verdict(recall: dict) -> str:
return "not_recovered"
def _mentions_fact(text: str | None) -> bool:
"""v1.1 WP-B.1: whether `text` names the independent fact, as a whole word."""
low = (text or "").lower()
return any(re.search(rf"(?<![a-z]){term}(?![a-z])", low) for term in INDEPENDENT_FACT_TERMS)
def _independent_memory_verdict(check: dict) -> str:
"""v1.1 WP-B.1: whether memory alone recovered the story-only fact.
`recovered_through_memory_independent` requires every precondition, so no
other layer could have carried the fact. It also requires that a memory
covering the planting turn carries the fact and was injected into the recall
turn. A failed precondition is named and is never a recovery, and the M04
verdicts above are untouched.
"""
if check.get("independent_planted_depth") is None:
return "precondition_unknown:planted_depth"
for name in INDEPENDENT_PRECONDITIONS:
value = check.get(name)
if value is None:
return f"precondition_unknown:{name}"
if not value:
return f"precondition_failed:{name}"
if not check.get("memory_covering_planting_carries_fact"):
return "not_recovered:not_created"
if check.get("memory_forgotten"):
return "not_recovered:evicted"
if not check.get("memory_injected"):
return "not_recovered:not_injected"
return "recovered_through_memory_independent"
def _independent_recall(run: "Run", db_path: Path) -> dict:
"""v1.1 WP-B.1: ask for the story-only fact, and find out which layer answered.
The prompt-level facts come from the recall turn's own stored context. The
memory rows come from the campaign database, read-only. Ranking is not
recomputed here, because that needs the embedding model;
`tools/v11_b1_memory.py diagnose` does it afterwards against a copy of the
database.
"""
import sqlite3
import zlib
result = run.turn(INDEPENDENT_RECALL_TEXT)
action_id = (run.last_done.get("action") or {}).get("id")
snapshot = (run.server.call("GET", f"/adventures/{run.adv}/actions/{action_id}/context")
if result.get("accepted") and action_id else {}) or {}
sections = {}
for sec in snapshot.get("sections") or []:
sections.setdefault(sec.get("label"), []).append(sec.get("text", ""))
text_of = {label: "\n".join(parts) for label, parts in sections.items()}
floor = (snapshot.get("history") or {}).get("floor_depth")
depth = run.independent_depth
used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
covering = []
connection = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
try:
rows = connection.execute(
"SELECT id, text, source_start, source_end, forgotten, pinned, use_count, "
"branch_id, depth FROM memories WHERE adventure_id = ? AND source_start <= ? "
"AND source_end >= ? ORDER BY id", (run.adv, depth, depth)).fetchall()
blob = connection.execute(
"SELECT context_snapshot FROM actions WHERE id = ?", (action_id or -1,)).fetchone()
finally:
connection.close()
for row in rows:
memory_id, text, start, end, forgotten, pinned, use_count, branch_id, node_depth = row
covering.append({
"memory_id": memory_id, "text": text, "source_start": start, "source_end": end,
"forgotten": bool(forgotten), "pinned": bool(pinned), "use_count": use_count,
"branch_id": branch_id, "depth": node_depth,
"carries_fact": all(re.search(rf"(?<![a-z]){t}(?![a-z])", (text or "").lower())
for t in INDEPENDENT_FACT_TERMS),
"injected": memory_id in used,
})
carrying = [c for c in covering if c["carries_fact"]]
best = next((c for c in carrying if c["injected"]), carrying[0] if carrying else None)
stored_snapshot_readable = blob is not None and blob[0] is not None
if stored_snapshot_readable:
try:
json.loads(zlib.decompress(blob[0]))
except Exception: # noqa: BLE001
stored_snapshot_readable = False
document = run.state().get("document") or {}
violations = dict(run.independent_violations)
check = {
"independent_planted_depth": depth,
"recall_accepted": bool(result.get("accepted")),
"history_floor_depth": floor,
"planted_turn_outside_history": (None if not snapshot else
floor is not None and depth < floor),
"absent_from_state": ("absent_from_state" not in violations
and not _mentions_fact(json.dumps(document))
and not _mentions_fact(text_of.get(STATE_LABEL))),
"absent_from_summary": ("absent_from_summary" not in violations
and not _mentions_fact(text_of.get(SUMMARY_LABEL))),
"absent_from_knowledge": not any(
_mentions_fact(text_of.get(label)) for label in IMPORTED_KNOWLEDGE_LABELS),
"absent_from_later_narration": "absent_from_later_narration" not in violations,
"violations_first_turn": violations,
"covering_memories": covering,
"memory_covering_planting_carries_fact": bool(carrying),
"memory_forgotten": bool(best and best["forgotten"]),
"memory_injected": bool(best and best["injected"]),
"memory_text_in_memories_section": bool(
best and best["text"] and best["text"] in (text_of.get(MEMORIES_LABEL) or "")),
"memory_ids_used": used,
"recall_action_id": action_id,
"stored_snapshot_readable": stored_snapshot_readable,
}
check["verdict"] = _independent_memory_verdict(check)
return check
#: Signs the application stored protocol as story. The first is a state-section
#: heading with an indented entry under it, in any markdown, because
#: `## Established:` got past a plain substring match and the count read 1
@@ -1225,10 +1442,26 @@ PROTOCOL_LEAK_HEADING_RE = re.compile(
re.MULTILINE,
)
PROTOCOL_LEAK_EVENTS = '"events"'
#: v1.1 WP-A2: the two shapes the M11 closeout's identity run stored that the
#: two signs above cannot see — a line opening with a call to an event, and the
#: length hint echoed with the application's own wording. Copied, as above.
PROTOCOL_LEAK_CALL_RE = re.compile(
r"^[ \t]*(?:>[ \t]*)?(?:create_entity|set_entity_status|set_entity_attribute"
r"|set_entity_conditions|set_current_location|set_possession|clear_possession"
r"|add_fact|invalidate_fact|add_relationship|end_relationship"
r"|open_story_thread|resolve_story_thread|set_scene)[ \t]*\(",
re.IGNORECASE | re.MULTILINE,
)
PROTOCOL_LEAK_HINT_RE = re.compile(
r"\[Hard limit:[^\]]*(?:append the state block|turn must not exceed \d+ words)",
re.IGNORECASE,
)
def _leaks_protocol(text: str) -> bool:
return bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
return (bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
or bool(PROTOCOL_LEAK_CALL_RE.search(text))
or bool(PROTOCOL_LEAK_HINT_RE.search(text)))
def _protocol_leaks(bundle: dict) -> dict:
+175 -3
View File
@@ -15,10 +15,17 @@ what M9 recorded as "this machine cannot drive a file into the browser". The
narrower and more useful statement is that it refuses `/tmp`: a path under the
user's home works. `stage()` exists to put evidence files there, so knowledge
import can be exercised through the real file input rather than in two halves.
**Downloads (v1.1 WP-C).** The same snap Firefox saves a download without any
dialog when its profile says where to, and the folder is under `$HOME`.
`firefox_download_prefs` is that profile, `require_under_home` refuses a folder
the sandbox would not let it write, and `wait_for_download` decides when a file
has actually finished arriving — never the click that started it.
"""
from __future__ import annotations
import base64
import json
import os
import shutil
@@ -33,6 +40,10 @@ GECKODRIVER = shutil.which("geckodriver") or "/snap/bin/geckodriver"
#: Where files the browser must open are staged. Under $HOME because the snap
#: sandbox denies /tmp; see the module docstring.
STAGE = Path.home() / "m11-evidence"
#: The W3C key an element reference is returned under.
ELEMENT_KEY = "element-6066-11e4-a52e-4f735466cecf"
#: What Firefox names a download while it is still arriving.
PARTIAL_SUFFIXES = (".part",)
def stage(name: str, body: str | bytes) -> str:
@@ -51,14 +62,86 @@ def free_port() -> int:
return s.getsockname()[1]
def geckodriver_version() -> str:
try:
out = subprocess.run([GECKODRIVER, "--version"], capture_output=True, text=True, timeout=30)
return (out.stdout.splitlines() or ["?"])[0].strip()
except (OSError, subprocess.SubprocessError):
return "?"
class WebDriverError(RuntimeError):
pass
# ------------------------------------------------------------------ downloads
def require_under_home(path: Path) -> Path:
"""`path`, resolved, if it is inside the user's home; otherwise refuse.
The snap sandbox will not write elsewhere, and a download folder under
`/tmp` would also put evidence where a reboot deletes it.
"""
resolved = Path(path).expanduser().resolve()
home = Path.home().resolve()
if resolved != home and home not in resolved.parents:
raise WebDriverError(f"{resolved} is not under {home}; the browser cannot write there")
return resolved
def firefox_download_prefs(directory: Path) -> dict:
"""Profile preferences that save every download to `directory`, unasked."""
return {
"browser.download.folderList": 2, # 2 = the folder named below
"browser.download.dir": str(directory),
"browser.download.useDownloadDir": True,
"browser.download.start_downloads_in_tmp_dir": False,
"browser.download.always_ask_before_handling_new_types": False,
"browser.helperApps.neverAsk.saveToDisk": "application/json,application/octet-stream",
"browser.download.manager.showWhenStarting": False,
"browser.download.alwaysOpenPanel": False,
"browser.download.panel.shown": True,
}
def wait_for_download(directory: Path, before: set[str], *, timeout: float = 60,
poll: float = 0.2, stable_polls: int = 3) -> Path:
"""The file a download wrote into `directory`, once it has finished.
Finished means all of these, at once:
- a name that was not in `before` (the listing taken before the click);
- no in-progress file (`*.part`) left in the folder;
- more than zero bytes;
- the same size for `stable_polls` consecutive polls.
A first appearance is not a finished download, and a zero-byte or partial
file never counts. Raises `WebDriverError` when nothing finishes in time.
"""
deadline = time.monotonic() + timeout
last: dict[str, int] = {}
steady: dict[str, int] = {}
while time.monotonic() < deadline:
names = {p.name for p in directory.iterdir()} if directory.exists() else set()
partial = any(n.endswith(PARTIAL_SUFFIXES) for n in names)
fresh = sorted(n for n in names - before if not n.endswith(PARTIAL_SUFFIXES))
for name in fresh:
size = (directory / name).stat().st_size
steady[name] = steady.get(name, 0) + 1 if last.get(name) == size else 1
last[name] = size
if not partial and size > 0 and steady[name] >= stable_polls:
return directory / name
time.sleep(poll)
listing = sorted(p.name for p in directory.iterdir()) if directory.exists() else []
raise WebDriverError(f"no finished download in {directory} within {timeout}s; saw {listing}")
# -------------------------------------------------------------------- browser
class Browser:
"""One headless Firefox, driven over the wire protocol."""
def __init__(self, *, headless: bool = True, log: Path | None = None):
def __init__(self, *, headless: bool = True, log: Path | None = None,
download_dir: Path | None = None):
self.port = free_port()
handle = open(log, "ab") if log else subprocess.DEVNULL
self.proc = subprocess.Popen(
@@ -68,9 +151,15 @@ class Browser:
self.base = f"http://127.0.0.1:{self.port}"
self._wait_for_driver()
args = ["-headless"] if headless else []
options: dict = {"args": args}
self.download_dir = None
if download_dir is not None:
self.download_dir = require_under_home(download_dir)
self.download_dir.mkdir(parents=True, exist_ok=True)
options["prefs"] = firefox_download_prefs(self.download_dir)
answer = self._call("POST", "/session", {"capabilities": {"alwaysMatch": {
"browserName": "firefox",
"moz:firefoxOptions": {"args": args},
"moz:firefoxOptions": options,
# Never silently accept a bad certificate: the endpoint policy and
# the TLS trust union are release claims (H12, A06), and a browser
# that ignored certificates would hide a failure of either.
@@ -78,6 +167,7 @@ class Browser:
}}})["value"]
self.session = answer["sessionId"]
self.version = answer["capabilities"].get("browserVersion", "?")
self.capabilities = answer["capabilities"]
# ------------------------------------------------------------- plumbing
@@ -123,6 +213,9 @@ class Browser:
def go(self, url: str) -> None:
self._call("POST", self._s("/url"), {"url": url})
def reload(self) -> None:
self._call("POST", self._s("/refresh"), {})
@property
def url(self) -> str:
return self._call("GET", self._s("/url"))["value"]
@@ -134,10 +227,56 @@ class Browser:
def source(self) -> str:
return self._call("GET", self._s("/source"))["value"]
def screenshot(self, path) -> Path:
"""The viewport as a PNG, written where you ask (v1.1 WP-E).
Evidence for a change a reader judges by looking at it: a contrast ratio
says a boundary is measurable, and a picture says what it looks like.
"""
encoded = self._call("GET", self._s("/screenshot"))["value"]
target = Path(path)
target.parent.mkdir(parents=True, exist_ok=True)
target.write_bytes(base64.b64decode(encoded))
return target
def hover(self, element: str) -> None:
"""A real pointer over an element, so `:hover` actually applies.
Dispatching a mouseover event from JavaScript does not do this: CSS
`:hover` follows the browser's own pointer state, not a synthetic event,
so a measurement taken after `dispatchEvent` reads the resting style and
reports it as the hover style. This moves the pointer (v1.1 WP-E).
"""
self._call("POST", self._s("/execute/sync"), {
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
"args": [{ELEMENT_KEY: element}]})
self._call("POST", self._s("/actions"), {"actions": [{
"type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
"actions": [{"type": "pointerMove", "duration": 60,
"origin": {ELEMENT_KEY: element}, "x": 0, "y": 0}]}]})
def unhover(self) -> None:
"""Move the pointer off whatever it was over, and forget the input state."""
self._call("POST", self._s("/actions"), {"actions": [{
"type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
"actions": [{"type": "pointerMove", "duration": 30,
"origin": "viewport", "x": 0, "y": 0}]}]})
try:
self._call("DELETE", self._s("/actions"))
except WebDriverError:
pass
def js(self, script: str, *args):
return self._call("POST", self._s("/execute/sync"),
{"script": script, "args": list(args)})["value"]
def element_by_js(self, script: str, *args):
"""An element a script returns, as a reference `click` can use, or None."""
value = self.js(script, *args)
if isinstance(value, dict) and ELEMENT_KEY in value:
return value[ELEMENT_KEY]
return None
def find(self, css: str, *, required=True):
try:
answer = self._call("POST", self._s("/element"),
@@ -163,6 +302,17 @@ class Browser:
return self._call("GET", self._s(f"/element/{element}/property/{name}"))["value"]
def click(self, element: str) -> None:
"""A real click, on an element first scrolled to the middle of the view.
WebDriver scrolls a target only as far as its edge, and the play page's
composer is fixed to the bottom of the window: a control just under it
(a failure notice's details, a turn's Inspect button) is then covered,
and the click is intercepted. A reader scrolls it clear first; so does
this (v1.1 WP-C).
"""
self._call("POST", self._s("/execute/sync"), {
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
"args": [{ELEMENT_KEY: element}]})
self._call("POST", self._s(f"/element/{element}/click"), {})
def clear(self, element: str) -> None:
@@ -183,6 +333,21 @@ class Browser:
answer = self._call("GET", self._s("/element/active"))
return list(answer["value"].values())[0]
# -------------------------------------------------------------- windows
@property
def window(self) -> str:
return self._call("GET", self._s("/window"))["value"]
def new_tab(self) -> str:
return self._call("POST", self._s("/window/new"), {"type": "tab"})["value"]["handle"]
def switch_to(self, handle: str) -> None:
self._call("POST", self._s("/window"), {"handle": handle})
def close_window(self) -> None:
self._call("DELETE", self._s("/window"))
# ------------------------------------------------------------- waiting
def wait_for(self, css: str, *, timeout=90, gone=False):
@@ -196,12 +361,19 @@ class Browser:
f"{'still present' if gone else 'never appeared'}: {css}")
def wait_until(self, script: str, *, timeout=90, what=""):
if self.wait_js(script, timeout=timeout):
return True
raise WebDriverError(f"condition never held: {what or script}")
def wait_js(self, script: str, *, timeout=90) -> bool:
"""Whether `script` became true within `timeout`. For a check to record,
where `wait_until` is for a precondition that must hold."""
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
if self.js(f"return ({script})"):
return True
time.sleep(0.25)
raise WebDriverError(f"condition never held: {what or script}")
return False
class Site:
+960
View File
@@ -0,0 +1,960 @@
"""v1.1 WP-B.1: where an early story fact is lost on its way to the narrator.
One planted fact **F** has four stages to survive before the narrator can use
it from memory, and this module reports each one separately:
created a memory whose `source_start`..`source_end` covers the planting
depth carries F
retained that memory is not `forgotten`
ranked it is eligible on the active lineage and embedded, and where it
scores for the recall query against `memory_top_k`
injected the recall turn's own stored `memories.used` names it, and its text
is in that turn's `used_memories` section
A fact is only evidence about memory if memory is the **only** thing carrying it.
`isolation()` checks every other layer: the authoritative document, per-node state
snapshots, the active summary, imported knowledge, the narration after the
planting block, and the recent-history window. A run where any of those carries F
is reported as a failed precondition, never as a memory result.
**Nothing here changes behaviour.**
- It reads rows.
- It reuses production's own pure helpers (`memorybank._drop_redundant`,
`memorybank.classify_authority`, `vectors.cosine`, `lineage.path_of`), so its
ranking is production's ranking, not a second opinion.
- It checks itself against what the recall turn actually recorded.
- The only computed fields are ephemeral report data. No column or table is
added.
The deterministic stubs at the bottom stand in for the models when a test needs a
fixed answer. **Read what they model before reading any result they produce:**
- `BestCaseSummariser` keeps F if and only if F is in the excerpt it is given.
It is the ideal summariser, so a creation failure under it is the
application's, not the model's.
- `ConceptEmbedder` maps words to a small concept table, so that "the brass dial
that tells the hour" lands near "sundial". It models what an embedding is
supposed to do. It says nothing about how well `nomic-embed-text` does it,
which is what the real-model run is for.
"""
from __future__ import annotations
import hashlib
import json
import math
import re
from dataclasses import dataclass, field
from sqlalchemy import select
from app import memorybank, models, summaries, vectors
from app.context import builder, history, lineage
from app.knowledge import classes as knowledge_classes
VERDICTS = (
"not_created",
"created_but_evicted",
"retained_but_not_ranked",
"ranked_but_not_selected",
"selected_but_not_injected",
"injected",
)
#: Section labels in a stored context snapshot. Copied from the builder's
#: vocabulary so a renamed section fails loudly here.
HISTORY_LABELS = ("history", "recent_history")
SUMMARY_LABEL = "story_summary"
MEMORIES_LABEL = "used_memories"
STATE_LABEL = "narrative_state"
KNOWLEDGE_LABELS = (
knowledge_classes.SECTION_CANON,
knowledge_classes.SECTION_REFERENCE,
knowledge_classes.SECTION_INSPIRATION,
)
@dataclass(frozen=True)
class Fact:
"""A planted fact, and how to recognise it in a text.
`carry_groups`: a text carries the fact when every group matches, where a
group matches when any one of its terms appears as a whole word. A memory has
to name both the thing and where it is to carry "where the thing is".
`leak_terms`: any one of these in another layer means that layer carries the
fact. This is deliberately looser than `carry_groups`. For isolation, a
mention is enough to disqualify.
"""
fact_id: str
sentence: str
carry_groups: tuple[tuple[str, ...], ...]
leak_terms: tuple[str, ...]
def carried_by(self, text: str | None) -> bool:
low = (text or "").lower()
return all(any(_has_word(low, term) for term in group) for group in self.carry_groups)
def mentioned_by(self, text: str | None) -> bool:
low = (text or "").lower()
return any(_has_word(low, term) for term in self.leak_terms)
def _has_word(low: str, term: str) -> bool:
return re.search(rf"(?<![a-z]){re.escape(term.lower())}(?![a-z])", low) is not None
#: The fixture's planted fact. Chosen to be natural in a tavern scene and absent
#: from every existing fixture: no "sundial" or "teapot" appears anywhere in the
#: Westhaven campaign, its knowledge files or its beats.
FACT_F = Fact(
fact_id="F-amber-sundial",
sentence="Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf.",
carry_groups=(("sundial",), ("teapot",)),
leak_terms=("sundial", "teapot"),
)
#: The abandoned-line control fact.
FACT_G = Fact(
fact_id="G-iron-weathervane",
sentence="Edrin buried the iron weathervane beneath the mill's broken waterwheel.",
carry_groups=(("weathervane",), ("waterwheel",)),
leak_terms=("weathervane", "waterwheel"),
)
# ------------------------------------------------------------------ reading
def _lineage_actions(db, adventure):
path = lineage.path_of(db, adventure)
return (
db.query(models.Action)
.filter(models.Action.adventure_id == adventure.id, path.clause(models.Action))
.order_by(models.Action.depth, models.Action.id)
.all()
)
def covering_memories(db, adventure, depth: int, *, any_branch: bool = False):
"""Memories whose source range covers `depth`, oldest first."""
query = select(models.Memory).where(
models.Memory.adventure_id == adventure.id,
models.Memory.source_start <= depth,
models.Memory.source_end >= depth,
)
if not any_branch:
query = query.where(lineage.path_of(db, adventure).clause(models.Memory))
return db.execute(query.order_by(models.Memory.id)).scalars().all()
def planting_block_end(db, adventure, plant_depth: int) -> int:
"""The last depth of the memory block holding the planted turn.
Taken from the memory that covers it where one exists. Before one exists it
is the furthest a block could reach, so a later-narration check never counts
a turn inside the planting block as a repetition.
"""
rows = covering_memories(db, adventure, plant_depth)
if rows:
return max(row.source_end for row in rows)
return plant_depth + memorybank.MEMORY_INTERVAL
# ---------------------------------------------------------------- isolation
def isolation(db, adventure, fact: Fact, plant_depth: int, *,
recall_snapshot: dict | None = None,
recall_depth: int | None = None) -> dict:
"""Every layer other than memory that could carry F, checked.
Returns `{check: {"ok": bool, "detail": str}}` and `ok` over all of them.
With `recall_snapshot`, the recall turn's stored context, the prompt-level
checks (history window, summary section, knowledge sections) are made
against what the narrator was actually given.
"""
checks: dict[str, dict] = {}
document = adventure.narrative_state or {}
hits = [key for key in ("entities", "facts", "relationships", "threads", "scene",
"possessions")
if fact.mentioned_by(json.dumps(document.get(key), default=str))]
checks["state_document"] = {
"ok": not hits and not fact.mentioned_by(json.dumps(document, default=str)),
"detail": f"mentioned in {hits}" if hits else "absent",
}
snapshot_hits = []
later_hits = []
block_end = planting_block_end(db, adventure, plant_depth)
for action in _lineage_actions(db, adventure):
if fact.mentioned_by(json.dumps(action.narrative_state_after, default=str)):
snapshot_hits.append(action.depth)
if (action.type == "ai" and action.depth is not None and action.depth > block_end
and (recall_depth is None or action.depth < recall_depth)
and fact.mentioned_by(action.text)):
later_hits.append(action.depth)
checks["state_snapshots"] = {
"ok": not snapshot_hits,
"detail": f"mentioned in snapshots at depths {snapshot_hits[:10]}" if snapshot_hits
else "absent from every node's narrative_state_after on the active lineage",
}
checks["later_narration"] = {
"ok": not later_hits,
"detail": (f"narration after the planting block (ends at depth {block_end}) "
f"mentions the fact at depths {later_hits[:10]}") if later_hits
else f"no narrator turn after depth {block_end} mentions the fact",
}
active = summaries.current(db, adventure)
summary_text = active.text if active is not None else ""
if recall_snapshot is not None:
summary_text += "\n" + _section(recall_snapshot, SUMMARY_LABEL)
checks["summary"] = {
"ok": not fact.mentioned_by(summary_text),
"detail": "the active summary mentions the fact" if fact.mentioned_by(summary_text)
else ("absent from the active summary" if active is not None else "no summary yet"),
}
sources = db.execute(
select(models.KnowledgeSource.content).where(
models.KnowledgeSource.adventure_id == adventure.id)
).scalars().all()
knowledge_text = "\n".join(s or "" for s in sources)
if recall_snapshot is not None:
knowledge_text += "\n" + "\n".join(_section(recall_snapshot, l) for l in KNOWLEDGE_LABELS)
checks["knowledge"] = {
"ok": not fact.mentioned_by(knowledge_text),
"detail": "imported knowledge mentions the fact" if fact.mentioned_by(knowledge_text)
else f"absent from {len(sources)} imported source(s)",
}
if recall_snapshot is not None:
hist = recall_snapshot.get("history") or {}
floor = hist.get("floor_depth")
history_text = "\n".join(_section(recall_snapshot, l) for l in HISTORY_LABELS)
outside = floor is not None and plant_depth < floor
checks["recent_history"] = {
"ok": outside and not fact.carried_by(history_text),
"detail": (f"history window starts at depth {floor}; planted at {plant_depth}; "
f"fact text in history sections: {fact.carried_by(history_text)}"),
}
checks["state_section"] = {
"ok": not fact.mentioned_by(_section(recall_snapshot, STATE_LABEL)),
"detail": "the recall prompt's narrative_state section "
+ ("mentions the fact" if fact.mentioned_by(_section(recall_snapshot, STATE_LABEL))
else "does not mention the fact"),
}
return {"ok": all(c["ok"] for c in checks.values()), "checks": checks}
def _section(snapshot: dict, label: str) -> str:
return "\n".join(s.get("text", "") for s in (snapshot.get("sections") or [])
if s.get("label") == label)
# ------------------------------------------------------------------- stages
async def rank_bank(db, adventure, settings, query: dict, embed) -> dict:
"""Production's ranking, recomputed for `query`, for every eligible memory.
`query` is a `memorybank.retrieval_query` dict. The catalogue clause, the
scoring (`memorybank.score_candidates`) and the selection with its pins and
redundancy rule (`memorybank.select_memories`) are production's own
functions, so this is production's ranking, not a second opinion. Returns
every scored row, not just the top-k, because "where did F rank" is the
question.
"""
catalogue = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.authority,
models.Memory.embedding_blob, models.Memory.text).where(
models.Memory.adventure_id == adventure.id,
lineage.path_of(db, adventure).clause(models.Memory),
models.Memory.forgotten.is_(False),
models.Memory.embedded.is_(True),
)
).all()
top_k = max(1, settings.memory_top_k)
texts = [t for t in (query["input"], query["context"]) if t.strip()]
if not catalogue or not texts:
return {"query": query, "scored": [], "selected": [], "top_k": top_k}
vectors_by_text = dict(zip(texts, await embed(texts)))
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
held = {row.id: vectors.unpack(row.embedding_blob) for row in catalogue if row.embedding_blob}
terms_of = ({row.id: memorybank.lexical_terms(row.text or "") for row in catalogue}
if query["input_terms"] else {})
authority_of = {row.id: row.authority for row in catalogue}
pinned_of = {row.id: row.pinned for row in catalogue}
scored = memorybank.score_candidates(
[row.id for row in catalogue if row.id in held], held, terms_of,
input_vec, context_vec, query["input_terms"])
used, suppressed = memorybank.select_memories(scored, pinned_of, held, authority_of, top_k)
selected = {row[1] for row in used}
suppressed_by = dict(suppressed)
return {
"query": query,
"top_k": top_k,
"scored": [
{"rank": i + 1, "memory_id": memory_id, "similarity": round(semantic, 4),
"semantic_score": round(semantic, 4), "lexical_score": round(lexical, 4),
"final_score": round(final, 4),
"pinned": pinned_of[memory_id], "selected": memory_id in selected,
"suppressed_as_duplicate_of": suppressed_by.get(memory_id)}
for i, (final, memory_id, semantic, lexical) in enumerate(scored)
],
"selected": sorted(selected),
}
def production_query(adventure, exclude_action_id: int | None) -> dict:
"""The retrieval query a turn used, built by production's own `retrieval_query`."""
return memorybank.retrieval_query(adventure, exclude_action_id)
def variant_query(base: dict, player_input: str) -> dict:
"""`base` with a different player input: "what if the player had asked this
here", with the scene and narration context the recall turn really had."""
return {"input": player_input, "context": base["context"],
"input_terms": sorted(memorybank.lexical_terms(player_input))}
def eviction_order(db, adventure) -> list[int]:
"""The order `_evict_over_capacity` would take unpinned active memories in.
Production's own `memorybank.eviction_order`, run to the end of the bank."""
rows = db.execute(
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
models.Memory.source_end, models.Memory.last_used_at,
models.Memory.created_at, models.Memory.use_count).where(
models.Memory.adventure_id == adventure.id,
models.Memory.forgotten.is_(False),
)
).all()
return memorybank.eviction_order(rows, len(rows))
def coverage(ranges: list[tuple[int, int]], tip: int | None) -> dict:
"""How much of the story `ranges` (active memories' source ranges) describe.
`largest_gap` is the longest run of depths, between the first memory's start
and `tip`, that no memory covers. It is how the eviction rule is judged in
general, not only for the planted fact."""
if not ranges:
return {"first_start": None, "last_end": None, "largest_gap": None}
ordered = sorted(ranges)
gaps = []
reach = ordered[0][1]
for start, end in ordered[1:]:
gaps.append(max(0, start - reach - 1))
reach = max(reach, end)
return {"first_start": ordered[0][0], "last_end": reach,
"largest_gap": max(gaps, default=0)}
async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
recall_action: models.Action, embed) -> dict:
"""The four stages for `fact`, judged at `recall_action`, the recall turn's AI node.
Ranking is recomputed with the query that turn used, and checked against the
turn's own stored `memories.used`. Injection is read from that snapshot, so
it reports what the narrator was actually given, not a re-run.
"""
snapshot = recall_action.context_snapshot or {}
out: dict = {"fact_id": fact.fact_id, "plant_depth": plant_depth,
"recall_depth": recall_action.depth}
covering = covering_memories(db, adventure, plant_depth)
carrying = [m for m in covering if fact.carried_by(m.text)]
elsewhere = [m for m in db.execute(select(models.Memory).where(
models.Memory.adventure_id == adventure.id)).scalars().all()
if fact.carried_by(m.text) and m not in carrying]
creation_input = []
for memory in covering:
block = memorybank.source_block(db, memory)
raw = "\n\n".join(a.text for a in block)
excerpt = memorybank.memory_excerpt(raw) # what `summarize_block` sends
creation_input.append({
"memory_id": memory.id, "source_start": memory.source_start,
"source_end": memory.source_end, "block_tokens": builder.count_tokens(raw),
"fact_in_block": fact.carried_by(raw),
"fact_in_summariser_excerpt": fact.carried_by(excerpt),
"memory_text": memory.text,
})
memory = carrying[0] if carrying else None
out["created"] = {
"yes": memory is not None,
"memory_id": getattr(memory, "id", None),
"source_start": getattr(memory, "source_start", None),
"source_end": getattr(memory, "source_end", None),
"memory_text": getattr(memory, "text", None),
"covering_memories": creation_input,
"no_covering_memory": not covering,
"carried_by_other_memories": [
{"memory_id": m.id, "source_start": m.source_start, "source_end": m.source_end}
for m in elsewhere],
}
if memory is None:
out["verdict"] = "not_created"
return out
order = eviction_order(db, adventure)
active = db.execute(select(models.Memory.id).where(
models.Memory.adventure_id == adventure.id,
models.Memory.forgotten.is_(False))).scalars().all()
on_lineage = db.execute(select(models.Memory.id).where(
models.Memory.id == memory.id,
lineage.path_of(db, adventure).clause(models.Memory))).scalar() is not None
out["retained"] = {
"yes": not memory.forgotten,
"forgotten": memory.forgotten,
"pinned": memory.pinned,
"embedded": memory.embedded,
"on_active_lineage": on_lineage,
"use_count": memory.use_count,
"last_used_at": str(memory.last_used_at) if memory.last_used_at else None,
"created_at": str(memory.created_at),
"active_memories": len(active),
"memory_bank_capacity": settings.memory_bank_capacity,
"eviction_position": (order.index(memory.id) + 1) if memory.id in order else None,
"reason": ("evicted: marked forgotten by capacity eviction" if memory.forgotten
else "active"),
}
if memory.forgotten:
out["verdict"] = "created_but_evicted"
return out
# The recall turn's AI node is excluded, so the newest action is the recall
# player action, exactly as the turn saw it before it wrote its reply.
query = production_query(adventure, recall_action.id)
ranking = await rank_bank(db, adventure, settings, query, embed)
row = next((r for r in ranking["scored"] if r["memory_id"] == memory.id), None)
stored_used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
out["ranked"] = {
"yes": row is not None and row["rank"] <= ranking["top_k"],
"eligible": row is not None,
"lexical_score": row["lexical_score"] if row else None,
"semantic_score": row["semantic_score"] if row else None,
"final_score": row["final_score"] if row else None,
"selected_top_k": ranking["selected"],
"pin_effect": "always selected" if memory.pinned else "none",
"rank": row["rank"] if row else None,
"of": len(ranking["scored"]),
"top_k_cutoff": ranking["top_k"],
"selected": bool(row and row["selected"]),
"suppressed_as_duplicate_of": row["suppressed_as_duplicate_of"] if row else None,
"query": query,
"replica_matches_stored_selection": sorted(stored_used) == ranking["selected"],
}
if row is None or row["rank"] > ranking["top_k"] and not row["selected"]:
out["verdict"] = "retained_but_not_ranked"
return out
if not row["selected"]:
out["verdict"] = "ranked_but_not_selected"
return out
section = _section(snapshot, MEMORIES_LABEL)
injected = memory.id in stored_used and memory.text in section
out["injected"] = {
"yes": injected,
"context_component": MEMORIES_LABEL,
"in_stored_memories_used": memory.id in stored_used,
"text_in_section": memory.text in section,
"token_count": builder.count_tokens(section) if section else 0,
}
out["verdict"] = "injected" if injected else "selected_but_not_injected"
return out
# ---------------------------------------------------------------- the stubs
@dataclass
class BestCaseSummariser:
"""The ideal memory writer: F survives if, and only if, F reached it.
A memory keeps every sentence of the excerpt that carries a planted fact, and
adds one sentence naming the block's own distinct detail so memories differ.
Summary updates never repeat a planted fact, so the summary layer stays out of
the experiment. Every excerpt it was given is kept, for the creation-window
diagnostic.
"""
facts: tuple[Fact, ...] = (FACT_F, FACT_G)
excerpts: list = field(default_factory=list)
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
if "Current story summary:" in user:
return "The travellers kept moving through the country around Westhaven."
excerpt = user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
self.excerpts.append(excerpt)
kept = [s.strip() for s in re.split(r"(?<=[.!?])\s+", excerpt)
if any(f.carried_by(s) for f in self.facts)]
detail = re.findall(r"\bat the ([a-z]+ [a-z]+)\b", excerpt.lower())
tail = f"The travellers spent time at the {detail[-1]}." if detail else \
"The travellers pressed on."
return " ".join(dict.fromkeys(kept + [tail]))
#: Words that mean the same thing to `ConceptEmbedder`. The point is only that a
#: paraphrase lands near the original; the table is the model of that.
CONCEPTS = {
"timepiece": ("sundial", "dial", "hour", "hours", "clock", "timepiece"),
"vessel": ("teapot", "pot", "kettle", "tea", "jar"),
"hid": ("hid", "hide", "hidden", "slipped", "tucked", "put", "stashed"),
"weathervane": ("weathervane", "vane"),
"waterwheel": ("waterwheel", "wheel", "mill"),
}
_WORD_TO_CONCEPT = {w: c for c, words in CONCEPTS.items() for w in words}
DIMENSIONS = 96
@dataclass
class ConceptEmbedder:
"""A deterministic embedding: concepts in fixed dimensions, other words hashed."""
calls: int = 0
async def embed(self, texts):
self.calls += 1
return [self.vector(t) for t in texts]
@staticmethod
def vector(text: str) -> list[float]:
v = [0.0] * DIMENSIONS
v[0] = 0.2 # every text shares a little, as real embeddings do
concept_names = list(CONCEPTS)
for word in re.findall(r"[a-z]+", text.lower()):
concept = _WORD_TO_CONCEPT.get(word)
if concept is not None:
v[1 + concept_names.index(concept)] += 3.0
elif len(word) > 3:
bucket = int(hashlib.sha256(word.encode()).hexdigest(), 16)
v[1 + len(concept_names) + bucket % (DIMENSIONS - 1 - len(concept_names))] += 1.0
norm = math.sqrt(sum(x * x for x in v)) or 1.0
return [x / norm for x in v]
# ---------------------------------------------------------------- scenarios
#: Filler places. No word here is in `CONCEPTS`, and none names a planted fact.
PLACES = (
"north gate", "salt market", "ferry landing", "chapel steps", "rope walk",
"fish stalls", "old bridge", "tanner yard", "lamp street", "weir path",
"grain store", "boat yard", "watch house", "cloth hall", "eel traps",
"sheep fold", "smith forge", "stone quay", "reed beds", "toll booth",
)
PARAPHRASE_QUERY = "I ask Mara where she tucked the little brass dial that tells the hour."
UNRELATED_QUERY = "I ask the ferryman what rope costs at the landing this season."
#: Built only from words every fixture memory holds ("travellers", "spent",
#: "time"), so its rarity weight is zero everywhere.
COMMON_WORDS_QUERY = "The travellers spent time."
def filler_prose(index: int, words: int) -> str:
"""Narration that moves on and never touches a planted fact."""
place = PLACES[index % len(PLACES)]
sentence = (f"At the {place} the travellers stopped, listened to the gulls over the "
f"grey water, and talked about the long road north.")
reps = max(1, round(words / len(sentence.split())))
return " ".join([sentence] * reps)
@dataclass
class Scenario:
"""One deterministic campaign. Depths: the opening is 0, turn *n*'s player
action is 2n-1 and its reply 2n."""
name: str
turns: int = 52
capacity: int = 80
top_k: int = 5
budget: int = 4096
prose_words: int = 60
plant_turn: int = 1
recall_text: str = "I ask Mara where she hid the amber sundial."
pin_first_memory: bool = False
lineage_control: bool = False
diagnose_recall: bool = True
#: The narration of the last turn before recall, when a fixture needs the
#: scene to say something (WP-B.2's context-dependent question). Must not
#: name a planted fact.
pre_recall_reply: str = ""
SCENARIOS = {
"independent_default": Scenario("independent_default"),
"past_capacity": Scenario("past_capacity", capacity=6),
"past_capacity_pinned": Scenario("past_capacity_pinned", capacity=6, pin_first_memory=True),
# Closer to the shipped ratio (memory_top_k 5 against capacity 80): most of
# the bank is not retrieved on a given turn.
"past_capacity_low_top_k": Scenario("past_capacity_low_top_k", capacity=8, top_k=2),
"long_block_fact_early": Scenario("long_block_fact_early", turns=10, prose_words=850,
plant_turn=1, budget=16384),
"long_block_fact_late": Scenario("long_block_fact_late", turns=10, prose_words=850,
plant_turn=3, budget=16384),
"lineage_control": Scenario("lineage_control", lineage_control=True),
# v1.1 WP-B.2: the ranking failure B.1 saw on the real model, made
# deterministic. Longer narration fills the v1.0.0 query, and `memory_top_k`
# is the real run's 4. Below capacity, isolation valid.
"ranking_crowded": Scenario("ranking_crowded", prose_words=150, top_k=4),
# v1.1 WP-B.2: all three B.1 failures at once. Every block is longer than the
# summariser's excerpt, narration crowds the query at `memory_top_k` 4, and
# the bank passes a capacity of 8 long before recall at depth 106.
"independent_full": Scenario("independent_full", prose_words=850, top_k=4,
capacity=8, budget=16384),
# v1.1 WP-B.2: a question that names neither the sundial nor the teapot and
# cannot be answered without the scene. The last narration puts Mara at the
# tavern's top shelf; the player asks "her" what she put "up there".
"ranking_context_dependent": Scenario(
"ranking_context_dependent", prose_words=150, top_k=4,
recall_text="I ask her what she keeps up there.",
pre_recall_reply=("Mara stands on a stool at the tavern's top shelf, running a cloth "
"around the old kettle up there, and she will not meet your eye.")),
}
class ScriptNarrator:
"""Stands in for the narrator: returns `next_reply`, with an empty state block."""
next_reply = ""
last_usage = None
prompts: list = []
def __init__(self, *a, **k):
pass
async def generate(self, parts, *, temperature, max_tokens):
ScriptNarrator.prompts.append((parts.system, parts.story))
yield ("text", ScriptNarrator.next_reply)
def run_scenario(scenario: Scenario) -> dict:
"""Plays `scenario` through the real turn route and returns everything measured.
Uses the database `app.database` is already bound to, creating and dropping
its tables, the way the suite's fixtures do. Patches are applied here and
removed before returning, so this runs the same under pytest and from the CLI.
"""
import asyncio
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, limits
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
from app.routers import adventures as adventure_routes
summariser = BestCaseSummariser()
embedder = ConceptEmbedder()
patches = [
(memorybank, "summary_provider", lambda s: summariser),
(memorybank, "embedding_provider", lambda s: embedder),
# Post-turn work is settled explicitly after each turn, so eviction
# happens at a known point rather than whenever a background task runs.
(memorybank, "schedule_post_turn", lambda adventure: None),
(adventure_routes.turns, "OpenAICompatibleProvider", ScriptNarrator),
(limits, "check_row_cap", lambda *a, **k: None),
]
saved = [(obj, name, getattr(obj, name)) for obj, name, _ in patches]
for obj, name, value in patches:
setattr(obj, name, value)
ScriptNarrator.prompts = []
Base.metadata.create_all(bind=engine)
memorybank._vector_cache.clear()
with SessionLocal() as db:
user = models.User(is_guest=False, email=f"b1-{scenario.name}@example.com")
db.add(user)
db.flush()
db.add(models.Settings(
user_id=user.id, model="script", endpoint_url="http://127.0.0.1:9/v1",
embedding_model="concept-embed", context_token_budget=scenario.budget,
max_output_tokens=500, memory_bank_capacity=scenario.capacity,
memory_top_k=scenario.top_k,
))
adventure = models.Adventure(
user_id=user.id, title=f"B.1 {scenario.name}", memory_bank_enabled=True,
auto_summarize=True, persona_name="Aldric",
)
db.add(adventure)
db.flush()
db.add(models.Action(adventure_id=adventure.id, type="start",
text="Rain over Westhaven, and the tavern door banging in the wind."))
db.commit()
adv, user_id = adventure.id, user.id
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
client = TestClient(app)
result: dict = {"scenario": scenario.__dict__.copy(), "trace": []}
def call(method, path, body=None, expect=200):
response = client.request(method, f"/api/adventures/{adv}{path}", json=body)
assert response.status_code == expect, (path, response.status_code, response.text[:300])
return response.json() if response.content else None
def marks():
with SessionLocal() as db:
rows = db.execute(select(models.Memory.id, models.Memory.forgotten,
models.Memory.embedded).where(
models.Memory.adventure_id == adv)).all()
summaries_n = db.query(models.Summary).filter_by(adventure_id=adv).count()
return tuple(sorted(rows)), summaries_n
def settle():
for _ in range(12):
before = marks()
asyncio.run(memorybank.run_post_turn(adv))
if marks() == before:
return
def memories():
with SessionLocal() as db:
return [dict(row._mapping) for row in db.execute(select(
models.Memory.id, models.Memory.text, models.Memory.source_start,
models.Memory.source_end, models.Memory.forgotten, models.Memory.pinned,
models.Memory.use_count, models.Memory.last_used_at, models.Memory.branch_id,
models.Memory.created_at).where(models.Memory.adventure_id == adv)
.order_by(models.Memory.id)).all()]
def per_turn_isolation(adventure_id):
with SessionLocal() as db:
adventure = db.get(models.Adventure, adventure_id)
active = summaries.current(db, adventure)
return {
"f_in_state": FACT_F.mentioned_by(json.dumps(adventure.narrative_state or {},
default=str)),
"f_in_summary": FACT_F.mentioned_by(active.text if active is not None else ""),
}
def turn(kind, text, reply):
ScriptNarrator.next_reply = f"{reply}\n```state\n{{\"events\": []}}\n```"
response = client.post(f"/api/adventures/{adv}/actions", json={"type": kind, "text": text})
assert response.status_code == 200, response.text[:300]
assert '"type": "error"' not in response.text, response.text[-300:]
plant_depth = None
f_memory_id = None
pinned_id = None
known: dict[int, dict] = {}
g: dict = {}
try:
for n in range(1, scenario.turns + 1):
if n == scenario.plant_turn:
turn("story", FACT_F.sentence, filler_prose(n, scenario.prose_words))
with SessionLocal() as db:
plant_depth = db.query(models.Action.depth).filter_by(
adventure_id=adv, text=FACT_F.sentence).scalar()
elif scenario.lineage_control and n == 21:
call("POST", "/checkpoints", {"name": "before the mill"}, expect=201)
turn("story", FACT_G.sentence, filler_prose(n, scenario.prose_words))
with SessionLocal() as db:
g["plant_depth"] = db.query(models.Action.depth).filter_by(
adventure_id=adv, text=FACT_G.sentence).scalar()
elif scenario.lineage_control and n == 30:
# Line A carries G's memory. Mark it, then abandon it: Undo back
# to before G was planted and write something else.
g["line_a"] = call("POST", "/checkpoints", {"name": "line A, after the mill"},
expect=201)["id"]
with SessionLocal() as db:
g_rows = [m for m in db.execute(select(models.Memory).where(
models.Memory.adventure_id == adv)).scalars() if FACT_G.carried_by(m.text)]
g["memory_ids"] = [m.id for m in g_rows]
with SessionLocal() as db:
g["last_action_id_before_divergence"] = db.query(models.Action.id).filter_by(
adventure_id=adv).order_by(models.Action.id.desc()).limit(1).scalar()
for _ in range(9):
call("POST", "/undo")
turn("do", f"I turn away from the mill and walk to the {PLACES[n % len(PLACES)]}.",
filler_prose(n + 100, scenario.prose_words))
g["diverged_at_turn"] = n
elif n == scenario.turns and scenario.pre_recall_reply:
turn("do", "I head back to the tavern.", scenario.pre_recall_reply)
else:
turn("do", f"I walk on to the {PLACES[n % len(PLACES)]}.",
filler_prose(n, scenario.prose_words))
settle()
rows = memories()
created = [r["id"] for r in rows if r["id"] not in known]
newly_forgotten = [r["id"] for r in rows
if r["forgotten"] and not known.get(r["id"], {}).get("forgotten")]
for r in rows:
known[r["id"]] = r
if f_memory_id is None and plant_depth is not None:
for r in rows:
if (r["source_start"] is not None and r["source_start"] <= plant_depth
<= r["source_end"] and FACT_F.carried_by(r["text"])):
f_memory_id = r["id"]
if scenario.pin_first_memory and pinned_id is None:
candidate = next((r for r in rows if r["id"] != f_memory_id), None)
if candidate is not None:
call("PATCH", f"/memories/{candidate['id']}", {"pinned": True})
pinned_id = candidate["id"]
f_row = known.get(f_memory_id) if f_memory_id else None
result["trace"].append({
"turn": n,
"active": sum(1 for r in rows if not r["forgotten"]),
"total": len(rows),
"created": created,
"evicted": newly_forgotten,
"created_and_evicted_same_turn": sorted(set(created) & set(newly_forgotten)),
"f_memory_id": f_memory_id,
"f_forgotten": bool(f_row and f_row["forgotten"]),
"f_use_count": f_row["use_count"] if f_row else None,
"coverage": coverage([(r["source_start"], r["source_end"]) for r in rows
if not r["forgotten"] and r["source_start"] is not None],
None),
# Isolation on every turn, not only at recall (WP-B.2 acceptance).
**per_turn_isolation(adv),
})
turn("do", scenario.recall_text, filler_prose(999, scenario.prose_words))
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv)
settings = db.query(models.Settings).filter_by(user_id=user_id).first()
recall_action = (db.query(models.Action)
.filter(models.Action.adventure_id == adv,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
result["plant_depth"] = plant_depth
result["recall_depth"] = recall_action.depth
result["isolation"] = isolation(
db, adventure, FACT_F, plant_depth,
recall_snapshot=recall_action.context_snapshot,
recall_depth=recall_action.depth)
result["diagnosis"] = asyncio.run(diagnose(
db, adventure, settings, FACT_F, plant_depth,
recall_action=recall_action, embed=embedder.embed))
result["summariser_excerpts"] = len(summariser.excerpts)
# Provenance as the recall turn recorded it, resolved back to rows.
used = (recall_action.context_snapshot.get("memories") or {}).get("used") or []
f_entry = next((m for m in used
if m.get("id") == result["diagnosis"]["created"]["memory_id"]), None)
f_memory = db.get(models.Memory, f_entry["id"]) if f_entry else None
block = memorybank.source_block(db, f_memory) if f_memory is not None else []
result["provenance"] = {
"recorded": f_entry and {k: f_entry.get(k) for k in
("id", "source", "semantic_score", "lexical_score",
"final_score", "authority")},
"range_covers_plant": bool(f_entry and f_entry["source"]["source_start"]
<= plant_depth <= f_entry["source"]["source_end"]),
"matches_row": bool(f_memory is not None and f_entry["source"] == {
"branch_id": f_memory.branch_id, "depth": f_memory.depth,
"source_start": f_memory.source_start, "source_end": f_memory.source_end}),
"source_block_depths": [a.depth for a in block],
"source_block_holds_planting": any(a.text == FACT_F.sentence for a in block),
}
memory_id = result["diagnosis"]["created"]["memory_id"]
if memory_id is not None and not result["diagnosis"]["retained"]["forgotten"]:
variants = {}
base = production_query(adventure, recall_action.id)
# WP-B.2's negative control needs a decoy: another memory that
# holds a word the question adds, and nothing about F.
decoy = next(((m.id, found.group(1)) for m in db.execute(
select(models.Memory).where(models.Memory.adventure_id == adventure.id,
models.Memory.forgotten.is_(False),
models.Memory.id != memory_id)
.order_by(models.Memory.id)).scalars()
if (found := re.search(r"at the ([a-z]+ [a-z]+)\.", m.text or ""))), None)
queries = [
("direct", variant_query(base, scenario.recall_text)),
("paraphrase", variant_query(base, PARAPHRASE_QUERY)),
("unrelated", variant_query(base, UNRELATED_QUERY)),
# The player's words with the context taken away.
("input_only", {**variant_query(base, scenario.recall_text), "context": ""}),
# Words every memory in these fixtures holds, and nothing else.
("common_words", variant_query(base, COMMON_WORDS_QUERY)),
]
if decoy is not None:
queries.append(("rare_word_with_paraphrase", variant_query(
base, PARAPHRASE_QUERY[:-1] + f", out by the {decoy[1]}.")))
for label, query in queries:
ranking = asyncio.run(rank_bank(db, adventure, settings, query, embedder.embed))
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
decoy_row = next((r for r in ranking["scored"]
if decoy is not None and r["memory_id"] == decoy[0]), None)
text = query["input"]
variants[label] = {"query": text, "rank": row and row["rank"],
"selected_count": len(ranking["selected"]),
"decoy_memory_id": decoy and decoy[0],
"decoy_rank": decoy_row and decoy_row["rank"],
"decoy_lexical_score": decoy_row and decoy_row["lexical_score"],
"of": len(ranking["scored"]),
"similarity": row and row["similarity"],
"lexical_score": row and row["lexical_score"],
"final_score": row and row["final_score"],
"selected": bool(row and row["selected"]),
"top_k": ranking["top_k"]}
result["ranking_variants"] = variants
if memory_id is not None:
result["f_first_used_turn"] = next(
(t["turn"] for t in result["trace"] if (t["f_use_count"] or 0) > 0), None)
result["f_last_use_increase_turn"] = max(
(b["turn"] for a, b in zip(result["trace"], result["trace"][1:])
if (b["f_use_count"] or 0) > (a["f_use_count"] or 0)), default=None)
evicted_turn = next((t["turn"] for t in result["trace"] if t["f_forgotten"]), None)
first_evictions = next((t["evicted"] for t in result["trace"] if t["evicted"]), [])
result["eviction"] = {
"capacity": scenario.capacity,
"f_evicted_at_turn": evicted_turn,
"f_use_count_when_evicted": next(
(t["f_use_count"] for t in result["trace"] if t["f_forgotten"]), None),
"first_eviction_turn": next(
(t["turn"] for t in result["trace"] if t["evicted"]), None),
"first_evicted_ids": first_evictions,
"f_memory_was_first_evicted": bool(f_memory_id and f_memory_id in first_evictions),
"created_and_evicted_same_turn": sorted(
{i for t in result["trace"] for i in t["created_and_evicted_same_turn"]}),
"pinned_memory_id": pinned_id,
"pinned_memory_forgotten": bool(pinned_id and known[pinned_id]["forgotten"]),
}
if scenario.lineage_control:
path_clause = lineage.path_of(db, adventure).clause(models.Memory)
stored = db.execute(select(models.Memory.id).where(
models.Memory.id.in_(g.get("memory_ids") or [-1]))).scalars().all()
eligible = db.execute(select(models.Memory.id).where(
models.Memory.id.in_(g.get("memory_ids") or [-1]), path_clause)).scalars().all()
used_after = set()
injected_text = False
# Only turns played after the divergence. Before it, G was on the
# active line, and a memory of it being used then is correct.
for action in (db.query(models.Action)
.filter(models.Action.adventure_id == adv,
models.Action.type == "ai",
models.Action.id > g["last_action_id_before_divergence"])
.options(undefer(models.Action.context_snapshot))):
snap = action.context_snapshot or {}
for m in (snap.get("memories") or {}).get("used") or []:
if m.get("id") in (g.get("memory_ids") or []):
used_after.add(action.id)
if FACT_G.mentioned_by(_section(snap, MEMORIES_LABEL)):
injected_text = True
g.update(stored=stored, eligible_on_active_line=eligible,
turns_whose_memories_used_named_g=sorted(used_after),
g_text_ever_in_used_memories=injected_text)
if scenario.lineage_control:
call("POST", f"/checkpoints/{g['line_a']}/restore")
with SessionLocal() as db:
adventure = db.get(models.Adventure, adv)
eligible = db.execute(select(models.Memory.id).where(
models.Memory.id.in_(g.get("memory_ids") or [-1]),
lineage.path_of(db, adventure).clause(models.Memory))).scalars().all()
g["eligible_after_returning_to_line_a"] = eligible
result["lineage_control"] = g
return result
finally:
for obj, name, value in saved:
setattr(obj, name, value)
app.dependency_overrides.clear()
adventure_routes.turns._active_turns.clear()
memorybank._vector_cache.clear()
Base.metadata.drop_all(bind=engine)
+551
View File
@@ -0,0 +1,551 @@
"""v1.1 WP-B.2: how faithfully a real model's memories keep the facts of their block.
**Diagnostic only. Nothing here is imported by the application, and nothing here
is a release gate.**
Why it exists: the first isolation-valid real-model WP-B run failed at creation.
The whole planting block reached the summariser, and the memory it wrote left the
planted fact out (`V1.1-WP-B2-REPORT.md` §L.2). B2.4 tried the plan's bounded
remedy, a memory prompt instructing the model to keep named facts and objects. It
was measured with this module, did not correct the failure, and was **not
shipped** (§T). The module stays so the limitation can be measured again, on this
model or a different one.
- **Fixtures.** Short story blocks, each built around a fact a later scene could
turn on, with the ordinary texture a real block carries around it. They are
genre-neutral (office, contemporary, a science-fiction-neutral station), plus
the attempt-2 planting block itself. Each names the facts a memory must keep,
whom each belongs to, and what it must not invent.
- **The checker** (`evaluate`) is deterministic and reads only the memory text.
It is a heuristic, and says so:
- a fact counts as kept when one sentence names every part of it;
- attribution is the nearest named character before the fact's verb;
- it also reports word count, a leading "Memory:" and second-person "you".
- **The comparison** (`compare`) sends each fixture to a real model through the
application's own provider and `memorybank.memory_user_prompt`. It scores
memories under the shipped prompt and under the rejected B2.4 experiment.
A scripted summariser cannot show what a prompt makes a model do. So the
deterministic tests prove only that the checker is right and the fixtures reach
the summariser; model quality is measured here, with inference, and reported.
# the real-model measurement (inference: ask first)
.venv/bin/python -m tools.memory_fidelity --endpoint <v1 URL> \\
--model qwen2.5:3b-instruct-16k --samples 5 --out "$HOME/v11-evidence/<label>"
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
#: The memory prompt WP-B.2 B2.4 tried and **rejected**. Kept verbatim so the
#: experiment in `V1.1-WP-B2-REPORT.md` §T can be repeated; the application
#: never uses it.
B24_EXPERIMENT_PROMPT = (
"You compress interactive-fiction story excerpts into memories. Respond with "
"1-2 plain sentences in past tense, in at most 50 words.\n\n"
"Keep the concrete facts a later scene could turn on: specific people, "
"objects and places; where something is; who has, hid, found, knows, saw or "
"promised what; injuries, clues and commitments. A distinctive fact comes "
"before mood, scenery, routine movement and small talk. Drop those first, "
"however much of the excerpt they fill, and never let a later passage crowd "
"out an earlier fact.\n\n"
"Keep each fact with the person it belongs to. Never move an action, promise, "
"possession, statement or piece of knowledge from one character to another, "
"and never add a fact the excerpt does not state.\n\n"
'Write in the third person. The narration calls the protagonist "you". The '
'protagonist\'s own actions and words are the lines that begin with ">", '
'written as "I" or "You", and what they establish is part of the story just '
"as the narration is. The protagonist is named in the Cast: refer to them by "
'that name, never as "you" or "I". If the Cast gives no name for them, call '
'them "the player". Name the other characters too rather than writing "he", '
'"she" or "they" on their own — this memory will be read on its own, much '
"later, with nothing around it to say who a pronoun meant.\n\n"
"No preamble, no commentary."
)
TARGET_WORDS = 50
@dataclass(frozen=True)
class Fact:
"""One fact a memory must keep.
`groups`: every group must be matched in one sentence, by any of its terms.
`verbs`: the relation. When one is in that sentence, the nearest named
character before it is who the memory says the fact belongs to.
`actor`: whom it belongs to. `None` for a fact with no owner.
"""
name: str
groups: tuple[tuple[str, ...], ...]
actor: str | None = None
verbs: tuple[str, ...] = ()
@dataclass(frozen=True)
class Fixture:
fixture_id: str
genre: str
requirement: str # which measurement this fixture serves
protagonist: str
others: tuple[str, ...]
actions: tuple[tuple[str, str], ...] # (type, text), oldest first
facts: tuple[Fact, ...]
min_facts: int | None = None # default: all
#: Regexes a faithful memory must not match. Each is anchored on the wrong
#: character as the subject ("Marcus promised"), because a faithful memory
#: may name that character elsewhere in the same sentence ("promised Marcus").
forbidden: tuple[str, ...] = ()
#: Hand-written memories for the checker's own tests: one that should pass,
#: and failures that should not, each with the reason it must report.
faithful: str = ""
unfaithful: tuple[tuple[str, str], ...] = field(default_factory=tuple)
@property
def cast(self) -> tuple[str, ...]:
return (self.protagonist, *self.others)
@property
def raw(self) -> str:
return "\n\n".join(text for _, text in self.actions)
# ------------------------------------------------------------------ checker
def _term(term: str) -> re.Pattern:
body = r"\s+".join(re.escape(part) for part in term.lower().split())
return re.compile(rf"(?<![a-z]){body}(?:s|es|ed|d)?(?![a-z])")
def _sentences(text: str) -> list[str]:
return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text or "") if s.strip()]
def _first(low: str, terms) -> int | None:
found = [m.start() for t in terms for m in [_term(t).search(low)] if m]
return min(found) if found else None
def _attributed_to(sentence: str, fact: Fact, cast: tuple[str, ...]) -> str | None:
"""The character this sentence gives the fact to, or None if it names none."""
low = sentence.lower()
at = _first(low, fact.verbs) if fact.verbs else None
names = [(m.start(), name) for name in cast for m in _term(name).finditer(low)]
if at is not None:
before = [(pos, name) for pos, name in names if pos < at]
if before:
return max(before)[1]
return min(names)[1] if names else None
def evaluate(fixture: Fixture, memory: str) -> dict:
"""What a memory kept, whom it gave each fact to, what it invented, and how
it is framed."""
memory = memory or ""
sentences = _sentences(memory)
facts = {}
for fact in fixture.facts:
holding = [s for s in sentences
if all(_first(s.lower(), group) is not None for group in fact.groups)]
owners = sorted({o for s in holding
if (o := _attributed_to(s, fact, fixture.cast)) is not None})
facts[fact.name] = {
"kept": bool(holding),
"attributed_to": owners,
"attribution_ok": (fact.actor is None or not holding
or (owners != [] and set(owners) == {fact.actor})),
}
kept = sum(1 for f in facts.values() if f["kept"])
needed = len(fixture.facts) if fixture.min_facts is None else fixture.min_facts
inventions = [p for p in fixture.forbidden if re.search(p, memory, re.I)]
words = len(memory.split())
misattributed = [name for name, f in facts.items() if f["kept"] and not f["attribution_ok"]]
return {
"facts": facts,
"kept": kept,
"needed": needed,
"retained": kept >= needed,
"misattributed": misattributed,
"inventions": inventions,
"words": words,
"over_target": words > TARGET_WORDS,
# Framing the shipped prompt's rules exist to prevent.
"memory_prefix": memory.lstrip().lower().startswith("memory:"),
"second_person": re.search(r"\byou(r|rs|rself)?\b", memory, re.I) is not None,
"passed": kept >= needed and not misattributed and not inventions,
}
# ----------------------------------------------------------------- fixtures
OFFICE_TEXTURE = (
"The open-plan floor hums with keyboards and the air conditioning rattles in "
"its vent. Someone has left a birthday card on the printer, and the coffee "
"machine gurgles through another pot."
)
FIXTURES: tuple[Fixture, ...] = (
Fixture(
fixture_id="object_place_office",
genre="office",
requirement="distinctive object and place",
protagonist="Dana", others=("Priya",),
actions=(
("start", "Monday at the insurance office. Dana is covering the late shift."),
("do", "> You check the queue of unanswered claims."),
("ai", f"{OFFICE_TEXTURE} Priya walks past your desk carrying a stack of folders, "
"and slides the red backup drive into the bottom drawer of the grey filing "
"cabinet in the archive room before locking it. The phones ring twice and stop. "
"Rain streaks the tall windows while the floor slowly empties."),
("do", "> You ask Priya whether the audit is still on for Thursday."),
("ai", "Priya shrugs, says nobody tells her anything, and goes back to her own desk. "
"The cleaners arrive with their carts and the lights dim on a timer."),
("do", "> You log off and pack your bag."),
),
facts=(Fact("drive in the cabinet", (("backup drive", "drive"), ("drawer", "filing cabinet", "cabinet")),
actor="Priya", verbs=("slid", "slide", "put", "placed", "locked", "hid", "stored", "left")),),
forbidden=(r"\bdana\s+(had\s+)?(took|takes|has|holds|held|stole|locked|slid|hid)\b[^.]*\bdrive\b",),
faithful="Priya locked the red backup drive in the bottom drawer of the grey filing cabinet "
"in the archive room while Dana covered the late shift.",
unfaithful=(
("Dana covered a quiet late shift at the insurance office while rain fell and the "
"cleaners arrived.", "not retained"),
("Dana locked the red backup drive in the bottom drawer of the filing cabinet.",
"misattributed"),
),
),
Fixture(
fixture_id="player_fact_station",
genre="science-fiction-neutral",
requirement="player-established concrete fact",
protagonist="Reyes", others=("Okafor",),
actions=(
("start", "Deck four of the relay station, halfway through the night cycle. Reyes is on maintenance duty."),
("do", "> You walk the corridor checking the pressure seals."),
("ai", "The corridor lights pulse a dim blue. Condensation beads on the pipes, and "
"somewhere below a pump cycles on with a shudder. Chief Okafor passes with a "
"tablet under one arm and nods without stopping."),
("do", "> I watch Chief Okafor seal the coolant sample in locker nine and log it under a false name."),
("ai", "The night cycle drags on. The ventilation hisses, a door chimes somewhere "
"down the ring, and the viewport shows the same slow turn of stars it always "
"does. You finish the seal checks and sign the maintenance sheet, and the "
"corridor settles back into its usual hum."),
("do", "> You head back to your bunk."),
),
facts=(Fact("sample in locker nine", (("coolant sample", "sample"), ("locker",)),
actor="Okafor", verbs=("seal", "sealed", "locked", "put", "stored", "hid", "logged", "placed")),),
forbidden=(r"\breyes\s+(had\s+)?(sealed|seals|hid|stored|locked|logged)\b[^.]*\bsample\b",),
faithful="Reyes saw Chief Okafor seal the coolant sample in locker nine and log it under a false name.",
unfaithful=(
("Reyes finished the seal checks on deck four during a quiet night cycle.", "not retained"),
),
),
Fixture(
fixture_id="promise_contemporary",
genre="contemporary",
requirement="promise / commitment",
protagonist="Dana", others=("Marcus",),
actions=(
("start", "A Saturday afternoon at the flat Dana is about to rent from Marcus."),
("do", "> You look around the empty living room."),
("ai", "Sunlight falls across bare floorboards. The radiator ticks, a neighbour's "
"radio plays through the wall, and Marcus jingles a ring of keys while he "
"talks about the boiler and the bins."),
("do", '> You say "Marcus, I will bring you the signed lease by Friday."'),
("ai", "Marcus nods and writes something on the back of an envelope. Outside a bus "
"pulls away, a dog barks twice, and the afternoon light moves slowly up the wall."),
("do", "> You thank him and leave."),
),
facts=(Fact("lease by Friday", (("lease",), ("friday",)), actor="Dana",
verbs=("promise", "promised", "bring", "agreed", "said", "would")),),
forbidden=(r"\bmarcus\s+(promised|agreed|will\s+bring|would\s+bring)\b[^.]*\blease\b",),
faithful="Dana promised Marcus she would bring him the signed lease for the flat by Friday.",
unfaithful=(
("Marcus promised to bring Dana the signed lease by Friday.", "misattributed"),
("Dana viewed the empty flat on a sunny Saturday while Marcus talked about the boiler.",
"not retained"),
),
),
Fixture(
fixture_id="attribution_station",
genre="science-fiction-neutral",
requirement="attribution",
protagonist="Reyes", others=("Lena", "Tomas"),
actions=(
("start", "The survey ship's cargo bay, between jumps."),
("do", "> You ask who can open the sealed vault."),
("ai", "Lena folds her arms. She is the only one aboard who knows the vault door "
"code, and she makes it clear she is keeping it to herself. Tomas taps the "
"access badge clipped to his jacket; without it the bay lift will not move."),
("do", "> You look from one of them to the other."),
("ai", "The bay lights flicker as the drive spools. Crates creak against their "
"straps, and the air smells of cold metal and oil."),
("do", "> You wait for one of them to speak."),
),
facts=(
Fact("code", (("code",),), actor="Lena", verbs=("knows", "knew", "keeps", "kept", "holds", "held")),
Fact("badge", (("badge",),), actor="Tomas",
verbs=("carries", "carried", "has", "had", "holds", "held", "wore", "wears", "tapped", "taps")),
),
forbidden=(r"\breyes\b[^.]*\b(knew|knows)\b[^.]*\bcode\b",),
faithful="Lena alone knew the vault door code and kept it to herself; Tomas carried the access "
"badge that the bay lift needed.",
unfaithful=(
("Tomas knew the vault door code, and Lena carried the access badge.", "misattributed"),
),
),
Fixture(
fixture_id="clutter_office",
genre="office",
requirement="clutter pressure",
protagonist="Dana", others=("Priya", "Owen"),
actions=(
("start", "The quarterly offsite at a conference hotel by the motorway."),
("do", "> You find a seat near the back."),
("ai", "The conference room smells of carpet cleaner and burnt coffee. Chairs scrape, "
"a projector fan whines, and someone at the front struggles with the clicker. "
"Owen talks about his weekend at length, the traffic on the ring road, a new "
"sandwich place, the football, and whether it will rain for the barbecue. The "
"slides cycle through charts nobody reads. Outside the window lorries hiss past "
"on the wet motorway, and the hotel's muzak drifts in whenever the door opens."),
("do", "> You go to the refreshment table."),
("ai", "Pastries sweat under cling film. Priya stirs her tea, glances around, and "
"quietly tells you that she saw Owen shred the signed supplier contract in the "
"copy room last night. Then she talks about the weather, the parking, and the "
"long drive home, and laughs at a joke from across the room. The afternoon "
"session is announced, people drift back to their seats, and the projector "
"fan starts whining again over a long talk about quarterly targets."),
("do", "> You take your seat for the afternoon session."),
),
facts=(Fact("contract shredded", (("contract",), ("shred", "shredded", "destroyed")),
actor="Owen", verbs=("shred", "shredded", "destroyed")),),
forbidden=(r"\b(priya|dana)\b\s+(had\s+)?(shred|shredded|destroyed)\b",),
faithful="At the offsite, Priya told Dana she had seen Owen shred the signed supplier contract "
"in the copy room the night before.",
unfaithful=(
("Dana sat through a dull offsite of charts, pastries and Owen's talk about the weekend.",
"not retained"),
("Priya shredded the signed supplier contract in the copy room.", "misattributed"),
),
),
Fixture(
fixture_id="no_invention_office",
genre="office",
requirement="no invention",
protagonist="Dana", others=("Owen",),
actions=(
("start", "A short planning meeting in the small room on the third floor."),
("do", "> You sit down opposite Owen."),
("ai", "A black briefcase sits unclaimed by the door; nobody mentions it. Owen says "
"the budget review has moved from Tuesday to Thursday, and asks you to tell "
"the team."),
("do", "> You agree to pass it on."),
("ai", "Owen thanks you, checks his phone, and the meeting ends after ten minutes. "
"The briefcase is still by the door when you leave."),
("do", "> You walk back to your desk."),
),
facts=(Fact("review moved", (("budget review", "review"), ("thursday",))),),
forbidden=(
r"\b(took|taken|stole|hid|hidden|grabbed|pocketed|carried|carries|owns|owned|belong\w*)\b[^.]*\bbriefcase\b",
r"\bbriefcase\b[^.]*\b(belong\w*|his|her|owen's|dana's|secret|clue)\b",
r"\b(clue|secret|password|code)\b",
),
faithful="Owen told Dana the budget review had moved from Tuesday to Thursday, and Dana agreed "
"to tell the team.",
unfaithful=(
("Owen told Dana the budget review had moved to Thursday and left his secret briefcase by the door.",
"invented"),
),
),
Fixture(
fixture_id="multiple_facts_station",
genre="science-fiction-neutral",
requirement="multiple concrete facts",
protagonist="Reyes", others=("Hale", "Varga", "Moreau"),
actions=(
("start", "The mess hall of the mining outpost after the shift change."),
("do", "> You sit with the day crew."),
("ai", "Trays clatter and the recycler drones. Engineer Hale admits, half joking, that "
"she hid the spare fuse inside the airlock control panel. Doctor Varga says only "
"she knows the reactor override phrase, and changes the subject. Pilot Moreau "
"grumbles that he owes Hale two shifts of cover."),
("do", "> You finish your meal."),
("ai", "The lights dim for the rest cycle and the crew drifts off to their bunks."),
("do", "> You head to your quarters."),
),
facts=(
Fact("fuse in the panel", (("fuse",), ("panel", "airlock")), actor="Hale",
verbs=("hid", "hidden", "put", "stashed", "kept")),
Fact("override phrase", (("override",),), actor="Varga", verbs=("knows", "knew", "keeps")),
Fact("owes cover", (("owes", "owed", "debt"), ("shift", "cover")), actor="Moreau",
verbs=("owes", "owed")),
),
min_facts=2,
faithful="Hale hid the spare fuse inside the airlock control panel, and Varga alone knew the "
"reactor override phrase.",
unfaithful=(
("Varga hid the spare fuse inside the airlock panel, and Hale knew the reactor override phrase.",
"misattributed"),
),
),
Fixture(
# The planting block of WP-B.2 real-model attempt 2, verbatim: the
# harness's own fixture campaign, with no identifier in it. It is the
# failure the reference model has not overcome.
fixture_id="regression_attempt_2",
genre="fantasy (the harness campaign)",
requirement="the actual failed-run block",
protagonist="Aldric", others=("Mara", "Edrin"),
actions=(
("start", "Rain over Westhaven. Aldric sits in the Crooked Lantern with a silver key in his "
"pocket and no-one to give it to."),
("do", "> I tell Mara quietly that the silver key opens the crypt beneath the Old Abbey — "
"SILVER-KEY-CRYPT-OLD-ABBEY."),
("ai", "The lantern's warmth feels cold against Aldric's hands as he fidgets with the silver "
"key. Mara observes him without a word, her eyes unreadable. The rain outside "
"persists, a backdrop of nature's indifference to the human drama transpiring. "
"Aldric's fingers close over the key, the familiar weight making him feel both more "
"alone and more determined. He pulls the key from his pocket, the metal feeling cold "
"against his skin. \n\n> \"I have the key,\" Aldric whispers to Mara. \"It leads to "
"the crypt beneath the abbey.\"\n\nEdrin, having finished his ale, stands. \"Well, now "
"that we know the key leads to somewhere... do you think we should see what it "
"opens?\"\n\nThe silver key feels heavy in Aldric's hands. He looks at Mara, then at "
"the key, and finally at Edrin. He decides. \"Let's do this.\""),
("do", "> I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top "
"shelf, and she makes me promise to tell no one."),
("ai", "Aldric steps closer to the smoky bar, his gaze shifting from the silver key in his "
"pocket to Mara. \"Let's not overthink this,\" he says, his voice steady despite the "
"gnawing doubt that tugs at him. \"Let's see what the crypt holds.\"\n\nEdrin raises "
"an eyebrow, his curiosity piqued. \"Alright, then. But if this leads somewhere "
"dangerous, we're splitting up. If we make it back here, I'll tell no one what we "
"found.\"\n\nMara nods her agreement, her eyes still unreadable. \"Agreed. But "
"remember, we're not the first to look there. We need to tread carefully.\"\n\nWith a "
"slight nod, they step out of the tavern. The rain has only intensified, pelting them "
"as they walk toward the Old Abbey's shadowed grounds. The crypt, sealed since its "
"founding, stands before them, its door ajar, inviting or perhaps foreboding. As "
"Aldric approaches, he feels a strange mixture of excitement and fear, his hands "
"curling around the silver key's handle. \n\n> He turns the key, and the lock clicks "
"open with a satisfying, if ominous, hiss. The door swings inward, revealing a narrow "
"passageway. The air inside is musty and cold, the only light filtering in through "
"the cracks. \n\nAldric takes a tentative step inside, his breath visible in the air. "
"Mara and Edrin follow, their footsteps echoing in the silence. The walls are lined "
"with stone and cobwebs, and the smell of damp and decay is overwhelming. They "
"advance cautiously, each of them aware of the potential dangers lurking within."),
("do", "> I ask Mara what she has heard about the abbey."),
),
facts=(Fact("sundial in the teapot", (("sundial",), ("teapot",)), actor="Mara",
verbs=("slip", "slipped", "hid", "hidden", "put", "placed", "tucked", "stashed")),),
forbidden=(
r"\b(aldric|edrin)\b\s+(had\s+)?(slipped|hid|placed|put|tucked|stashed)\b[^.]*\bsundial\b",
),
faithful="Aldric watched Mara slip the amber sundial inside the cracked teapot on the tavern's "
"top shelf and promised her to tell no one; then Aldric, Mara and Edrin entered the "
"crypt beneath the Old Abbey.",
unfaithful=(
# The memory attempt 2 actually stored, verbatim.
("Aldric sits in the Crooked Lantern with a silver key in his pocket, no-one to give it to. "
"Mara observed him quietly, her eyes unreadable. Edrin finished his ale and stood, asking if "
"they should see what the crypt beneath the Old Abbey holds. Aldric decided to go, promising "
"not to tell anyone. They walked to the Old Abbey's grounds, the crypt door ajar, inviting "
"and foreboding. Inside, the air was musty and cold, with the smell of damp and decay. They "
"advanced cautiously, each aware of potential dangers. The silver key, the key to the crypt, "
"felt heavy in Aldric's hands.", "not retained"),
("Aldric slipped the amber sundial inside the cracked teapot on the tavern's top shelf.",
"misattributed"),
),
),
)
FIXTURES_BY_ID = {f.fixture_id: f for f in FIXTURES}
def cast_brief_for(fixture: Fixture) -> str:
"""The cast brief `memorybank.cast_brief` would build for this fixture."""
from app import memorybank
lines = [memorybank._cast_line(fixture.protagonist, "", protagonist=True)]
lines += [memorybank._cast_line(name, "") for name in fixture.others]
return "Cast:\n" + "\n".join(lines)
def user_prompt_for(fixture: Fixture) -> str:
"""Exactly the user message the application sends for this block."""
from app import memorybank
return memorybank.memory_user_prompt(cast_brief_for(fixture), memorybank.memory_excerpt(fixture.raw))
# --------------------------------------------------------- the measurement
async def compare(endpoint: str, model: str, samples: int) -> dict:
"""Every fixture, `samples` times, under the shipped prompt and the B2.4 experiment."""
from app import memorybank, models
settings = models.Settings(endpoint_url=endpoint, model=model, summary_model="",
api_mode=models.Settings.__table__.c.api_mode.default.arg,
model_timeout_seconds=300)
provider = memorybank.summary_provider(settings)
arms = {"shipped": memorybank.MEMORY_SYSTEM_PROMPT, "b2.4-experiment": B24_EXPERIMENT_PROMPT}
out: dict = {"endpoint_model": model, "samples": samples, "fixtures": {}}
for fixture in FIXTURES:
user = user_prompt_for(fixture)
row: dict = {"requirement": fixture.requirement, "genre": fixture.genre, "arms": {}}
for arm, system in arms.items():
runs = []
for _ in range(samples):
text = (await provider.complete(system, user) or "").strip()
runs.append({"memory": text, **evaluate(fixture, text)})
summary = {
"passed": sum(r["passed"] for r in runs),
"retained": sum(r["retained"] for r in runs),
"misattributed": sum(bool(r["misattributed"]) for r in runs),
"invented": sum(bool(r["inventions"]) for r in runs),
"memory_prefix": sum(r["memory_prefix"] for r in runs),
"second_person": sum(r["second_person"] for r in runs),
"words_median": sorted(r["words"] for r in runs)[len(runs) // 2],
"words_max": max(r["words"] for r in runs),
"over_target": sum(r["over_target"] for r in runs),
}
row["arms"][arm] = {"runs": runs, **summary}
print(f"{fixture.fixture_id:28} {arm:16} "
+ " ".join(f"{k} {v}" for k, v in summary.items()), flush=True)
out["fixtures"][fixture.fixture_id] = row
return out
def main() -> int:
import argparse
import asyncio
import json
import os
import sys
import tempfile
from pathlib import Path
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--endpoint", required=True)
parser.add_argument("--model", required=True)
parser.add_argument("--samples", type=int, default=5)
parser.add_argument("--out", required=True)
args = parser.parse_args()
handle = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
handle.close()
os.environ["AIDND_DB_PATH"] = handle.name # nothing is written; never the real database
os.environ.pop("AIDND_DATABASE_URL", None)
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
try:
result = asyncio.run(compare(args.endpoint, args.model, args.samples))
finally:
Path(handle.name).unlink(missing_ok=True)
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
(out / "fidelity.json").write_text(json.dumps(result, indent=2))
print(f"written to {out / 'fidelity.json'}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+125
View File
@@ -0,0 +1,125 @@
"""v1.1 WP-B.1: run the deterministic memory-retention scenarios, or diagnose a real campaign.
# the deterministic scenarios, against an isolated database in --out
.venv/bin/python -m tools.v11_b1_memory scenarios --out "$HOME/v11-evidence/b1/<label>"
# the four stages for a finished real campaign (reads its database; embeds
# the recall query with the campaign's own configured embedding model)
AIDND_TEST_ENDPOINT=... AIDND_TEST_EMBED_MODEL=nomic-embed-text:latest \\
.venv/bin/python -m tools.v11_b1_memory diagnose --db <campaign.db> \\
--plant-depth 3 --out "$HOME/v11-evidence/b1/<label>"
Run from `backend/`. Nothing here changes memory behaviour; see
`tools/memory_diagnostic.py` for what is measured and what the stubs model.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sys
from pathlib import Path
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
sub = parser.add_subparsers(dest="command", required=True)
scen = sub.add_parser("scenarios")
scen.add_argument("--out", required=True)
scen.add_argument("--only", action="append", default=[])
diag = sub.add_parser("diagnose")
diag.add_argument("--db", required=True)
diag.add_argument("--plant-depth", type=int, required=True)
diag.add_argument("--out", required=True)
args = parser.parse_args()
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
if args.command == "scenarios":
db_path = out / "scenarios.db"
if db_path.exists():
db_path.unlink()
os.environ["AIDND_DB_PATH"] = str(db_path)
else:
# A copy, so diagnosis never writes to the evidence database.
copy = out / "diagnosed-copy.db"
shutil.copy2(args.db, copy)
os.environ["AIDND_DB_PATH"] = str(copy)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from tools import memory_diagnostic as md # after the database is chosen
if args.command == "scenarios":
names = args.only or list(md.SCENARIOS)
summary = {}
for name in names:
result = md.run_scenario(md.SCENARIOS[name])
(out / f"{name}.json").write_text(json.dumps(result, indent=2, default=str))
d = result.get("diagnosis") or {}
summary[name] = {
"verdict": d.get("verdict"),
"isolation_ok": (result.get("isolation") or {}).get("ok"),
"plant_depth": result.get("plant_depth"),
"recall_depth": result.get("recall_depth"),
"f_evicted_at_turn": (result.get("eviction") or {}).get("f_evicted_at_turn"),
}
print(f"{name:26} verdict={d.get('verdict')!s:26} "
f"isolation_ok={summary[name]['isolation_ok']} "
f"plant={result.get('plant_depth')} recall={result.get('recall_depth')}")
(out / "summary.json").write_text(json.dumps(summary, indent=2))
return 0
import asyncio
from sqlalchemy.orm import undefer
from app import memorybank, models
from app.database import SessionLocal
endpoint = os.environ.get("AIDND_TEST_ENDPOINT", "")
embed_model = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
with SessionLocal() as db:
adventure = db.query(models.Adventure).order_by(models.Adventure.id).first()
settings = db.query(models.Settings).filter_by(user_id=adventure.user_id).first()
if endpoint:
settings.endpoint_url = endpoint
if embed_model:
settings.embedding_model = embed_model
recall_action = (db.query(models.Action)
.filter(models.Action.adventure_id == adventure.id,
models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first())
embed = memorybank.embedding_provider(settings).embed
iso = md.isolation(db, adventure, md.FACT_F, args.plant_depth,
recall_snapshot=recall_action.context_snapshot,
recall_depth=recall_action.depth)
diagnosis = asyncio.run(md.diagnose(db, adventure, settings, md.FACT_F, args.plant_depth,
recall_action=recall_action, embed=embed))
variants = {}
memory_id = diagnosis["created"]["memory_id"]
if memory_id is not None and not diagnosis.get("retained", {}).get("forgotten"):
base = md.production_query(adventure, recall_action.id)
for label, query in (("recall_turn", base),
("paraphrase", md.variant_query(base, md.PARAPHRASE_QUERY)),
("unrelated", md.variant_query(base, md.UNRELATED_QUERY))):
ranking = asyncio.run(md.rank_bank(db, adventure, settings, query, embed))
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
variants[label] = {"rank": row and row["rank"], "of": len(ranking["scored"]),
"similarity": row and row["similarity"],
"lexical_score": row and row["lexical_score"],
"final_score": row and row["final_score"],
"selected": bool(row and row["selected"])}
db.rollback()
report = {"isolation": iso, "diagnosis": diagnosis, "ranking_variants": variants}
(out / "diagnosis.json").write_text(json.dumps(report, indent=2, default=str))
print(json.dumps({"isolation_ok": iso["ok"], "verdict": diagnosis["verdict"]}, indent=2))
return 0
if __name__ == "__main__":
sys.exit(main())
+187
View File
@@ -0,0 +1,187 @@
"""v1.1: does a real v1.0.0 database open unchanged?
# the same database, opened by each tree, snapshotted read-only
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
--tree <v1.0.0 worktree>/backend --label v100 --out "$HOME/v11-evidence/compat"
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
--label v11 --exercise --out "$HOME/v11-evidence/compat"
Run from `backend/`. The source database is never opened. It is copied into
`--out` first, and the copy is what the application opens.
A read-only snapshot is taken through the API, the same way a reader sees the
campaign:
- the export bundle, which carries the whole tree, the head, the Save Points,
state, events, summaries, memories and knowledge, and has no timestamp of its
own;
- the narrative state and its events;
- the Save Points, the imported knowledge, the memories, the derived status and
the settings;
- the database schema and `PRAGMA user_version`, before and after the
application opened it.
Two snapshots of the same database from two trees are then compared. Identical
means v1.1 read it exactly as v1.0.0 did, and a matching schema and version mean
nothing migrated.
`--exercise` then uses the v1.1 copy: undo, redo, a Save Point restore, a
context dry run (knowledge retrieval), an export, and an import of that export.
It first points the copy's endpoint at a loopback port that refuses, so nothing
here reaches an inference server.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sqlite3
import sys
from pathlib import Path
def _schema(path: Path) -> dict:
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
version = connection.execute("PRAGMA user_version").fetchone()[0]
rows = connection.execute(
"SELECT type, name, sql FROM sqlite_master WHERE name NOT LIKE 'sqlite_%' "
"ORDER BY type, name").fetchall()
finally:
connection.close()
return {"user_version": version, "objects": [list(r) for r in rows]}
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--db", required=True)
parser.add_argument("--tree", default="", help="a backend/ directory to import the app from")
parser.add_argument("--label", required=True)
parser.add_argument("--exercise", action="store_true")
parser.add_argument("--out", required=True)
args = parser.parse_args()
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
copy = out / f"{args.label}.db"
if copy.exists():
print(f"{copy} exists; choose a new --label or --out")
return 2
shutil.copy2(args.db, copy)
schema_before = _schema(copy)
if args.tree:
sys.path.insert(0, str(Path(args.tree).resolve()))
os.environ["AIDND_DB_PATH"] = str(copy)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from fastapi import Depends
from fastapi.testclient import TestClient
from app import auth, limits, models
from app.database import SessionLocal, get_db
from app.main import app
print(f"app imported from {Path(sys.modules['app'].__file__).parent}")
limits.check_row_cap = lambda *a, **k: None
with SessionLocal() as db:
owner = db.query(models.Adventure.user_id).order_by(models.Adventure.id).first()
user_id = owner[0] if owner else db.query(models.User.id).first()[0]
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
def call(client, method, url, body=None, expect=200):
response = client.request(method, f"/api{url}", json=body)
if response.status_code != expect:
raise SystemExit(f"{method} {url}: HTTP {response.status_code} {response.text[:300]}")
return response.json() if response.content else None
report: dict = {"label": args.label, "schema_before": schema_before}
with TestClient(app) as client:
with SessionLocal() as db:
adventure_ids = [a for (a,) in db.query(models.Adventure.id)
.filter(models.Adventure.user_id == user_id)
.order_by(models.Adventure.id)]
snapshot = {"settings": call(client, "GET", "/settings"), "adventures": {}}
for adv in adventure_ids:
snapshot["adventures"][str(adv)] = {
"export": call(client, "GET", f"/adventures/{adv}/export"),
"state": call(client, "GET", f"/adventures/{adv}/state"),
"state_events": call(client, "GET", f"/adventures/{adv}/state/events"),
"checkpoints": call(client, "GET", f"/adventures/{adv}/checkpoints"),
"knowledge": call(client, "GET", f"/adventures/{adv}/knowledge"),
"memories": call(client, "GET", f"/adventures/{adv}/memories"),
"derived": call(client, "GET", f"/adventures/{adv}/derived"),
"newest_actions": call(client, "GET", f"/adventures/{adv}/actions?limit=5"),
}
report["snapshot"] = snapshot
if args.exercise and adventure_ids:
adv = adventure_ids[0]
ex: dict = {}
call(client, "PUT", "/settings", {"endpoint_url": "http://127.0.0.1:9/v1",
"embedding_model": ""})
before = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
ex["before"] = {k: before[k] for k in ("total", "can_undo", "can_redo")}
undone = call(client, "POST", f"/adventures/{adv}/undo")
ex["after_undo"] = {k: undone[k] for k in ("total", "can_undo", "can_redo")}
redone = call(client, "POST", f"/adventures/{adv}/redo")
ex["after_redo"] = {k: redone[k] for k in ("total", "can_undo", "can_redo")}
points = call(client, "GET", f"/adventures/{adv}/checkpoints")
if points:
point = points[0]
restored = call(client, "POST",
f"/adventures/{adv}/checkpoints/{point['id']}/restore")
page = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
ex["restore"] = {"save_point": point["name"], "total": page["total"],
"can_redo": page["can_redo"],
"response_keys": sorted(restored or {})}
context = client.get(f"/api/adventures/{adv}/context")
body = context.json()
ex["context"] = {
"status": context.status_code,
"knowledge_used": len(((body.get("knowledge") or {}).get("used")) or []),
"canon_section": any(s["label"] == "campaign_canon"
for s in body.get("sections") or []),
"state_section": any(s["label"] == "narrative_state"
for s in body.get("sections") or []),
"summary": body.get("summary"),
"tokens": body.get("tokens"),
"window": body.get("window"),
}
bundle = call(client, "GET", f"/adventures/{adv}/export")
imported = call(client, "POST", "/adventures/import", bundle, expect=201)
new_id = imported["id"]
reimport = call(client, "GET", f"/adventures/{new_id}/export")
ex["import"] = {
"new_id": new_id,
"actions_in_bundle": len(bundle.get("actions") or []),
"actions_after_import": len(reimport.get("actions") or []),
"head_same": (bundle.get("headBranch") is not None
and bundle.get("headDepth") == reimport.get("headDepth")),
"checkpoints": [len(bundle.get("checkpoints") or []),
len(reimport.get("checkpoints") or [])],
"memories": [len(bundle.get("memories") or []),
len(reimport.get("memories") or [])],
"narrative_state_same": bundle.get("narrativeState") == reimport.get("narrativeState"),
}
report["exercise"] = ex
app.dependency_overrides.clear()
report["schema_after"] = _schema(copy)
(out / f"{args.label}.json").write_text(json.dumps(report, indent=2, sort_keys=True, default=str))
same_schema = report["schema_before"] == report["schema_after"]
print(f"schema unchanged by opening: {same_schema} "
f"(user_version {report['schema_before']['user_version']} -> "
f"{report['schema_after']['user_version']})")
if "exercise" in report:
print(json.dumps(report["exercise"], indent=2, default=str)[:3000])
return 0
if __name__ == "__main__":
sys.exit(main())
+307
View File
@@ -0,0 +1,307 @@
"""v1.1 release smoke test: the shipped image, as a reader would meet it.
python -m tools.v11_release_smoke --image <tag> --out <dir under $HOME>
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT` (an **HTTPS** Ollama on the
trusted LAN) and `AIDND_TEST_MODEL`. `--ca` names the private CA to install
inside the container, defaulting to this machine's own.
Supplemental release evidence, not a replacement for the gates: it asks whether
the artefact that ships actually runs, reaches its approved narrator, refuses an
unapproved one, and keeps a campaign across a container restart.
## The two things this is careful about
**The CA is installed, not bypassed.** `app/tlstrust.ssl_context()` is
`ssl.create_default_context()` — the platform's own store — unioned with
certifi's. So the private CA is mounted into
`/usr/local/share/ca-certificates/` and registered with
`update-ca-certificates`, and verification is then ordinary. Nothing sets
`verify=False`, and a check inside the container proves the handshake succeeds
through that store.
**Loopback means the published port.** The process inside the container listens
on `0.0.0.0` because that is the only address a published port can reach
(`docker-compose.yml` says so). What must be loopback-only is the *publish*, so
the container is started with `-p 127.0.0.1:<port>:8000` and the check is that
the host's LAN address refuses the same port.
"""
from __future__ import annotations
import argparse
import json
import os
import socket
import subprocess
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import Browser, free_port, require_under_home # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
NAME = "v11-release-smoke"
VOLUME = "v11-release-smoke-data"
#: An endpoint the policy must refuse whatever else is true: a public host.
PUBLIC_ENDPOINT = "https://api.openai.com/v1"
class Checks:
def __init__(self) -> None:
self.rows: list[dict] = []
def record(self, name: str, ok: bool, detail: str = "") -> bool:
self.rows.append({"check": name, "result": "PASS" if ok else "FAIL",
"detail": detail})
print(f" {'ok ' if ok else 'FAIL'} {name}" + (f" — {detail}" if detail else ""),
flush=True)
return ok
@property
def failed(self) -> list[dict]:
return [r for r in self.rows if r["result"] == "FAIL"]
def run(*args: str, **kwargs) -> subprocess.CompletedProcess:
return subprocess.run(args, capture_output=True, text=True, **kwargs)
def api(base: str, method: str, path: str, payload=None, timeout=900):
data = json.dumps(payload).encode() if payload is not None else None
request = urllib.request.Request(
f"{base}/api{path}", data=data, method=method,
headers={"Content-Type": "application/json"} if data else {})
with urllib.request.urlopen(request, timeout=timeout) as response:
body = response.read().decode()
return json.loads(body) if body else None
def stream_turn(base: str, adv: int, text: str) -> list[dict]:
request = urllib.request.Request(
f"{base}/api/adventures/{adv}/actions",
data=json.dumps({"type": "do", "text": text}).encode(),
method="POST", headers={"Content-Type": "application/json"})
events: list[dict] = []
with urllib.request.urlopen(request, timeout=900) as response:
for raw in response:
line = raw.decode(errors="replace").strip()
if line.startswith("data:"):
try:
events.append(json.loads(line[5:].strip()))
except json.JSONDecodeError:
pass
return events
def lan_address() -> str | None:
"""This machine's own LAN address, for the loopback-only check."""
probe = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
try:
probe.connect(("192.0.2.1", 9)) # TEST-NET-1: routed nowhere, sends nothing
return probe.getsockname()[0]
except OSError:
return None
finally:
probe.close()
def wait_ready(base: str, *, timeout: float = 180) -> bool:
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
urllib.request.urlopen(f"{base}/api/settings", timeout=3)
return True
except Exception:
time.sleep(1)
return False
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--image", required=True)
parser.add_argument("--out", required=True)
parser.add_argument("--ca", default="/usr/local/share/ca-certificates/draco.crt")
parser.add_argument(
"--add-host", default="", metavar="NAME:ADDRESS",
help=("resolve the narrator's hostname inside the container. A `.local` "
"name is mDNS, and a container has no mDNS resolver, so the "
"endpoint policy refuses an address it cannot classify and "
"`PUT /api/settings` answers 400. Mapping the name — rather than "
"using the address — keeps the hostname the certificate is issued "
"for, which is the thing this test verifies."))
args = parser.parse_args()
if not (ENDPOINT and MODEL):
print("set AIDND_TEST_ENDPOINT (https://…) and AIDND_TEST_MODEL")
return 2
if not ENDPOINT.startswith("https://"):
print("the smoke test needs an HTTPS endpoint: that is what it verifies")
return 2
ca = Path(args.ca)
if not ca.exists():
print(f"no CA at {ca}")
return 2
out = require_under_home(Path(args.out).expanduser())
out.mkdir(parents=True, exist_ok=True)
checks = Checks()
port = free_port()
base = f"http://127.0.0.1:{port}"
started = datetime.now()
run("docker", "rm", "-f", NAME)
run("docker", "volume", "rm", VOLUME)
run("docker", "volume", "create", VOLUME)
print(f"starting {args.image} on 127.0.0.1:{port} with a fresh volume …")
start = run(
"docker", "run", "-d", "--name", NAME,
"-p", f"127.0.0.1:{port}:8000",
"-v", f"{VOLUME}:/data",
"-v", f"{ca}:/usr/local/share/ca-certificates/{ca.name}:ro",
*(("--add-host", args.add_host) if args.add_host else ()),
args.image,
"sh", "-c",
"update-ca-certificates >/dev/null 2>&1; "
"exec uvicorn app.main:app --host 0.0.0.0 --port 8000",
)
if start.returncode != 0:
print(start.stderr[:400])
return 1
container = start.stdout.strip()[:12]
try:
checks.record("the container starts", True, container)
ready = wait_ready(base)
if not checks.record("the application answers on loopback", ready, base):
logs = run("docker", "logs", NAME)
(out / "container.log").write_text(logs.stdout + logs.stderr)
return 1
published = run("docker", "port", NAME).stdout.strip()
checks.record("the port is published on loopback only",
"127.0.0.1" in published and "0.0.0.0" not in published, published)
lan = lan_address()
if lan:
try:
urllib.request.urlopen(f"http://{lan}:{port}/api/settings", timeout=4)
reachable = True
except Exception:
reachable = False
checks.record("the LAN address does not serve the application", not reachable,
f"port {port} on this machine's LAN address")
page = urllib.request.urlopen(base + "/", timeout=30)
html = page.read().decode(errors="replace")
checks.record("the first page loads", page.status == 200 and "<div id=\"root\"" in html,
f"HTTP {page.status}, {len(html)} bytes")
remote = [chunk for chunk in html.split('"')
if chunk.startswith("http://") or chunk.startswith("https://")]
checks.record("the shell references no remote origin", not remote, str(remote[:3]))
csp = page.headers.get("content-security-policy") or ""
checks.record("a CSP is served", bool(csp), csp[:80])
# The approved endpoint, verified through the private CA *inside* the
# container, with the application's own trust context and no bypass.
probe = run("docker", "exec", NAME, "python", "-c",
"import json,urllib.request,ssl,sys;"
"sys.path.insert(0,'/app/backend');"
"from app.tlstrust import ssl_context;"
f"r=urllib.request.urlopen('{ENDPOINT}/models',"
" timeout=20, context=ssl_context());"
"print(r.status)")
checks.record("the approved HTTPS narrator verifies through the private CA",
probe.returncode == 0 and "200" in probe.stdout,
(probe.stdout + probe.stderr).strip()[:160])
api(base, "PUT", "/settings", {"endpoint_url": ENDPOINT, "model": MODEL,
"max_output_tokens": 300,
"model_timeout_seconds": 600})
try:
api(base, "PUT", "/settings", {"endpoint_url": PUBLIC_ENDPOINT, "model": MODEL})
refused = False
detail = "accepted"
except urllib.error.HTTPError as exc:
refused = 400 <= exc.code < 500
detail = f"HTTP {exc.code}"
checks.record("a public endpoint is refused", refused, detail)
# Put the approved one back, whatever happened above.
api(base, "PUT", "/settings", {"endpoint_url": ENDPOINT, "model": MODEL,
"max_output_tokens": 300,
"model_timeout_seconds": 600})
created = api(base, "POST", "/adventures", {
"title": "Release Smoke",
"opening": "Rain over Westhaven, and the abbey bell tolling.",
"persona_name": "Aldric"})
adv = created["id"]
checks.record("a campaign is created", bool(adv), f"id {adv}")
events = stream_turn(base, adv, "I ask Mara what the bell means.")
errors = [e for e in events if e.get("type") == "error"]
page_after = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {}
checks.record("one real narrator turn is accepted",
not errors and (page_after.get("total") or 0) >= 2,
errors[0].get("detail", "")[:160] if errors else
f"{page_after.get('total')} actions")
before = [(a.get("type"), (a.get("text") or "")[:120])
for a in (page_after.get("actions") or [])]
state_before = api(base, "GET", f"/adventures/{adv}/state") or {}
print("restarting the container …")
run("docker", "restart", NAME)
ready = wait_ready(base)
checks.record("the container restarts and serves again", ready)
page_reopened = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {}
after = [(a.get("type"), (a.get("text") or "")[:120])
for a in (page_reopened.get("actions") or [])]
checks.record("the transcript survived the restart", after == before,
f"{len(before)} -> {len(after)} actions")
state_after = api(base, "GET", f"/adventures/{adv}/state") or {}
checks.record("the narrative state survived the restart",
state_after == state_before)
browser = Browser(headless=True, log=out / "geckodriver.log")
try:
browser.go(f"{base}/play/{adv}")
browser.wait_for(".story-controls", timeout=60)
story = browser.js(
"const el = document.querySelector('.story');"
" return el ? el.textContent.trim().length : 0;")
checks.record("Firefox renders the reopened campaign",
isinstance(story, int) and story > 0, f"{story} characters of story")
browser.screenshot(out / "reopened-campaign.png")
finally:
browser.quit()
finally:
logs = run("docker", "logs", NAME)
(out / "container.log").write_text(logs.stdout + logs.stderr)
run("docker", "rm", "-f", NAME)
run("docker", "volume", "rm", VOLUME)
report = {
"image": args.image,
"started": started.isoformat(timespec="seconds"),
"seconds": round((datetime.now() - started).total_seconds()),
"endpoint_class": "trusted-LAN HTTPS with a private CA",
"checks": checks.rows,
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
"failed": len(checks.failed),
}
(out / "smoke-report.json").write_text(json.dumps(report, indent=2))
print(f"\n{report['passed']} passed, {report['failed']} failed "
f"-> {out / 'smoke-report.json'}")
return 1 if checks.failed else 0
if __name__ == "__main__":
raise SystemExit(main())
+245
View File
@@ -0,0 +1,245 @@
"""v1.1 WP-A2: replay real stored narration through the v1.0.0 and current extractors.
The A2 extractor changes remove more text from a narrator's reply than v1.0.0
did. Removing story is worse than leaving protocol (`TECHNICAL-DESIGN.md`
§15.4), so every change is shown to a person rather than summarised. This tool
takes every real reply the evidence kept, runs it through both extractors, and
writes each turn whose prose differs with:
- the v1.0.0 prose and the current prose;
- every line removed, and the rule that explains it;
- any removal no rule explains, which fails the replay.
**Input.** A reply is read from the turn's stored `raw_output` where the
evidence database kept one: that is exactly what the narrator sent. A bundle
carries no raw reply, so a bundle's turns are replayed from their stored text,
which is v1.0.0's output already. For those the old prose is the input itself,
and the comparison is still exact.
**The v1.0.0 extractor** is read from the release tag with `git show`, not
copied, so this tool compares against what shipped. It shares `events` and
`render` with the current tree. A2 does not change `events.SPECS` or
`render.SECTION_HEADINGS`, and the report verifies that with `git diff`.
Evidence stays outside the repository. Replayed text is fiction from the
acceptance fixtures, but it is still somebody's run.
.venv/bin/python -m tools.v11_replay_extractor \\
--db "$HOME/m11-evidence/**/*.db" \\
--bundle "$HOME/m11-evidence/closeout-3652dc6/identity/turn-99/bundle.json" \\
--out "$HOME/v11-evidence/a2-replay"
"""
from __future__ import annotations
import argparse
import difflib
import glob
import hashlib
import importlib.util
import json
import sqlite3
import subprocess
import sys
import types
import zlib
from pathlib import Path
RELEASE = "v1.0.0"
EXTRACT_PATH = "backend/app/narrative/extract.py"
#: A removal larger than this share of the v1.0.0 prose is flagged for review
#: even when every line is explained, because a rule that eats most of a reply
#: is the shape a false positive takes.
LARGE_REMOVAL_SHARE = 0.25
def load_release_extractor(repo: Path) -> types.ModuleType:
"""`app.narrative.extract` as it was at the release tag."""
source = subprocess.run(
["git", "-C", str(repo), "show", f"{RELEASE}:{EXTRACT_PATH}"],
check=True, capture_output=True, text=True,
).stdout
import app.narrative # noqa: F401 the package the relative import needs
spec = importlib.util.spec_from_loader("app.narrative._extract_release", loader=None)
module = importlib.util.module_from_spec(spec)
module.__package__ = "app.narrative"
exec(compile(source, f"{RELEASE}:{EXTRACT_PATH}", "exec"), module.__dict__)
return module
def _unpack(blob):
if blob is None:
return None
try:
return json.loads(zlib.decompress(bytes(blob)).decode("utf-8"))
except (zlib.error, ValueError, UnicodeDecodeError):
return None
def turns_from_db(path: str):
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
rows = connection.execute(
"SELECT id, depth, text, context_snapshot FROM actions WHERE type = 'ai'"
).fetchall()
finally:
connection.close()
for action_id, depth, text, blob in rows:
snapshot = _unpack(blob) or {}
raw = snapshot.get("raw_output")
if isinstance(raw, str) and raw.strip():
yield {"source": path, "id": action_id, "depth": depth,
"input": raw, "input_kind": "raw_output"}
elif text:
yield {"source": path, "id": action_id, "depth": depth,
"input": text, "input_kind": "stored_text"}
def turns_from_bundle(path: str):
bundle = json.loads(Path(path).read_text())
for action in bundle.get("actions") or []:
if action.get("type") == "ai" and action.get("text"):
yield {"source": path, "id": action.get("id"), "depth": action.get("depth"),
"input": action["text"], "input_kind": "stored_text"}
def removed_lines(before: str, after: str) -> list[str]:
"""Lines present in `before` and gone from `after`, in order.
Compared with trailing whitespace ignored. The extractor strips the end of
every reply it cuts, so a story line that becomes the last line loses a
trailing space. A first version of this tool reported that space as a
rewritten line of story, which it is not.
"""
old = [line.rstrip() for line in before.split("\n")]
new = [line.rstrip() for line in after.split("\n")]
matcher = difflib.SequenceMatcher(a=old, b=new, autojunk=False)
gone: list[str] = []
for tag, a0, a1, _b0, _b1 in matcher.get_opcodes():
if tag in ("delete", "replace"):
gone.extend(old[a0:a1])
return gone
def replay(inputs, old, new) -> dict:
seen: set[str] = set()
unchanged = 0
changed: list[dict] = []
duplicates = 0
for turn in inputs:
digest = hashlib.sha256(turn["input"].encode()).hexdigest()
if digest in seen:
duplicates += 1
continue
seen.add(digest)
old_prose, _old_parsed, _old_raw = old.split(turn["input"])
new_prose, _new_parsed, _new_raw = new.split(turn["input"])
if old_prose == new_prose:
unchanged += 1
continue
lines = []
unexplained = 0
for line in removed_lines(old_prose, new_prose):
if not line.strip():
continue
rule = new.explain_removed_line(line)
if rule is None:
unexplained += 1
lines.append({"line": line, "rule": rule})
added = [line for line in removed_lines(new_prose, old_prose) if line.strip()]
share = 1 - len(new_prose) / max(1, len(old_prose))
flags = []
if unexplained:
flags.append("unexplained_removal")
if added:
flags.append("text_added_or_rewritten")
if share > LARGE_REMOVAL_SHARE:
flags.append("large_removal")
changed.append({
**{k: turn[k] for k in ("source", "id", "depth", "input_kind")},
"sha256": digest,
"old_prose": old_prose,
"new_prose": new_prose,
"removed": lines,
"added_or_rewritten": added,
"removed_chars": len(old_prose) - len(new_prose),
"removed_share": round(share, 4),
"flags": flags,
})
return {
"replayed": unchanged + len(changed),
"duplicates_skipped": duplicates,
"unchanged": unchanged,
"changed": len(changed),
"flagged": sum(1 for c in changed if c["flags"]),
"turns": changed,
}
def write_markdown(result: dict, path: Path) -> None:
out = [
"# A2 extractor replay",
"",
f"- replayed (unique replies): **{result['replayed']}**",
f"- duplicates skipped: {result['duplicates_skipped']}",
f"- unchanged: {result['unchanged']}",
f"- changed: **{result['changed']}**",
f"- flagged: **{result['flagged']}**",
"",
]
for index, turn in enumerate(result["turns"], 1):
out += [
f"## {index}. {Path(turn['source']).parent.name}/{Path(turn['source']).name}"
f" action {turn['id']} depth {turn['depth']} ({turn['input_kind']})",
"",
f"- removed chars: {turn['removed_chars']} ({turn['removed_share']:.1%})",
f"- flags: {', '.join(turn['flags']) or 'none'}",
"",
"Removed lines:",
"",
]
for item in turn["removed"]:
out.append(f"- `{item['rule'] or 'UNEXPLAINED'}` — {item['line']!r}")
out += ["", "<details><summary>v1.0.0 prose</summary>", "", "```text",
turn["old_prose"], "```", "</details>", "",
"<details><summary>current prose</summary>", "", "```text",
turn["new_prose"], "```", "</details>", ""]
path.write_text("\n".join(out))
def main(argv=None) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--db", action="append", default=[],
help="an evidence database, or a glob of them")
parser.add_argument("--bundle", action="append", default=[],
help="an exported bundle whose campaign has no database here")
parser.add_argument("--out", required=True)
args = parser.parse_args(argv)
repo = Path(__file__).resolve().parents[2]
from app.narrative import extract as current
old = load_release_extractor(repo)
dbs = sorted({p for pattern in args.db for p in glob.glob(pattern, recursive=True)})
def inputs():
for path in dbs:
yield from turns_from_db(path)
for path in args.bundle:
yield from turns_from_bundle(path)
result = replay(inputs(), old, current)
result["databases"] = dbs
result["bundles"] = args.bundle
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
(out / "replay.json").write_text(json.dumps(result, indent=2, ensure_ascii=False))
write_markdown(result, out / "replay.md")
print(json.dumps({k: result[k] for k in
("replayed", "duplicates_skipped", "unchanged", "changed", "flagged")}))
return 1 if result["flagged"] else 0
if __name__ == "__main__":
sys.exit(main())
+381
View File
@@ -0,0 +1,381 @@
"""v1.1 release Gate 9: a real v1.0.0 campaign, opened by the candidate.
python -m tools.v11_upgrade_check --v100 <worktree> --out <dir under $HOME>
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and
`AIDND_TEST_EMBED_MODEL`: the campaign has to be *played*, because memories,
summaries and narrative state are things a narrator produces. A schema-only
fixture would prove nothing about an upgrade, which is why §11 item 9 asks for a
database the v1.0.0 application built.
Three phases, each its own server process, so everything that survives crosses
as bytes on disk:
1. **v1.0.0 builds and plays.** The `432f041` tree serves the application: a
campaign is created, canon knowledge imported, turns played, a Save Point
taken, a turn undone so the head is not at the tip and Redo is available, and
a narration length chosen. Then that server stops, and a census is taken.
2. **The candidate opens the same file.** Nothing is copied; the candidate's own
migrations run against it. The census is taken again and compared field by
field.
3. **Bundles cross both ways.** The v1.0.0 export is imported by the candidate.
The candidate's export is offered back to v1.0.0, and whatever happens is
reported — the format string is unchanged, which is a reason to test backward
import, not a reason to assume it.
Settings hold only what the caller's environment names, and the evidence
directory lives under `$HOME`.
## Shapes this had to be written against, not guessed
- A turn is **SSE**: `POST /adventures/{id}/actions`, and a failed turn is an
`error` *event* inside an HTTP 200. Reading the status code would call every
failure a success.
- History is `GET /{id}/actions?limit=N` -> `{actions, total, has_more,
can_undo, can_redo}`. There is **no head-id field**, so the active head is
compared as the newest action plus the two flags.
- Save Points are **checkpoints**. Creating one after an Undo names the undone
position, deliberately.
- There is **no summaries route**; summaries and post-turn health both come from
`GET /{id}/derived`.
- Campaign switches are `PATCH /adventures/{id}` with `memory_bank_enabled` and
`auto_summarize` — not the names a reader would guess.
- Knowledge import is **multipart**, as `m11_long_run` and `m11_offline` build it.
"""
from __future__ import annotations
import argparse
import json
import os
import shutil
import sqlite3
import subprocess
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import free_port, require_under_home # noqa: E402
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
TURN_TIMEOUT = 900
#: What the evidence database is left holding. The campaign has to be played
#: against a real narrator, but nothing about the *upgrade* depends on which
#: host that was, and §11 item 9 asks for no real hostname in this database.
PLACEHOLDER_ENDPOINT = "http://127.0.0.1:11434/v1"
#: What the upgrade must preserve, compared exactly on both sides. `settings` is
#: included because a migration that silently rewrote an endpoint would be a real
#: defect; `schema_version` is read from the file rather than the API.
CENSUS = ("transcript", "newest_action", "total", "can_undo", "can_redo",
"checkpoints", "state", "memories", "summaries", "knowledge",
"narration_length", "memory_bank_enabled", "auto_summarize",
"settings", "schema_version")
CANON_MD = """# Westhaven
The abbey bell is rung only for a death. Mara keeps the harbour ledger.
Aldric carries a silver key he will not explain.
"""
class App:
"""One application process, from whichever tree it is given."""
def __init__(self, tree: Path, db: Path, log: Path, label: str):
self.tree, self.db, self.label = tree, db, label
self.port = free_port()
self.log = log
handle = open(log, "ab")
self.proc = subprocess.Popen(
[str(Path(__file__).resolve().parent.parent / ".venv/bin/uvicorn"),
"app.main:app", "--host", "127.0.0.1", "--port", str(self.port)],
cwd=str(tree / "backend"), stdout=handle, stderr=subprocess.STDOUT,
env={**os.environ, "AIDND_DB_PATH": str(db),
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
)
self.url = f"http://127.0.0.1:{self.port}"
deadline = time.monotonic() + 120
while time.monotonic() < deadline:
if self.proc.poll() is not None:
raise RuntimeError(f"{label} exited early; see {log}")
try:
urllib.request.urlopen(self.url + "/api/settings", timeout=2)
print(f" {label} serving {db.name} on {self.url}", flush=True)
return
except Exception:
time.sleep(0.2)
raise RuntimeError(f"{label} never became ready; see {log}")
def call(self, method: str, path: str, payload=None, timeout=120):
data = json.dumps(payload).encode() if payload is not None else None
request = urllib.request.Request(
f"{self.url}/api{path}", data=data, method=method,
headers={"Content-Type": "application/json"} if data else {})
with urllib.request.urlopen(request, timeout=timeout) as response:
body = response.read().decode()
return json.loads(body) if body else None
def stream(self, path: str, payload) -> list[dict]:
"""A turn. A failed turn is an event in the stream, not a status code."""
request = urllib.request.Request(
f"{self.url}/api{path}", data=json.dumps(payload).encode(),
method="POST", headers={"Content-Type": "application/json"})
events: list[dict] = []
with urllib.request.urlopen(request, timeout=TURN_TIMEOUT) as response:
for raw in response:
line = raw.decode(errors="replace").strip()
if line.startswith("data:"):
try:
events.append(json.loads(line[5:].strip()))
except json.JSONDecodeError:
pass
return events
def upload(self, adv: int, name: str, body: str, classification: str) -> dict:
boundary = "----v11upgrade"
parts = (
f"--{boundary}\r\nContent-Disposition: form-data; name=\"classification\""
f"\r\n\r\n{classification}\r\n"
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
f"--{boundary}--\r\n"
).encode()
request = urllib.request.Request(
f"{self.url}/api/adventures/{adv}/knowledge", data=parts, method="POST",
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"})
with urllib.request.urlopen(request, timeout=120) as response:
return json.loads(response.read().decode())
def stop(self) -> None:
if self.proc.poll() is None:
self.proc.terminate()
try:
self.proc.wait(timeout=30)
except subprocess.TimeoutExpired:
self.proc.kill()
def schema_version(db: Path) -> int:
connection = sqlite3.connect(f"file:{db}?mode=ro", uri=True)
try:
return connection.execute("PRAGMA user_version").fetchone()[0]
finally:
connection.close()
def census(app: App, adv: int, db: Path) -> dict:
page = app.call("GET", f"/adventures/{adv}/actions?limit=500") or {}
actions = page.get("actions") or []
derived = app.call("GET", f"/adventures/{adv}/derived") or {}
adventure = app.call("GET", f"/adventures/{adv}") or {}
settings = app.call("GET", "/settings") or {}
checkpoints = app.call("GET", f"/adventures/{adv}/checkpoints") or []
memories = app.call("GET", f"/adventures/{adv}/memories") or []
state = app.call("GET", f"/adventures/{adv}/state") or {}
knowledge = app.call("GET", f"/adventures/{adv}/knowledge") or []
newest = actions[-1] if actions else {}
return {
"transcript": [(a.get("type"), (a.get("text") or "")[:300]) for a in actions],
# No head id is exposed; the head is the newest action on the read line
# plus the two flags the page carries.
"newest_action": ((newest.get("type"), (newest.get("text") or "")[:300])
if newest else None),
"total": page.get("total"),
"can_undo": page.get("can_undo"),
"can_redo": page.get("can_redo"),
"checkpoints": sorted((c.get("name"), c.get("depth"), c.get("branch_id"),
c.get("on_path"), c.get("resolved"))
for c in checkpoints),
"state": state.get("document") if isinstance(state, dict) else state,
"memories": sorted((m.get("text") or "")[:200] for m in memories),
"summaries": sorted((s.get("text") or "")[:200]
for s in (derived.get("summaries") or [])),
"knowledge": sorted((k.get("filename") or k.get("title"),
k.get("classification")) for k in knowledge),
"narration_length": adventure.get("narration_length"),
"memory_bank_enabled": adventure.get("memory_bank_enabled"),
"auto_summarize": adventure.get("auto_summarize"),
"settings": {k: settings.get(k)
for k in ("endpoint_url", "model", "max_output_tokens")},
"schema_version": schema_version(db),
}
def play(app: App, adv: int, text: str) -> bool:
events = app.stream(f"/adventures/{adv}/actions", {"type": "do", "text": text})
errors = [e for e in events if e.get("type") == "error"]
if errors:
print(f" turn refused: {errors[0].get('detail', '')[:150]}", flush=True)
return False
return True
def build_v100_campaign(app: App) -> int:
created = app.call("POST", "/adventures", {
"title": "Upgrade Evidence",
"opening": "Rain over Westhaven, and the abbey bell tolling.",
"canon_rules": ["The dead do not return."],
"persona_name": "Aldric",
})
adv = created["id"]
app.call("PUT", "/settings", {
"endpoint_url": ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL,
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
# Memory and summaries are per-campaign switches defaulting to off, so the
# census would otherwise have nothing to compare.
app.call("PATCH", f"/adventures/{adv}", {
"narration_length": "brief", "memory_bank_enabled": True,
"auto_summarize": True})
app.upload(adv, "westhaven-canon.md", CANON_MD, "canon")
# Enough turns that memories and summaries actually exist. A memory needs
# MEMORY_INTERVAL (6) actions plus SETTLE_SLACK (1) settled past the anchor,
# and a summary needs SUMMARY_INTERVAL (15) uncovered actions — both counted
# in *actions*, and a turn writes two. A first pass at this gate played five
# turns, wrote neither, and compared 0 against 0, which proves nothing about
# whether the upgrade preserves them.
beats = [
"I ask Mara what the bell means.",
"I show her the silver key.",
"I follow her to the harbour ledger.",
"I ask who else knows about the key.",
"I read the ledger's last page aloud.",
"I ask the ferryman about the fen road.",
"I wait out the rain and watch the harbour.",
"I ask Mara about the abbey's sealed crypt.",
"I count the entries against the tide table.",
"I ask who signed for the last shipment.",
"I walk the quay to the chandler's door.",
"I ask the chandler what he remembers of that night.",
"I show the chandler the key.",
"I return to Mara with what he said.",
"I ask Mara what she means to do now.",
"I agree to meet her at first light.",
"I take the long way back along the ridge.",
"I check whether anyone followed me.",
"I write down what I have learned so far.",
"I sleep, and wake before the bell.",
]
played = 0
for text in beats:
if play(app, adv, text):
played += 1
print(f" {played} turns accepted by v1.0.0", flush=True)
app.call("POST", f"/adventures/{adv}/checkpoints", {"name": "before the ledger"})
# One Undo, so the head is not at the retained tip and Redo is available.
app.call("POST", f"/adventures/{adv}/undo", {})
# §11 item 9: the database must carry **loopback or placeholder settings with
# no real hostnames**. The turns above needed a real narrator, so the
# endpoint is reset to loopback once the story exists — before the census is
# taken and before either bundle is exported.
app.call("PUT", "/settings", {
"endpoint_url": PLACEHOLDER_ENDPOINT, "model": MODEL,
"embedding_model": EMBED_MODEL,
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
return adv
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--v100", required=True, help="the v1.0.0 worktree")
parser.add_argument("--out", required=True)
args = parser.parse_args()
if not (ENDPOINT and MODEL and EMBED_MODEL):
print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL")
return 2
out = require_under_home(Path(args.out).expanduser())
shutil.rmtree(out, ignore_errors=True)
out.mkdir(parents=True)
v100_tree = Path(args.v100).expanduser().resolve()
candidate_tree = Path(__file__).resolve().parent.parent.parent
db = out / "campaign.db"
results: dict = {"started": datetime.now().isoformat(timespec="seconds"),
"v100_tree": str(v100_tree), "candidate": str(candidate_tree)}
failures: list[str] = []
print("phase 1 — v1.0.0 builds and plays the campaign")
app = App(v100_tree, db, out / "v100-server.log", "v1.0.0")
try:
adv = build_v100_campaign(app)
before = census(app, adv, db)
v100_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
(out / "v100-export.json").write_text(json.dumps(v100_bundle))
finally:
app.stop()
results["adventure"], results["before"] = adv, before
print(f" {before['total']} actions, redo={before['can_redo']}, "
f"checkpoints={len(before['checkpoints'])}, memories={len(before['memories'])}, "
f"summaries={len(before['summaries'])}, knowledge={len(before['knowledge'])}, "
f"schema={before['schema_version']}")
print("\nphase 2 — the candidate opens that same database file")
app = App(candidate_tree, db, out / "candidate-server.log", "candidate")
try:
after = census(app, adv, db)
v11_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
(out / "v11-export.json").write_text(json.dumps(v11_bundle))
try:
imported = app.call("POST", "/adventures/import", v100_bundle, timeout=600)
results["v100_bundle_into_v11"] = {"status": "imported",
"id": (imported or {}).get("id")}
print(" the v1.0.0 bundle imported into the candidate")
except urllib.error.HTTPError as exc:
results["v100_bundle_into_v11"] = {
"status": "refused", "code": exc.code,
"detail": exc.read().decode()[:300]}
failures.append("v100_bundle_into_v11")
print(f" the candidate REFUSED the v1.0.0 bundle: {exc.code}")
finally:
app.stop()
results["after"] = after
print(" comparing the census, field by field:")
for field in CENSUS:
same = before.get(field) == after.get(field)
print(f" {'ok ' if same else 'DIFF'} {field}")
if not same:
failures.append(field)
results.setdefault("differences", {})[field] = {
"before": before.get(field), "after": after.get(field)}
print("\nphase 3 — the candidate's bundle offered back to v1.0.0")
app = App(v100_tree, out / "backward.db", out / "v100-backward.log", "v1.0.0")
try:
try:
back = app.call("POST", "/adventures/import", v11_bundle, timeout=600)
results["v11_bundle_into_v100"] = {"status": "imported",
"id": (back or {}).get("id")}
print(" v1.0.0 ACCEPTED the v1.1 bundle")
except urllib.error.HTTPError as exc:
detail = exc.read().decode()[:400]
results["v11_bundle_into_v100"] = {"status": "refused", "code": exc.code,
"detail": detail}
# Reported, not failed: §11 item 9 asks for the result, and the
# owner's brief asks whether a refusal breaks the compatibility
# promise — a judgement, not an assertion this script may make.
print(f" v1.0.0 REFUSED the v1.1 bundle: {exc.code} {detail[:160]}")
finally:
app.stop()
results["failures"] = failures
(out / "upgrade-report.json").write_text(json.dumps(results, indent=2, default=str))
print(f"\n{'PASS' if not failures else 'FAIL'}: {len(failures)} field(s) differ "
f"-> {out / 'upgrade-report.json'}")
return 1 if failures else 0
if __name__ == "__main__":
raise SystemExit(main())
+211
View File
@@ -0,0 +1,211 @@
"""v1.1 WP-A1: real turns, and what the server said it read.
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 AIDND_TEST_MODEL=<model> \\
.venv/bin/python -m tools.v11_window_accounting \\
--bundle "$HOME/m11-evidence/m04-final/bundle.json" --turns 4 \\
--out "$HOME/v11-evidence/a1-accounting/<label>"
Run from `backend/`. The evidence that matters is at the edge of the window, and
a new campaign takes dozens of turns to reach it. So this imports a long
campaign, by default the v1 evidence run's 207-action bundle, and every turn is
assembled against a full window from the first. A real narrator is then asked
for `--turns` turns, and each one is written out with:
configured budget, verified window, and where the window came from
the application's estimate of what it sent (its count, plus the text the
provider adds)
the server's own prompt-token count, from the usage it reported
the reply allocation and the safety reserve
the observed margin: window - reply allocation - the server's count
the accounting status: fits, exceeded, truncation_suspected or unknown
The first turn on a cold model cannot verify the window: `/api/ps` knows nothing
until the model is loaded. That is M11's behaviour, and the row says so rather
than hiding the turn.
The database lives in `--out`, not in `/tmp`, so the stored snapshots behind
every row can be read again. The endpoint is read from the environment and is
never written into a committed file.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
import time
from pathlib import Path
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
TURNS = [
"I look around carefully and take stock of where I am.",
"I ask the nearest person what has happened since I was last here.",
"I check what I am carrying.",
"I move on towards the place I meant to reach.",
"I wait and listen.",
"I say, \"Tell me the part you left out.\"",
]
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--bundle", default="",
help="a campaign bundle to import, so turns start at a full window")
parser.add_argument("--turns", type=int, default=4)
parser.add_argument("--budget", type=int, default=16384)
parser.add_argument("--max-output", type=int, default=500)
parser.add_argument("--timeout", type=int, default=900)
parser.add_argument("--out", required=True)
parser.add_argument("--unload-first", action="store_true",
help="ask the configured server to unload the model before turn 1, "
"so the first turn starts cold (the A1 corrective test)")
args = parser.parse_args()
if not (ENDPOINT and MODEL):
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
return 2
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
db_path = out / "accounting.db"
if db_path.exists():
print(f"{db_path} exists; choose a new --out")
return 2
os.environ["AIDND_DB_PATH"] = str(db_path)
os.environ.pop("AIDND_DATABASE_URL", None)
os.environ.pop("DATABASE_URL", None)
from fastapi import Depends
from fastapi.testclient import TestClient
from sqlalchemy.orm import undefer
from app import auth, limits, models
from app.database import Base, SessionLocal, engine, get_db
from app.main import app
limits.check_row_cap = lambda *a, **k: None
Base.metadata.create_all(bind=engine)
with SessionLocal() as db:
user = models.User(is_guest=False, email="accounting@example.com")
db.add(user)
db.flush()
db.add(models.Settings(
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
context_token_budget=args.budget, max_output_tokens=args.max_output,
model_timeout_seconds=args.timeout,
))
db.commit()
user_id = user.id
app.dependency_overrides[auth.get_current_user] = (
lambda db=Depends(get_db): db.get(models.User, user_id)
)
client = TestClient(app)
if args.bundle:
bundle = json.loads(Path(args.bundle).read_text())
# The evidence campaign had its memory bank and auto-summarise on. Here
# they would only add post-turn model calls between the measured turns,
# on the same host, and fail noisily as the in-process client closes.
# This tool measures the turn's own prompt, which neither changes.
bundle["memoryBankEnabled"] = False
bundle["autoSummarize"] = False
imported = client.post("/api/adventures/import", json=bundle)
imported.raise_for_status()
adv = imported.json()["id"]
else:
created = client.post("/api/adventures", json={
"title": "Window accounting", "opening": "A quiet road at dusk."})
created.raise_for_status()
adv = created.json()["id"]
if args.unload_first:
# The configured endpoint only, under the same policy and TLS trust as a turn.
import asyncio
import httpx
from app import contextwindow, endpoints, tlstrust
reason = endpoints.rejection_reason(ENDPOINT)
if reason:
print(f"endpoint refused: {reason}")
return 2
base = contextwindow.native_base(ENDPOINT)
with httpx.Client(verify=tlstrust.ssl_context(), timeout=120) as http:
unloaded = http.post(f"{base}/api/generate", json={"model": MODEL, "keep_alive": 0})
resident = [m.get("name") for m in http.get(f"{base}/api/ps").json().get("models", [])]
contextwindow.cache_clear()
print(f"unload: HTTP {unloaded.status_code} {unloaded.text[:120]} | resident now: {resident}")
rows: list[dict] = []
timeline = (out / "turns.jsonl").open("a")
print(f"model {MODEL}, budget {args.budget}, reply {args.max_output}")
print(f"{'#':>2} {'status':22} {'window':>13} {'estimate':>8} {'server':>7} "
f"{'reserve':>7} {'margin':>7} {'sec':>5}")
for index in range(args.turns):
text = TURNS[index % len(TURNS)]
started = time.monotonic()
response = client.post(f"/api/adventures/{adv}/actions",
json={"type": "do", "text": text})
seconds = round(time.monotonic() - started, 1)
error = None
if response.status_code != 200 or '"type": "error"' in response.text:
error = response.text[-400:]
with SessionLocal() as db:
action = (
db.query(models.Action)
.filter(models.Action.adventure_id == adv, models.Action.type == "ai")
.options(undefer(models.Action.context_snapshot))
.order_by(models.Action.id.desc()).first()
)
snapshot = (action.context_snapshot or {}) if action else {}
tokens = snapshot.get("tokens") or {}
window = snapshot.get("window") or {}
accounting = snapshot.get("accounting") or {}
row = {
"turn": index + 1,
"action_id": action.id if action else None,
"seconds": seconds,
"error": error,
"configured_budget": tokens.get("configured_budget"),
"effective_budget": tokens.get("budget"),
"window_verified": window.get("verified"),
"window_tokens": window.get("tokens"),
"window_source": window.get("source"),
"preflight_attempted": (window.get("preflight") or {}).get("attempted"),
"preflight_loaded": (window.get("preflight") or {}).get("loaded"),
"preflight_verified_before": (window.get("preflight") or {}).get("verified_before"),
"preflight_verified_after": (window.get("preflight") or {}).get("verified_after"),
"preflight_detail": (window.get("preflight") or {}).get("detail"),
"app_prompt_tokens": tokens.get("total"),
"transport_tokens": tokens.get("transport"),
"app_estimate": tokens.get("estimate"),
"output_reserve": tokens.get("output_reserve"),
"safety_reserve": tokens.get("safety_reserve"),
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
"difference": accounting.get("difference"),
"observed_margin": accounting.get("observed_margin"),
"status": accounting.get("status"),
"history_included": (snapshot.get("history") or {}).get("included"),
"history_total": (snapshot.get("history") or {}).get("total"),
}
rows.append(row)
timeline.write(json.dumps(row) + "\n")
timeline.flush()
print(f"{row['turn']:>2} {str(row['status'] if not error else 'ERROR'):22} "
f"{str(row['window_tokens'])+('v' if row['window_verified'] else '?'):>13} "
f"{str(row['app_estimate']):>8} {str(row['server_prompt_tokens']):>7} "
f"{str(row['safety_reserve']):>7} {str(row['observed_margin']):>7} {seconds:>5}")
if error:
print(f" error: {error[:200]}")
timeline.close()
(out / "summary.json").write_text(json.dumps({
"model": MODEL, "budget": args.budget, "max_output_tokens": args.max_output,
"bundle": args.bundle, "rows": rows,
}, indent=2))
app.dependency_overrides.clear()
return 0 if all(r["error"] is None for r in rows) else 1
if __name__ == "__main__":
sys.exit(main())
+353
View File
@@ -0,0 +1,353 @@
"""v1.1 WP-D criterion 3: the backup completes **through the UI**.
python -m tools.wpd_backup_ui --out <dir under $HOME> [--case real|large|both] [--show]
Run from `backend/`, with `frontend/dist` already built.
WP-D's first pass drove `POST /api/backups` — the endpoint the *Back up now*
button calls — on a 2.3 MB campaign database and a 117 MB one. That is evidence
about the implementation, and the plan's criterion 3 asks for something else:
that the backup *completes through the UI* on both. A reader does not call an
endpoint. This drives the reader-facing control in a real Firefox, against the
production build served by FastAPI, exactly as WP-C's harness does.
**No narrator and no inference.** A backup needs neither, so nothing here
touches a model host.
What it refuses to call a pass, per the brief:
- the click does nothing (no toast, no file);
- the request fails (an error toast);
- no backup file appears on disk;
- the finished copy does not pass a full `PRAGMA integrity_check`;
- the UI reports an error.
Every wait is on a condition the page or the filesystem can show. Nothing here
sleeps and then asserts.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import shutil
import sqlite3
import sys
import time
from datetime import datetime
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.m11_webdriver import ( # noqa: E402
Browser, Site, geckodriver_version, require_under_home,
)
BACKEND = Path(__file__).resolve().parent.parent
HOME = Path.home()
#: The two databases criterion 3 names. Both are copied before use: the first is
#: the M11 evidence campaign and must not be written to, and the second is the
#: 117 MB application database built for WP-D's timing.
DEFAULT_REAL = HOME / "m11-evidence/m04-final/campaign.db"
DEFAULT_LARGE = HOME / "v11-evidence/wp-d/endpoint/campaign-100mb/campaign.db"
BLOCK = '[data-testid="database-backup"]'
BUTTON = f'{BLOCK} button.primary'
class Checks:
"""Results, with the discipline that an unrun check is not a passing one."""
def __init__(self) -> None:
self.rows: list[dict] = []
def record(self, case: str, name: str, ok: bool, detail: str = "") -> bool:
self.rows.append({"case": case, "check": name,
"result": "PASS" if ok else "FAIL", "detail": detail})
print(f" {'ok ' if ok else 'FAIL'} {case:6} {name}"
+ (f" — {detail}" if detail else ""), flush=True)
return ok
@property
def failed(self) -> list[dict]:
return [r for r in self.rows if r["result"] == "FAIL"]
# ----------------------------------------------------------------- helpers
def integrity_of(path: Path) -> tuple[str, float]:
"""The full check on a finished copy, and what it cost, measured here.
The application runs its own `integrity_check` before keeping the file; this
is an independent second opinion on the artefact the UI produced, and it is
where criterion 3's 'time for integrity_check' comes from.
"""
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
try:
started = time.perf_counter()
rows = connection.execute("PRAGMA integrity_check").fetchall()
elapsed = time.perf_counter() - started
finally:
connection.close()
return ", ".join(str(r[0]) for r in rows), elapsed
def digest(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()[:16]
def backups_in(db_path: Path) -> dict[str, int]:
directory = db_path.parent / "backups"
if not directory.exists():
return {}
return {p.name: p.stat().st_size for p in directory.iterdir() if p.is_file()}
def wait_for_new_backup(db_path: Path, before: set[str], *, timeout: float = 300,
poll: float = 0.2) -> Path | None:
"""The backup file the application wrote, once it is really there.
A file condition, not a sleep: a name that was not there before, not a
`.partial`, and a size that has stopped growing.
"""
directory = db_path.parent / "backups"
deadline = time.monotonic() + timeout
sizes: dict[str, int] = {}
while time.monotonic() < deadline:
if directory.exists():
for path in directory.iterdir():
if not path.is_file() or path.name in before:
continue
if path.name.endswith(".partial"):
continue
size = path.stat().st_size
if size > 0 and sizes.get(path.name) == size:
return path
sizes[path.name] = size
time.sleep(poll)
return None
def toasts(browser: Browser) -> list[dict]:
"""What the page is telling the reader — the message apart from its mark.
A toast renders a decorative mark before the message:
`<span class="toast-mark" aria-hidden="true">❖</span><span>…</span>`. So the
button's `textContent` begins with that character, and the first run of this
tool matched `textContent.startswith('Backup written:')` and reported six
*passing* behaviours as failures — the UI had said exactly what it should,
and the assertion was reading the mark. `message` is the message span alone;
`text` is kept whole for the evidence record.
"""
return browser.js("""
const host = document.querySelector('.toast-host');
if (!host) return [];
return [...host.querySelectorAll('button.toast')].map(b => {
const span = b.querySelector('span:not(.toast-mark)');
return {
error: b.classList.contains('toast-error'),
message: (span ? span.textContent : b.textContent).trim(),
text: b.textContent.trim(),
};
});
""") or []
def open_backup_panel(browser: Browser, site: Site, checks: Checks, case: str) -> bool:
"""Reach the control the way a reader does: the nav link, then the panel."""
browser.go(site.url)
browser.wait_for(".topnav", timeout=60)
link = browser.find('.nav-links a[href="/settings"]', required=False)
if link is None:
return checks.record(case, "the Settings link is in the navigation", False)
browser.click(link)
opened = browser.wait_js(f"!!document.querySelector('{BLOCK}')", timeout=60)
if not checks.record(case, "Settings opens from the navigation link", opened,
browser.url):
return False
summary = browser.find(f"{BLOCK} summary", required=False)
if summary is None:
return checks.record(case, "the backup panel has a disclosure", False)
browser.click(summary)
# The panel is a <details>: it loads what is already on disk when it opens,
# so waiting for the directory line proves the application answered.
shown = browser.wait_js(
f"document.querySelector('{BLOCK}').open === true"
f" && !!document.querySelector('{BUTTON}')", timeout=30)
checks.record(case, "the backup panel opens and shows its control", shown)
listed = browser.wait_js(f"!!document.querySelector('{BLOCK} code')", timeout=30)
checks.record(case, "the panel reports where backups are written", listed)
return shown
def click_back_up_now(browser: Browser, site: Site, db: Path, checks: Checks,
case: str, label: str) -> dict:
"""One press of the button, judged by what the page and the disk then show."""
before = set(backups_in(db))
# Clear anything still on screen, so the toast this press produces is the
# one that is read back rather than a leftover from the previous press.
browser.js("""
for (const b of document.querySelectorAll('.toast-host button.toast')) b.click();
return true;
""")
button = browser.find(BUTTON, required=False)
if button is None:
checks.record(case, f"{label}: the control is on the page", False)
return {}
started = time.perf_counter()
browser.click(button)
# Either outcome ends the wait, so a failure is reported as a failure rather
# than as a timeout.
settled = browser.wait_js(
"(() => { const t = [...document.querySelectorAll('.toast-host button.toast')];"
" return t.length > 0; })()", timeout=300)
elapsed = time.perf_counter() - started
shown = toasts(browser)
errors = [t for t in shown if t["error"]]
written = [t for t in shown if t["message"].startswith("Backup written:")]
checks.record(case, f"{label}: the click produced a visible result", settled,
json.dumps(shown)[:200])
checks.record(case, f"{label}: the UI reports no error",
not errors, json.dumps(errors)[:300])
checks.record(case, f"{label}: the UI reports the backup was written",
bool(written), json.dumps(written)[:200])
idle = browser.wait_js(
"(() => { const b = document.querySelector(%s);"
" return !!b && b.textContent.trim() === 'Back up now'; })()"
% json.dumps(BUTTON), timeout=120)
checks.record(case, f"{label}: the control returns from 'Backing up…'", idle)
produced = wait_for_new_backup(db, before)
checks.record(case, f"{label}: a backup file was physically produced",
produced is not None, str(produced))
if produced is None:
return {"seconds": round(elapsed, 3), "toasts": shown}
named = any(produced.name in t["message"] for t in written)
checks.record(case, f"{label}: the UI names the file that appeared", named,
f"{produced.name} — {written[0]['message'] if written else ''}")
relisted = browser.wait_js(
"[...document.querySelectorAll('%s .backup-list code')]"
".some(c => c.textContent.trim() === %s)" % (BLOCK, json.dumps(produced.name)),
timeout=60)
checks.record(case, f"{label}: the new backup appears in the panel's list", relisted)
verdict, integrity_seconds = integrity_of(produced)
checks.record(case, f"{label}: the finished copy passes full integrity_check",
verdict == "ok", f"{verdict} in {integrity_seconds * 1000:.1f} ms")
return {
"file": str(produced),
"bytes": produced.stat().st_size,
"sha256_16": digest(produced),
"seconds": round(elapsed, 3),
"integrity": verdict,
"integrity_seconds": round(integrity_seconds, 4),
"toasts": shown,
}
def run_case(case: str, source: Path, out: Path, checks: Checks, *, show: bool,
twice: bool) -> dict:
print(f"\n=== {case}: {source} ===", flush=True)
work = out / case
shutil.rmtree(work, ignore_errors=True)
(work / "data").mkdir(parents=True)
db = work / "data" / "campaign.db"
copy_started = time.perf_counter()
shutil.copy2(source, db)
copy_seconds = time.perf_counter() - copy_started
size = db.stat().st_size
print(f" copied {size:,} bytes in {copy_seconds:.2f}s -> {db}", flush=True)
site = Site(BACKEND, db, work / "server.log")
browser = Browser(headless=not show, log=work / "geckodriver.log")
result: dict = {"database": str(source), "bytes": size,
"served_at": site.url, "firefox": browser.version}
try:
if not open_backup_panel(browser, site, checks, case):
return result
result["first"] = click_back_up_now(browser, site, db, checks, case, "backup")
browser.screenshot(work / "back-up-now.png")
if twice and result["first"].get("file"):
kept = Path(result["first"]["file"])
before_bytes, before_digest = kept.stat().st_size, digest(kept)
result["second"] = click_back_up_now(browser, site, db, checks, case,
"second backup")
still_there = kept.exists()
checks.record(case, "the earlier backup still exists", still_there)
if still_there:
checks.record(
case, "and is byte-for-byte what it was",
kept.stat().st_size == before_bytes and digest(kept) == before_digest,
f"{before_bytes:,} bytes, sha256:{before_digest}")
verdict, _ = integrity_of(kept)
checks.record(case, "and still passes integrity_check", verdict == "ok",
verdict)
if result["second"].get("file"):
checks.record(case, "the second backup is a different file",
result["second"]["file"] != result["first"]["file"],
Path(result["second"]["file"]).name)
finally:
browser.quit()
site.stop()
return result
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--out", required=True,
help="evidence directory, which must be under $HOME")
parser.add_argument("--case", default="both", choices=["real", "large", "both"])
parser.add_argument("--show", action="store_true", help="run Firefox visibly")
parser.add_argument("--real-db", default=str(DEFAULT_REAL))
parser.add_argument("--large-db", default=str(DEFAULT_LARGE))
args = parser.parse_args()
out = require_under_home(Path(args.out).expanduser())
out.mkdir(parents=True, exist_ok=True)
if not (BACKEND.parent / "frontend/dist/index.html").exists():
print("frontend/dist is not built", file=sys.stderr)
return 2
checks = Checks()
started = datetime.now()
results: dict[str, dict] = {}
wanted = [("real", Path(args.real_db).expanduser(), True)] if args.case != "large" else []
if args.case != "real":
wanted.append(("large", Path(args.large_db).expanduser(), False))
print(f"WP-D criterion 3 — the backup through the UI. geckodriver "
f"{geckodriver_version()}", flush=True)
for case, source, twice in wanted:
if not source.exists():
checks.record(case, "the database is present", False, str(source))
continue
results[case] = run_case(case, source, out, checks, show=args.show, twice=twice)
report = {
"started": started.isoformat(timespec="seconds"),
"seconds": round((datetime.now() - started).total_seconds()),
"cases": results,
"checks": checks.rows,
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
"failed": len(checks.failed),
}
(out / "backup-ui-report.json").write_text(json.dumps(report, indent=2))
print(f"\n{report['passed']} passed, {report['failed']} failed "
f"-> {out / 'backup-ui-report.json'}")
return 1 if checks.failed else 0
if __name__ == "__main__":
raise SystemExit(main())
+26 -1
View File
@@ -151,7 +151,32 @@ export const api = {
sendAction: (advId, payload, handlers, signal) =>
streamSSE(`/adventures/${advId}/actions`, payload, handlers, signal),
retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal),
exportAdventure: (id) => request(`/adventures/${id}/export`),
// v1.1 WP-D: the bundle, plus what the server says about importing it back.
// The body is the bundle and nothing else — the browser saves exactly those
// bytes — so the size and the warning come back in headers.
exportAdventure: async (id) => {
const resp = await fetch(`/api/adventures/${id}/export`, {
headers: { 'Content-Type': 'application/json' },
})
if (!resp.ok) {
let detail = resp.statusText
try { detail = (await resp.json()).detail || detail } catch { /* non-JSON */ }
throw new Error(detail)
}
const bundle = await resp.json()
const number = (name) => {
const raw = Number(resp.headers.get(name))
return Number.isFinite(raw) && raw > 0 ? raw : null
}
return {
bundle,
exportBytes: number('X-Export-Bytes'),
importLimitBytes: number('X-Import-Limit-Bytes'),
// Absent header (an older server) means nothing is claimed either way.
importable: resp.headers.get('X-Importable-By-This-Version') !== 'false',
warning: resp.headers.get('X-Export-Warning') || null,
}
},
importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }),
// M9. A verified copy of the whole database, which is a different tool from
+19
View File
@@ -45,6 +45,25 @@ export function classifyError(message) {
const detail = String(message || '').trim() || 'No detail was reported.'
const low = detail.toLowerCase()
// ---- A correction the story refused ----
//
// v1.1 WP-C, found in the browser. The State panel's refusals ("That
// correction can't be applied — no fact 'f1' to invalidate.") matched none of
// the rules below and fell through to "Generation failed", with a Retry button
// and a line about what you typed. No turn was attempted and nothing was
// typed: the reader corrected the story's state and the rules refused it. So
// it is said as that, and the reason stays under the technical details.
if (low.includes("correction can't be applied")) {
return {
kind: KIND.STATE,
title: 'That correction was not applied',
detail,
hint: 'Nothing in the story or its state was changed. The reason is in the technical details.',
retryable: false,
keptInput: false,
}
}
// ---- Model / endpoint: the story cannot be told at all ----
// "No model configured — set one in Settings."
+4 -2
View File
@@ -66,10 +66,12 @@ export default function Campaigns() {
const exportOne = async (campaign) => {
try {
const bundle = await api.exportAdventure(campaign.id)
const { bundle, warning } = await api.exportAdventure(campaign.id)
const safe = (campaign.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
downloadJSON(bundle, `${safe}.json`)
toast('Campaign exported.')
// v1.1 WP-D: the file is written either way. A campaign too large for this
// version to import back says so now, not when it is needed.
toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
} catch (err) {
toast(classifyError(err.message).detail, 'error')
}
+7 -3
View File
@@ -55,9 +55,13 @@ export function FailureNotice({ failure, partial, onRetry, onDismiss }) {
{failure.hint && <p className="failure-hint">{failure.hint}</p>}
<p className="failure-kept">
Your story is unchanged, and what you typed is still in the box below.
</p>
{/* Every failed turn keeps what was typed (A05). A refused correction
had nothing typed in the box, so it does not claim to (v1.1 WP-C). */}
{failure.keptInput !== false && (
<p className="failure-kept">
Your story is unchanged, and what you typed is still in the box below.
</p>
)}
{partial ? (
<details className="failure-partial">
+35 -2
View File
@@ -29,8 +29,8 @@ beforeEach(() => { vi.restoreAllMocks() })
describe('the failure notice states only what is true', () => {
it('claims the typed words were kept — so the caller must actually keep them', async () => {
// This assertion is the contract the defect broke. The notice is
// unconditional, so every path that shows it owes the reader their text.
// This assertion is the contract the defect broke. Every failed turn shows
// it, so every path that shows it owes the reader their text.
await renderWith(
<FailureNotice failure={classifyError('Could not connect to http://127.0.0.1:9/v1')}
partial={null} onRetry={vi.fn()} onDismiss={vi.fn()} />,
@@ -64,6 +64,39 @@ describe('the failure notice states only what is true', () => {
expect(onRetry).toHaveBeenCalled()
})
// v1.1 WP-C, found in the browser: a correction the story refused fell through
// to "Generation failed", offered to try the turn again, and claimed the
// typed words were kept — none of which a State-panel correction involves.
const REFUSAL = "That correction can't be applied — no fact 'f1' to invalidate."
it('classifies a refused correction as a state refusal, not a failed turn', () => {
const refusal = classifyError(REFUSAL)
expect(refusal.kind).toBe(KIND.STATE)
expect(refusal.title).toBe('That correction was not applied')
expect(refusal.retryable).toBe(false)
expect(refusal.keptInput).toBe(false)
expect(refusal.detail).toBe(REFUSAL)
})
it('shows a refused correction with its reason, and no turn to retry', async () => {
await renderWith(
<FailureNotice failure={classifyError(REFUSAL)} partial={null}
onRetry={vi.fn()} onDismiss={vi.fn()} />,
)
expect(screen.getByText('That correction was not applied')).toBeInTheDocument()
expect(screen.getByText(/no fact 'f1' to invalidate/)).toBeInTheDocument()
expect(screen.queryByRole('button', { name: 'Try that turn again' })).toBeNull()
expect(screen.queryByText(/what you typed is still in the box/)).toBeNull()
})
it('still claims the typed words were kept for a failed turn', async () => {
await renderWith(
<FailureNotice failure={classifyError('boom')} partial={null}
onRetry={vi.fn()} onDismiss={vi.fn()} />,
)
expect(screen.getByText(/what you typed is still in the box/)).toBeInTheDocument()
})
it('labels partial prose as not kept rather than showing it as story', async () => {
await renderWith(
<FailureNotice failure={classifyError('boom')} partial="The tavern door swung"
@@ -79,10 +79,12 @@ export function CampaignSettingsPanel({ adventure, setAdventure, onError, moment
const exportCampaign = async () => {
try {
const bundle = await api.exportAdventure(adventure.id)
const { bundle, warning } = await api.exportAdventure(adventure.id)
const safe = (adventure.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
downloadJSON(bundle, `${safe}.json`)
toast('Campaign exported.')
// v1.1 WP-D: see Campaigns.jsx. The export is delivered; the warning says
// this version could not import the file back.
toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
} catch (err) {
onError(classifyError(err.message).detail)
}
@@ -47,6 +47,63 @@ function Section({ title, count, children, open = false, testId }) {
)
}
/* v1.1 WP-A1: what the server said it read, set against what was sent.
*
* Only a turn that was actually sent has this — the next-turn view has not
* been sent yet. The two conditions that mean something went wrong are shown
* as alerts, because the failure they describe is otherwise silent: Ollama
* answers 200 whether or not it cut the front of the prompt off. */
const ACCOUNTING = {
fits: {
title: 'The server read the whole prompt',
body: 'Its count stayed inside the room kept for the reply and the safety margin.',
},
exceeded: {
title: 'The prompt was larger than the server allowed for',
body: 'The server counted more tokens than the safety margin covers, so the '
+ 'reply may have been cut short. The turn is kept.',
alert: true,
},
truncation_suspected: {
title: 'The server may have cut the start of the prompt',
body: 'It read far fewer tokens than were sent, which is what happens when a '
+ 'prompt is larger than the window the model was loaded with. The '
+ 'narrator’s rules and the campaign canon are at the start. The turn is kept.',
alert: true,
},
unknown: {
title: 'The server did not say how much it read',
body: 'Nothing here can confirm whether the whole prompt was used.',
},
}
function AccountingReport({ accounting }) {
if (!accounting) return null
const copy = ACCOUNTING[accounting.status] || ACCOUNTING.unknown
return (
<div
className={copy.alert ? 'notice error' : 'ctx-accounting'}
role={copy.alert ? 'alert' : undefined}
data-testid="ctx-accounting"
data-status={accounting.status}
>
<strong>{copy.title}</strong>
<p>{copy.body}</p>
{accounting.server_prompt_tokens != null && (
<p className="notice-detail">
Sent {accounting.estimate?.toLocaleString()} by this app’s count; the
server read {accounting.server_prompt_tokens.toLocaleString()}.
{accounting.observed_margin != null
&& ` ${accounting.observed_margin.toLocaleString()} tokens were left beside the reply.`}
</p>
)}
{accounting.window_verified === false && (
<p className="notice-detail">The model’s window was not verified for this turn.</p>
)}
</div>
)
}
function TokenBar({ sections, total }) {
if (!sections.length || total <= 0) return null
return (
@@ -122,6 +179,12 @@ export function ContextPanel({
<span>{tokens.output_reserve.toLocaleString()}</span>
</div>
)}
{tokens.safety_reserve > 0 && (
<div className="ctx-token-line dim">
<span>Kept free as a safety margin</span>
<span>{tokens.safety_reserve.toLocaleString()}</span>
</div>
)}
<div className="ctx-token-line dim">
<span>What the model can hold</span>
<span>{tokens.budget.toLocaleString()}</span>
@@ -135,6 +198,8 @@ export function ContextPanel({
)}
</div>
<AccountingReport accounting={report.accounting} />
{failing.length > 0 && (
<div className="notice error" role="alert" data-testid="ctx-derived-failing">
<strong>Background work is failing</strong>
@@ -220,6 +220,50 @@ describe('context inspector (§23, §55)', () => {
expect(tokens).toHaveTextContent('8,000')
})
it('shows the safety margin kept free beside the reply (v1.1 A1)', async () => {
api.getAdventureContext.mockResolvedValue({
...REPORT, tokens: { ...REPORT.tokens, safety_reserve: 410 },
})
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
expect(screen.getByTestId('ctx-tokens')).toHaveTextContent(/safety margin\s*410/)
})
it('says nothing about accounting for a turn that has not been sent', async () => {
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
expect(screen.queryByTestId('ctx-accounting')).toBeNull()
})
it('alerts when the server may have cut the start of a sent prompt (v1.1 A1)', async () => {
vi.spyOn(api, 'getActionContext').mockResolvedValue({
...REPORT,
accounting: {
status: 'truncation_suspected', estimate: 6316, server_prompt_tokens: 2050,
observed_margin: 1546, window_verified: true,
},
})
await renderWith(
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
const el = screen.getByTestId('ctx-accounting')
expect(el).toHaveAttribute('data-status', 'truncation_suspected')
expect(el).toHaveAttribute('role', 'alert')
expect(el).toHaveTextContent('6,316')
expect(el).toHaveTextContent('2,050')
expect(el).toHaveTextContent(/turn is kept/)
})
it('reports an unknown count plainly, without claiming the prompt fit', async () => {
vi.spyOn(api, 'getActionContext').mockResolvedValue({
...REPORT, accounting: { status: 'unknown', server_prompt_tokens: null },
})
await renderWith(
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
const el = screen.getByTestId('ctx-accounting')
expect(el).toHaveAttribute('data-status', 'unknown')
expect(el).not.toHaveAttribute('role')
expect(el).toHaveTextContent(/did not say/)
expect(el.textContent).not.toMatch(/read the whole prompt/)
})
it('shows a retrieved passage with its source, class, heading and score', async () => {
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
const row = document.querySelector('[data-chunk-id="11"]')
+167
View File
@@ -0,0 +1,167 @@
/* v1.1 WP-D: an export that says whether this version could import it back.
*
* The file is delivered either way — a campaign too large to re-import is not a
* damaged one, and refusing to write it would destroy the copy the reader was
* making. What changes is what they are told, and both reader-facing Export
* controls have to tell them: the one on the campaign card and the one in the
* campaign's own settings.
*
* The server decides. These assert that the page shows what it was given and
* keeps delivering the file, not that it re-derives the size policy.
*/
import { screen, waitFor } from '@testing-library/react'
import userEvent from '@testing-library/user-event'
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { api } from '../api'
import * as components from '../components'
import Campaigns from './Campaigns'
import { CampaignSettingsPanel } from './Play/panels/CampaignSettingsPanel'
import { mockModelStatus, renderWith } from '../test/helpers'
const BUNDLE = { format: 'ai-dnd-adventure-v3', title: 'Long Campaign', actions: [] }
const WARNING =
"This export is larger than this version's 20 MB import limit (21,230,000 bytes). "
+ 'The file was exported successfully, but this version cannot import it.'
const CAMPAIGN = {
id: 4, title: 'Long Campaign', action_count: 900,
updated_at: '2026-09-15T10:00:00', snippet: 'Rain over the harbour.',
}
const ADVENTURE = {
id: 4, title: 'Long Campaign', ai_instructions: '', narration_length: 'brief',
canon_rules: [], persona_name: 'Aldric', persona_desc: '',
}
// restoreAllMocks does not undo stubGlobal, and the fetch stub below would
// otherwise outlive its own describe block.
beforeEach(() => { vi.restoreAllMocks(); vi.unstubAllGlobals() })
function exportReturns({ warning = null } = {}) {
return vi.spyOn(api, 'exportAdventure').mockResolvedValue({
bundle: BUNDLE,
exportBytes: warning ? 21_230_000 : 12_000,
importLimitBytes: 20 * 1024 * 1024,
importable: !warning,
warning,
})
}
async function library() {
mockModelStatus(api)
vi.spyOn(api, 'listAdventures').mockResolvedValue([CAMPAIGN])
await renderWith(<Campaigns />)
await screen.findByText('Long Campaign')
}
async function settingsPanel() {
mockModelStatus(api)
await renderWith(
<CampaignSettingsPanel adventure={ADVENTURE} setAdventure={vi.fn()}
onError={vi.fn()} moments={900} />,
)
}
describe('reading what the server said', () => {
// The tests below mock api.exportAdventure, so nothing there exercises the
// header names. These do: a typo in one of them would otherwise leave the
// whole suite green and the reader silently uninformed.
function serverSends(headers) {
const body = JSON.stringify(BUNDLE)
vi.stubGlobal('fetch', vi.fn().mockResolvedValue(new Response(body, {
status: 200,
headers: { 'Content-Type': 'application/json', ...headers },
})))
}
it('reports a warning the server sent, with the sizes it named', async () => {
serverSends({
'X-Export-Bytes': '21230000',
'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
'X-Importable-By-This-Version': 'false',
'X-Export-Warning': WARNING,
})
const result = await api.exportAdventure(4)
expect(result.bundle).toEqual(BUNDLE)
expect(result.warning).toBe(WARNING)
expect(result.importable).toBe(false)
expect(result.exportBytes).toBe(21_230_000)
expect(result.importLimitBytes).toBe(20 * 1024 * 1024)
})
it('claims nothing when an older server sends no headers', async () => {
serverSends({})
const result = await api.exportAdventure(4)
expect(result.bundle).toEqual(BUNDLE)
expect(result.warning).toBeNull()
expect(result.importable).toBe(true)
expect(result.exportBytes).toBeNull()
expect(result.importLimitBytes).toBeNull()
})
it('treats an ordinary export as importable', async () => {
serverSends({
'X-Export-Bytes': '12000',
'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
'X-Importable-By-This-Version': 'true',
})
const result = await api.exportAdventure(4)
expect(result.importable).toBe(true)
expect(result.warning).toBeNull()
expect(result.exportBytes).toBe(12_000)
})
})
describe('the campaign library export', () => {
it('delivers the file and says nothing more when it can be imported back', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns()
await library()
await userEvent.click(screen.getByRole('button', { name: 'Export' }))
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
expect(download.mock.calls[0][0]).toEqual(BUNDLE)
expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
expect(screen.queryByText(/cannot import/)).toBeNull()
})
it('still delivers the file when it is too large, and says so', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns({ warning: WARNING })
await library()
await userEvent.click(screen.getByRole('button', { name: 'Export' }))
// The file is written first: the warning is about importing it back, not
// about the export having failed.
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
expect(download.mock.calls[0][0]).toEqual(BUNDLE)
const notice = await screen.findByText(/cannot import it/)
expect(notice).toBeInTheDocument()
expect(notice.textContent).toContain('20 MB')
expect(notice.textContent).toContain('exported successfully')
expect(screen.queryByText('Campaign exported.')).toBeNull()
})
})
describe('the campaign settings export', () => {
it('delivers the file and confirms it when it can be imported back', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns()
await settingsPanel()
await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
expect(screen.queryByText(/cannot import/)).toBeNull()
})
it('still delivers the file when it is too large, and says so', async () => {
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
exportReturns({ warning: WARNING })
await settingsPanel()
await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
const notice = await screen.findByText(/cannot import it/)
expect(notice.textContent).toContain('20 MB')
expect(screen.queryByText('Campaign exported.')).toBeNull()
})
})
+26 -1
View File
@@ -32,6 +32,16 @@
font-variant-numeric: tabular-nums;
}
.ctx-warn { margin: 8px 0 0; color: var(--danger); font-size: 0.76rem; }
/* v1.1 A1: what the server said it read. The two failure states render as a
`.notice.error` alert instead; this is the quiet form for `fits` and
`unknown`. */
.ctx-accounting {
padding: 9px 13px;
border: 1px solid var(--border);
border-radius: 7px;
color: var(--text-dim);
}
.ctx-accounting p { margin: 4px 0 0; }
.token-bar {
display: flex;
@@ -49,7 +59,22 @@
.slice-4 { background: #6f9e8c; }
.slice-5 { background: #c48a6a; }
.slice-6 { background: #7c86b8; }
.slice-7 { background: var(--border-bright); }
/* v1.1 WP-E: pinned to the literal value this slice already rendered, instead
of borrowing --border-bright. A chart fill and a control edge have different
jobs: WP-E raised --border-bright to clear WCAG 1.4.11 (3:1) for control
boundaries, and that dragged this slice to #7a7aaa, an OKLab dE of 0.035
from .slice-6 (#7c86b8) — two neighbouring slices the same colour. The
other slices sit 0.100-0.119 from their nearest neighbour; at #3d3d55 this
one sits 0.251, the most separated in the set.
1.4.11's 3:1 does not govern this: it is a proportional fill in a labelled
breakdown, not the boundary of a control, and what it needs is to be
distinguishable from the seven slices beside it. Reassigning it to a freer
hue was considered and rejected — inside the palette's own chroma and
lightness bands the only hues that beat 0.100 are pinks near 14 degrees,
which is --danger's territory and would paint an ordinary prompt section in
the colour this application reserves for failure. */
.slice-7 { background: #3d3d55; }
.ctx-block {
border: 1px solid var(--border);
+9 -2
View File
@@ -3,8 +3,15 @@
--bg-panel: #131320;
--bg-panel-glass: rgba(19, 19, 32, 0.82);
--bg-input: #1a1a2a;
--border: #2b2b3d;
--border-bright: #3d3d55;
/* v1.1 WP-E: control boundaries carry WCAG 1.4.11 (3:1 non-text contrast) on
their own, rather than leaning on the control's text label. The floor is
measured against --bg-input (#1a1a2a), not --bg-panel: inputs and buttons
are drawn on --bg-input (styles/forms.css), and it is the lightest of the
three backgrounds a border sits on, so it is the worst case.
--border 3.21:1 and --border-bright 4.24:1 there; higher on the others
(backend/tools/contrast_audit.py, which now fails below 3:1). */
--border: #676792;
--border-bright: #7a7aaa;
--text: #e2ddd0;
--text-dim: #918c7d;
--accent: #d4a94e;
+9 -1
View File
@@ -1,6 +1,6 @@
# Adventure Storyteller — Production Build Milestones
**Status:** In implementation. M1-M8 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6, M7 and M8: 2026-09-06, each of the last four after an independent review and a corrective pass). **M9 — Export, Backup, Recovery, and Migration Hardening — is next, and has not been started.**
**Status:** **Complete. M1-M11 are closed, and v1.0.0 was released on 2026-09-14** (signed tag `v1.0.0` on signed commit `432f041`, which `main` also points at). This document is the v1 milestone history and is not extended. Post-v1 work is planned as v1.1 work packages in `V1.1-PLAN.md`, not as further milestones.
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
## 1. Purpose
@@ -1466,6 +1466,10 @@ applicable (report §T). The acceptance takes effect with the owner's signed
closeout commit. **No release tag exists**: tagging `v1.0.0` is a separate
decision, and the tag must point at that signed commit.
*Post-release note (2026-09-14):* both events have since happened. The closeout
commit was signed as `432f041`, `main` was fast-forwarded to it, and the signed
tag `v1.0.0` points at it. The paragraph above is kept as it stood at closeout.
**M1-M11 are all complete. There is no M12.** Post-v1 work is backlog, listed
below under *Post-v1 backlog*, and none of it is an unfinished v1 milestone.
@@ -1550,6 +1554,10 @@ Work recorded for after v1. None of it is a v1 requirement or an unfinished v1
milestone, and none of it has a brief. Each item needs one before work begins.
Sources are the M11 report's §P and §S.6.
**Triaged on 2026-09-14 in `V1.1-PLAN.md`**, which orders it into v1.1 work
packages, a v1.2 list and future work. The list below is kept as recorded at the
M11 closeout; `V1.1-PLAN.md` is where its disposition now lives.
- **Context-window safety margin.** The largest prompts leave 23-42 real tokens,
and Ollama cuts an over-window prompt with no error. Consider a deliberate
reserve, or counting with the narrator's own tokenizer.
+138
View File
@@ -451,6 +451,50 @@ The memory system should favor:
- uniqueness,
- continuity relevance.
### As implemented (v1.1 WP-B.2): what the summariser is shown
A memory is written from one block of `MEMORY_INTERVAL` (6) story actions. The
summariser is given the cast brief, then the block, inside a budget of
`MEMORY_EXCERPT_TOKENS` (2,000).
- **A block that fits** is sent whole, exactly as v1.0.0 sent it.
- **A longer block** was cut to its last 2,000 tokens in v1.0.0, so a fact near
its start never reached the summariser (WP-B.1). It is now sent as its
opening and its end, in order, with a visible marker between them
(`EXCERPT_OMISSION_MARKER`, "[… the middle of this stretch of story is left
out here …]"). The marker and its blank lines are paid for first, and the rest
is halved, the odd token going to the end: 992 + 993 + 15 = 2,000 tokens. The
rejoined text is measured, and the opening gives up tokens until the whole is
within budget.
- **The marker is never stored.** `summarize_block` removes it from anything the
model repeats back.
- **Limit.** A fact in the middle of a very long block is still left out. The
input stays bounded; it is not a summary of everything.
Existing memories are not rewritten. `tools/rewrite_memories.py`, which is
opt-in, uses the same function.
### As implemented (v1.1): what a memory can be relied on to keep
The memory prompt is v1.0.0's, unchanged. v1.1 changed what the summariser is
shown (above), not what it is told.
**Known limitation, accepted for v1.1.** The application now delivers the whole
relevant block to the summariser, and keeps, ranks and injects the memory it
writes (§18, §20, §21). But the reference summariser, `qwen2.5:3b-instruct-16k`,
can still be given a block that states a distinctive fact and write a memory
that:
- omits the fact, or the specific objects in it;
- attributes it to the wrong character;
- prefers the generic narration that follows it.
So **independent recovery from memory is proven for the application's
mechanisms, and is not guaranteed with the reference model.** A bounded prompt
change aimed at this was tried and rejected
(`reports/v1.1/V1.1-WP-B2-REPORT.md` §T). A stronger dedicated summariser,
structured fact extraction, or separate factual and narrative memory are
future options, and none is implemented.
## 16. Memory Retrieval
Retrieval should be local.
@@ -513,6 +557,27 @@ abbey crypt
prior discoveries
```
### As implemented (v1.1 WP-B.2)
v1.0.0 embedded the newest four actions cut to 600 tokens, so the player's
one-line question arrived after three turns of narration and barely moved the
embedding (WP-B.1: a direct question's cosine fell from 0.708 alone to 0.241 in
that query). The query is now two short texts, embedded in one call:
| Component | What it is | Bound |
| --- | --- | --- |
| **input** | the player's own action this turn (`do`, `say` or `story`) | last 200 tokens |
| **context** | the scene from the authoritative state (its summary, the location's name, the names of who is present), then the end of the newest narration | 60 + 120 tokens |
- A continue turn, and the Insights dry run, have no input; the context alone is
searched.
- A retry searches with the input being retried; the discarded attempt is not in
the context.
- The full entity list, threads and older narration are deliberately left out,
so a long scene or a large cast cannot outweigh the question by length.
- What was searched for is recorded per turn in the context snapshot
(`memories.query`: input, context, input terms, weights).
## 19. Memory Retrieval Filtering
Before ranking memories, filter by:
@@ -553,6 +618,41 @@ future work rather than something M6 delivered.
What M6 does implement, because similarity alone proved insufficient, is
redundancy suppression before the final selection: see §22.
### As implemented (v1.1 WP-B.2)
One transparent lexical term is added to similarity. For each eligible memory:
```text
semantic_score = 0.6 * cos(input, memory) + 0.4 * cos(context, memory)
(either cosine alone when the other text is empty)
lexical_score = sum of w(t) over the input's terms the memory holds
/ sum of w(t) over all the input's terms in [0, 1]
w(t) = ln((N + 1) / (df(t) + 1)) N eligible memories, df holding t
final_score = semantic_score + 0.15 * lexical_score
```
- **Terms** are the knowledge path's tokenizer and stop list (`knowledge.fts`),
with possessives dropped and a plural `s` folded. No stemmer, no dependency.
- **Rarity** is computed per turn over the eligible candidates only. There is no
index and no stored field. A word every candidate holds (a protagonist's
name) weighs 0; a word the question shares with one memory weighs most.
- **Only the player's input** is matched lexically, never the context.
- **The weight** was chosen by sweep over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5:
0.15 is the smallest at which the lexical term alone lifts the planting-era
memory into `memory_top_k` against the v1.0.0 query, while no rare-word
negative control lets an unrelated memory pass a semantically relevant one.
A memory can gain at most 0.15 from wording, so it cannot pass one more than
0.15 ahead of it in meaning.
- **Pins** are unchanged: always used, counted toward `memory_top_k`.
- **Ties** on the final score are broken by memory id.
- **Provenance.** Each used memory records `semantic_score`, `lexical_score` and
`final_score`; `similarity` keeps its v1.0.0 meaning, the semantic score, so
the inspector's "closeness" is unchanged.
Importance, recency, entity overlap and story-thread overlap remain
unimplemented. Redundancy suppression (§22) is unchanged and runs over the final
order.
## 21. Memory Budget
Retrieved memories should have a bounded token budget.
@@ -564,6 +664,44 @@ Recommended behavior:
- include only the highest-value items that fit,
- preserve source IDs for inspection.
### As implemented (v1.1 WP-B.2): selection and the bank's capacity
**Selection** is unchanged in shape: every eligible, embedded memory on the
active lineage is scored (§20), pinned memories are taken first, the rest fill
`memory_top_k` (default 5) best first, skipping repeats (§22). The Memories
section is priced into the protected context like any other live section, so
its budget is unchanged (F03). There is still no relevance floor: a full
`memory_top_k` is used whenever the bank holds that many.
**Capacity** (`memory_bank_capacity`, default 80) is unchanged. Eviction still
runs after each post-turn pass, over the whole adventure rather than one
lineage, and marks rows `forgotten` rather than deleting them. What changed is
the order (`memorybank.eviction_order`), because least-recently-used alone
discarded the only memory of an early stretch first (WP-B.1):
- **Coverage signal.** Memories with a source range say which stretch they
describe. Each is judged by the hole its removal would leave between the end
of the memory before it and the start of the memory after it. The smallest
hole goes first, so the bank thins where it is densest. A memory whose start
another memory shares leaves no hole.
- **Boundaries.** The earliest and the latest memory by position are not
coverage candidates: they are the only records of the opening and of the most
recent stretch. This also keeps a memory written this turn from being evicted
by the pass that wrote it (the frozen bank).
- **Recency signal.** Among equal holes, the least recently used goes first
(`coalesce(last_used_at, created_at)`), then the less used, then the lower id.
- **Pinned rows** are never taken, and count as coverage.
- **Fallback.** Memories with no range (typed by the player, or migrated) and
boundaries are taken least recently used first, as in v1.0.0, once no coverage
candidate remains. The bank stays bounded either way; only an all-pinned bank
may exceed capacity.
Measured on banks where nothing is ever retrieved, the kept bank starts at the
opening and its largest uncovered stretch stays within about 1.3 times the
average spacing (story length / capacity). The v1.0.0 order kept only the newest
stretch. The rule reads no text and no vectors, and lineage eligibility is
unaffected: it decides only `forgotten`.
## 22. Duplicate Suppression
Do not include the same fact repeatedly through:
@@ -69,6 +69,18 @@ out of tokens partway through it. All of that is removed before the prose is
stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4
has the rules.
**Implementation note (v1.1 WP-A2).** A small model also copies the protocol's
*instructions*: the vocabulary written as calls, the length hint, and the scene
line. The fix is on both sides:
- **Prompt:** the vocabulary is shown in the wire format, and the fixed example
uses genre-neutral placeholders.
- **Extractor:** it recognises those echoes only by strings and names the
application owns.
No event type, field, validation rule or proposal record changed.
`TECHNICAL-DESIGN.md` §15.4 lists the four rules.
### Where it lives
- `adventures.narrative_state` — the current authoritative document. This is
+50 -27
View File
@@ -2,11 +2,22 @@
**This file is the index. Start here.**
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
**milestones M1 through M11 complete**. **M11 was accepted at its closeout
(2026-09-14), and the v1 release gate passed** on the release-candidate tree.
There is no further planned milestone. The signed closeout commit and any
`v1.0.0` tag are separate events and the repository owner's to perform. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
**Current state:** **v1.0.0 remains the released version. Every v1.1 work
package is implemented, accepted and signed** — WP-A1/A2 (`d63804f`), WP-B.1
(`beb17ad`), WP-B.2 (`0c1ba83`), WP-C (`59b5ebc`), WP-D and WP-E (`87a4032`).
**Integrated release validation passed on candidate `87a4032`**
(`reports/v1.1/V1.1-RELEASE-REPORT.md`): the v1 contract holds (81 PASS, H09 NOT
APPLICABLE), every suite and build passes, and the browser, offline, long-run,
identity, recovery, upgrade and smoke gates are clean. WP-B ships with a
documented reference-model memory limitation. **No release commit, no `main`
update and no `v1.1.0` tag exist yet** — those are the owner's separate events.
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
through M11 complete and closed**. M11 was accepted at its closeout
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
owner then signed the closeout commit (`432f041`), fast-forwarded `main` to it,
and created and pushed the signed tag `v1.0.0` pointing at it. v1.1 development
is on the `v1.1-development` branch, from that commit, and is planned in
`V1.1-PLAN.md` as work packages, not milestones. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
independent review found a real defect and a corrective pass fixed it.
@@ -46,13 +57,18 @@ The closeout repeated the browser, offline and identity runs on the exact
release-candidate tree (`3652dc6`), and recorded the acceptance (report §S and
§T). **The v1 release gate passed.**
**Not yet happened**, and each a separate event:
- the owner signing the closeout commit;
- merging to `main`, which still points at M10;
- any `v1.0.0` tag.
**The release, 2026-09-14.** The three events the closeout left to the owner
have all happened:
- the closeout commit is signed, as `432f041`;
- `main` was fast-forwarded to `432f041`;
- the signed tag `v1.0.0` points at `432f041`, and is pushed.
**There is no further planned milestone and no M12.** Post-v1 work is backlog
(`BUILD-MILESTONES.md`, *Post-v1 backlog*).
The M11 report's §T still lists the last two as not done. That table records the
state at closeout and is deliberately left as written.
**There is no M12.** Post-v1 work is organised as **v1.1 work packages** in
`V1.1-PLAN.md`, which triages the *Post-v1 backlog* recorded in
`BUILD-MILESTONES.md`.
**Package version:** see `VERSION.md`, which records what each revision changed
and why.
@@ -122,7 +138,8 @@ Two standing qualifications:
| `SECURITY-THREAT-MODEL.md` | The trust boundary, and the inference endpoint policy as implemented. |
| `MEDIA-EXTENSION-CONTRACT.md` | The contract future image/video/audio/TTS/STT work must fit. |
| `BROWSER-UX-SPEC.md` | The browser surface, and what is deliberately not in it. |
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. |
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. Closed history; not extended. |
| `V1.1-PLAN.md` | **Post-v1 work.** The triaged backlog, the ordered v1.1 work packages with acceptance criteria, the v1.1 release criteria, and what is deferred. |
| `V1-ACCEPTANCE-TESTS.md` | The pass/fail contract v1 is measured against. |
| `TEST-CAMPAIGN-FIXTURE.md` | The standard campaign the acceptance tests are run on. |
| `VERSION.md` | Package revision history: what each closeout changed. |
@@ -133,7 +150,8 @@ Two standing qualifications:
1. This file.
2. `SPECIFICATION.md`
3. `TECHNICAL-DESIGN.md`
4. `BUILD-MILESTONES.md` — find the milestone you are being asked to do.
4. `V1.1-PLAN.md` — find the work package you are being asked to do.
`BUILD-MILESTONES.md` is the v1 history behind it.
5. `STORY-BRANCH-SEMANTICS.md`
6. `DATA-MODEL.md`
7. `CONTEXT-AND-MEMORY.md`
@@ -185,7 +203,9 @@ that is the one the next milestone's planning has to consult:
It stays in `reports/` after acceptance. The rotation moves a report to the
archive when the next milestone's report is written, and after M11 there is no
next milestone. Nothing here calls for moving it, so no new convention was
invented to do so.
invented to do so. `V1.1-PLAN.md` asks each v1.1 work package for a report of
its own; when the first is written, it joins this one in `reports/` and the
rotation is decided then.
M9's and M10's reports both moved to `archive/milestone-reports/` when this one
was written. M9's had been kept here past its turn because M9 was unaccepted;
@@ -337,27 +357,30 @@ Milestone M11 COMPLETE / ACCEPTED (2026-09-14)
|
v
v1 release gate PASSED (2026-09-14)
signed release commit and v1.0.0 tag:
the repository owner's, not yet done
|
v
Post-v1 backlog only; no milestone planned
v1.0.0 RELEASED (2026-09-14)
signed tag on signed commit 432f041;
main points at the same commit
|
v
v1.1 PLANNING (from 2026-09-14)
branch v1.1-development; V1.1-PLAN.md
work packages, not milestones
```
## Stop Rule
**One milestone at a time. Do not begin a milestone before its brief exists.**
**One work package at a time. Do not begin a work package before its brief
exists.**
**Every planned milestone is complete, and M11 is accepted.** The v1 release gate
passed on 2026-09-14 (M11 report §T).
**Every v1 milestone is complete, M11 is accepted, and v1.0.0 is released**
(2026-09-14, tag `v1.0.0` on `432f041`).
The next actions are the owner's:
1. sign the closeout commit;
2. decide on the `v1.0.0` tag, pointed at that signed commit.
Neither is a milestone. **Do not begin post-v1 work as though it were a v1
milestone.** Anything after v1 starts from the *Post-v1 backlog* in
`BUILD-MILESTONES.md`, with a brief of its own.
**Do not begin post-v1 work as though it were a v1 milestone, and do not create
M12.** v1.1 work is the ordered work packages in `V1.1-PLAN.md`. Each starts
only from a coding brief of its own, and stops at its own boundary for review.
The plan names the brief to write first.
All three questions the M8 debt raised against M9 are settled and recorded:
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
+116
View File
@@ -1230,6 +1230,75 @@ whether there is a number to cap to at all, and that is what the builder and the
declaration everywhere it appears, and the connection test says plainly that
nothing has checked it against the server.
**As implemented (v1.1 WP-A1): a safety reserve, and the server's own count.**
A ceiling in the application's tokens is not a ceiling in the narrator's. The
builder counts with `cl100k_base`, and the v1 evidence left the largest prompts
23-42 real tokens from the edge of a 16,384 window. Past the edge Ollama does not
refuse. Measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt returned
200 with `prompt_tokens` 2,050.
- **The reserve.** `contextwindow.safety_reserve(budget)` is
`max(256, ceil(5% of the effective budget))`, computed with integer rounding
up: 256 at 4,096, 410 at 8,192, 820 at 16,384. It is taken before any history
is chosen. The effective budget is the verified or declared window when there
is one, and the configured budget otherwise. It is fixed and documented, is not
a setting, and is not calibrated per model.
- **What replaced the 64-token margin.** M6's `OUTPUT_SAFETY_MARGIN` absorbed two
unrelated things.
- The application's own text added after pricing: separators between
sections, and the chat hint the provider appends to every request. This is
now priced exactly as `transport`.
- Tokenizer drift. This is now the reserve.
The reply allocation is exactly `max_output_tokens`. Protected context is
`sections + transport + reply + reserve`, and `ContextOverflow` is raised
before the model call when that does not fit.
- **The server's count.** Streaming requests set `stream_options.include_usage`.
Without it Ollama sends no usage, and none of the 514 AI turns in the v1
evidence has one. After the reply, `contextwindow.classify_usage` compares the
server's `prompt_tokens` with `tokens.estimate`: the assembled text plus what
the provider adds.
- **Accounting states.** The turn's snapshot records `accounting`, whose status
is one of the following, checked in this order:
| Status | Meaning |
| --- | --- |
| `unknown` | No positive integer count was reported. It is never read as `fits`. |
| `truncation_suspected` | The server read fewer tokens than the estimate by more than the reserve. |
| `exceeded` | The server's count plus the reply allocation is over the budget. |
| `fits` | Otherwise. |
The record also carries the server's count, the difference, the reserve and
the observed margin (`budget - reply - server count`).
- **Surfacing.** The record is returned on the turn's `done` event, logged as a
warning when it is `exceeded` or `truncation_suspected`, and shown in the
context inspector, where those two statuses are an alert.
- **The turn is kept.** A discrepancy found after the reply is recorded, never
enforced. The narration has already streamed to the reader, and the accepted
turn is not discarded.
- **Accounting belongs to one attempt.** It sits in `attempts.ATTEMPT_KEYS`
beside `usage`. When a retry or a take selection moves the shared prompt,
each take keeps the accounting for its own call.
- **A cold model is loaded before its turn is built (v1.1 A1 corrective).** A
model that is not resident cannot report its window. The A1 evidence caught
exactly that: 13,875 tokens were sent to a server that read 2,050.
`contextwindow.ensure_window` works like this:
- it probes;
- if the window is unverified and the server answered, it makes one bounded
`POST /api/generate` naming only the model, with no prompt. Ollama loads
the model and generates nothing ("done_reason": "load");
- it probes again, bypassing the cache;
- the turn is built to whatever that second probe says.
A failed load, or a window still unverified afterwards, changes nothing: the
configured budget stands and the accounting still catches a cut. The request
goes to the configured endpoint only, under the same policy and TLS path. It
writes nothing, and it is recorded as `window.preflight` in the turn's
snapshot. The context dry run never loads a model.
No schema change: the accounting lives in the snapshot JSON, and a turn from
v1.0.0 simply has none. The bundle format is unchanged, for the same reason.
### 15.3 The history window moves in blocks (post-M11)
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
@@ -1300,6 +1369,53 @@ headings. A lone heading followed by prose stays, and so do JSON a character
typed and a fact restated inside a sentence. That last case is how a narrator can
still carry authoritative state into its prose (M11 report §G.4 and §P).
**As implemented (v1.1 WP-A2): the source and the sink together.** The M11
closeout's identity run stored four shapes the extractor left. All of them were
application text. v1.1 changes both the prompt that taught them and the
extractor that missed them. Every new removal is anchored to something the
application owns, never to what prose looks like.
- **The prompt.**
- `events.vocabulary_for_prompt` shows each event as the object the model
must send (`{"type": "set_possession", "item": "<key>", "owner": "<key>"}`),
not as `set_possession(item, owner)`. The call notation was never the wire
format, and the narrator copied it.
- `EMIT_RULE`'s example uses the placeholders `character-1`, `item-1` and
`location-1`, not the fantasy fixture's `mara`, `silver-key`, `old-abbey` and
`aldric`. The narrator had proposed `silver-key` in an office meeting.
- The length hint's opening and closing words are named constants shared by
the builder and the extractor.
- **The extractor**, rules R1-R4 (`RULE_*` in `narrative/extract.py`):
- **R1:** a whole line that begins with a call to an event in `events.SPECS`,
optionally `>`-quoted. Not inside a fenced code block, not mid-sentence, and
not for a call-shaped name the vocabulary lacks.
- **R2:** a trailing bracket that opens `Hard limit:` and carries the hint's
own wording ("append the state block", or "turn must not exceed *N* words").
- **R3:** the renderer's scene line left as the reply's last line. It is
removed when it ends in the renderer's `(at <location>)`, or when protocol
was already cut from the same reply.
- **R4:** a ```` ```json ```` or bare ```` ``` ```` opener left as the last
line with nothing after it, counted as an opener rather than a closer.
- **Proven on real narration.** Every stored real reply in the v1 evidence was
replayed through the v1.0.0 and v1.1 extractors (`tools/v11_replay_extractor.py`).
Every changed line is attributed to one of the four rules, and a person reviewed
every change. The results are in the WP-A1/A2 report.
- **Deliberately still left:** a fact restated inside a sentence; model-invented
headings; and a bracket that starts `Hard limit:` but carries none of the
application's wording.
- **R5, the echoed instruction tail (v1.1 A2 corrective).** A v1.1 identity turn
ended in a reworded continue hint, "[… Continue the story here, directly.
Output only story text.]". Because nothing recognised it, nothing above it was
ever trailing, and the reminder, a reworded length hint and a scene line all
stayed. R5 makes two changes:
- **The continue hint is recognised by its own sentence.** A trailing bracket
containing "Output only story text" is an echoed instruction.
`CONTINUE_HINT_PHRASE` is pinned by a test to `CHAT_CONTINUE_HINT`.
- **A bracket opening with the length hint's own `Hard limit:` is removed only
directly above an echoed instruction already cut from the same reply's end.**
An in-world "[Hard limit: forty days]" stays when it is the last line, and
when a state block follows it. Any other bracket above an echo stays.
## 16. Database Direction
SQLite remains the selected v1 authoritative store.
+940
View File
@@ -0,0 +1,940 @@
# Adventure Storyteller — v1.1 Plan
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
signed v1.0.0 release commit `432f041`.
**WP-A1 and WP-A2** are committed and signed as `d63804f`
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention
diagnostic, is committed and signed as `beb17ad`
(`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and
staged for the owner's signed commit. It is **accepted with a documented
real-model limitation** (owner decision, 2026-09-15):
- B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship;
- deterministic independent-memory recovery passes;
- the one isolation-valid reference-model run failed at memory creation;
- the B2.4 prompt experiment did not fix that and was reverted.
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
**WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`).
**WP-D** (recovery honesty) and **WP-E**
(control-boundary contrast) are committed and signed as `87a4032`, each with its
own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`.
**Integrated release validation has since run on candidate `87a4032` and
passed** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): the 82 REQUIRED v1 tests hold
(81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a
`--no-cache` image whose SPA is file-for-file identical to the local build,
offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window
run passing M01-M04 with every turn `fits` and an A2 leak count of 0, identity
0 signals / 0 protocol shapes, recovery 16/16, a real v1.0.0 upgrade comparing
identical on all 15 fields with both bundle directions importing, and a release
smoke of 15/15.
WP-B's reference-model memory limitation is carried as an accepted residual, as
are the mid-reply instruction echo and the doubled full stop; K1 is classified as
v1.2 backlog. **No `v1.1.0` tag exists, `main` is unchanged, and no release
commit has been made** — those three remain the owner's separate events.
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
specification and acceptance contract are unchanged: every package below
improves how an existing requirement is met. No package adds a requirement.
---
## 1. Terminology
- **Work package (WP).** One independently reviewable unit of v1.1 work, with
its own coding brief, its own report, and its own review. Work packages are
lettered. **They are not milestones, and there is no M12.**
- **Brief.** The coding prompt for one work package. It holds only the
objective, the source documents, the scope and non-scope, the acceptance
criteria, and the stop condition. No work package begins before its brief
exists.
- **Reference narrator.** `qwen2.5:3b-instruct`, with and without `num_ctx`
16,384 baked in, and `nomic-embed-text` for embeddings. These are the models
the v1 evidence used (M11 report §E.1). v1.1 real-model evidence uses the
same models so that its numbers compare with v1's.
- **Reference hosts.** The CPU reference host serves Ollama over HTTPS with a
private CA, which is the A06 evidence path. The GPU inference host serves
plain HTTP on the LAN and is used for long runs (M11 report §E.1).
- **Planning package version** (`VERSION.md`, v4.x) and **product version**
(v1.0.0, v1.1.0) are separate numbers. `SECURITY-THREAT-MODEL.md`'s own
"Status: v1.1" line is that document's revision label from Phase 0B. It does
not refer to this release.
## 2. Baseline
| | |
| --- | --- |
| Release | **v1.0.0**, 2026-09-14 |
| Release commit | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, signed by the owner |
| Tag | `v1.0.0`, a signed annotated tag on that commit, pushed |
| `main` | `432f041`, the same commit |
| v1.1 branch | `v1.1-development`, created at `432f041` |
| v1 contract | 82 REQUIRED FOR V1 tests: 81 PASS and H09 NOT APPLICABLE (M11 report §T) |
| Schema | migrations up to 94 (`LATEST_VERSION` 94) |
| Bundle format | `ai-dnd-adventure-v3` |
| Provenance | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
## 3. v1.1 goals
1. **No silent loss of what the narrator is given.** A prompt must not reach the
server's window edge by an accident of tokenizer arithmetic. A turn the
server truncated must be reported.
2. **No protocol in the story.** The shapes of the application's own protocol
that a narrator copies must not be stored as narration. The fixed
instructions must not hand every campaign one genre's nouns to copy. Story
prose must not be removed in the process.
3. **Memory that remembers on its own.** An old fact must be recoverable from
the memory bank itself, with provenance, when nothing else carries it. This
must not weaken lineage safety.
4. **Release evidence without API-only gaps.** Every reader-facing history,
state, length, failure and export behaviour must be driven in a real browser.
5. **Recovery that says when it cannot recover.** Backups get a full integrity
check, and an export that import would refuse must say so.
6. **Control boundaries that meet WCAG 1.4.11.**
And throughout: **every v1.0.0 campaign opens unchanged in v1.1.**
## 4. Non-goals
- New product features: media generation, TTS, STT, story search, whole-transcript
copy, a discarded-history recovery screen, a restore button, tablet redesign.
- Changes to the history, branch, head or Save Point architecture (ADR 005,
ADR 012).
- Changes to the authoritative state model, its event vocabulary or its
validator (ADR 010, ADR 013). Changes to the prompt *wording* that describes
the vocabulary are in scope, in WP-A2.
- Changes to the knowledge authority classes or their retrieval.
- A new bundle format version.
- Any new runtime dependency, any network access beyond the configured
inference endpoint, and any first-use download, including per-model
tokenizers.
- Supporting inference servers other than Ollama, beyond what
`context_window_override` already allows.
- Re-writing stored narration or memories of existing campaigns automatically.
- Redesigning identity or state handling on the strength of post-M8 finding D,
which has not been reproduced.
## 5. Rules for every work package
The v1 milestone rules (`BUILD-MILESTONES.md` §2) apply to every work package,
and so do the following:
1. **Compatibility default.** Existing v1.0.0 campaign databases open
unchanged. Any migration is forward-only and additive, and is tested by
opening a real v1.0.0 database. Fresh-install and upgraded schemas must still
compare identical (`test_m11_migration.py`). v1.0.0 bundles still import.
2. **The bundle stays `ai-dnd-adventure-v3`** unless a brief makes the case under
M9's semantic test and the owner agrees. Adding a key inside stored evidence
is not a format change when an absent key unambiguously means "not
recorded".
3. **The v1 contract is the regression floor.** No REQUIRED FOR V1 test is
retired, relaxed or reclassified.
4. **Evidence discipline** (M11 report §O). Any product change made after a
black-box run was taken invalidates that run for release purposes. Harness
defects and product defects are reported separately. A check that cannot fail
is not evidence.
5. **Each package has a report** (`planning/reports/V1.1-WP-<id>-REPORT.md`).
It records what was built, the acceptance results with measurements, what
was not verified, and the as-implemented documentation changes. The report
rotation for `reports/` is decided when the first one is written
(`planning/README.md`).
6. **Real-model runs longer than a few minutes on the GPU host** require the
power, link and kernel logging in `DEVELOPMENT.md`, started before the run.
7. **No real hostnames, addresses or people's names** in committed files,
fixtures or reports. Evidence stays under `$HOME`, never `/tmp`.
8. **Commits and tags are the owner's.** A package ends with its changes staged
for a signed commit.
---
## 6. Backlog triage
Severity reflects user risk: **High** means story corruption or lost continuity
that the reader is not told about. **Medium** means silent degradation or a
real verification gap. **Low** means visible, rare or cosmetic.
### 6.1 The recorded post-v1 backlog
| # | Item | Source | Current evidence | Severity | Disposition |
| --- | --- | --- | --- | --- | --- |
| 1 | Context-window safety margin | M11 §P risk 3, §N; post-v1 backlog | The largest prompts left 23, 34 and 42 real tokens of headroom in the GPU runs. The application's `cl100k_base` count ran 16 tokens below the narrator's on every prompt measured. The only slack is `OUTPUT_SAFETY_MARGIN = 64` (`context/builder.py`), which is fixed and must also absorb the section separators. Ollama 0.34 cut an over-window synthetic prompt to 8,194 tokens with no error. Truncation drops the oldest tokens, which are the narrator's rules and the canon. The server's reported usage is already stored per turn (`snapshot["usage"]`, `routers/adventures/turns.py`) and read by nothing. | **High**: silent, invisible, and reachable by a model whose tokenizer diverges further | **v1.1, WP-A1** |
| 2 | Narrator restating prompt and state text | §P risk 4, §G.4 | The state's fact line was restated as a sentence on 74 of 104 turns in the evidence run. Phrases such as "Scene set, continue your adventure." and lines opening "Memory:" are model-invented, and none is an application string (verified by search). The restatements sit inside story sentences. | **Medium**: narration quality, and it is the path by which M04's fact reached memory, which masks item 5 | **v1.1, WP-A2** measures it. Reducing in-sentence restatement is **v1.2**, because removing it means judging prose |
| 3 | Protocol leakage: event-call syntax, stray headings, repeated length hints | §P risk 6, §S.6 | On 4 of 10 turns of the closeout identity run (4,096 window) and 0 of 104 in the 100-turn runs. Causes verified in code. `events.vocabulary_for_prompt` shows every event as `name(field, …)`, which is the notation copied. `_is_echoed_instruction` requires both "state block" and "events list", and the length hint's tail says only "state block". A lone trailing `Scene:` is not removed. Stored text is replayed as history (TECHNICAL-DESIGN §15.4). | **Medium-high**: silent and self-reinforcing, but rare at a full window | **v1.1, WP-A2** |
| 4 | Genre-specific example in the fixed state rule | §P risk 16; `narrative/extract.py` `EMIT_RULE` | The example names `mara`, `silver-key`, `old-abbey` and `aldric`. In the office-meeting identity run the narrator proposed giving `silver-key` to Alice, and the validator refused it. No state was corrupted. | **Medium**: every non-fantasy campaign gets fantasy nouns to copy | **v1.1, WP-A2** |
| 5 | Independent long-term-memory retention | §P risk 5, §G.4; F02 and M04 qualifications | No run showed a planting-era memory carrying the planted fact. Recovery ran from state to the narrator's restatement to memory. Code facts that bear on it, none yet proven to be the cause: a memory is at most 50 words for a 6-action block. The summariser sees the block's *last* 2,000 tokens (`truncate_to_last_tokens`), so an early fact in a long block can be cut. Eviction is least-recently-used above `memory_bank_capacity` (default 80), so a never-retrieved early memory is the first to go once a campaign passes about 480 actions; the 33-memory evidence run never reached that. Ranking is cosine similarity plus a pin (CONTEXT-AND-MEMORY §20). | **High** for long campaigns: lost continuity, seen only as the story forgetting | **v1.1, WP-B** |
| 6 | Browser automation gaps: Retry, Save Point, state correction, narration length, failed generation, export download | §S.3, §T qualifications, §P risk 11 | Proved through the API, the component suite and the 100-turn campaign, but never driven in a browser. The export control is exercised only as far as the click, because a `blob:` download does not leave the headless snap Firefox. No defect is known. | **Medium**: a verification gap on the release path | **v1.1, WP-C** |
| 7 | Practical export and bundle-size ceiling | §P risk 12, §K; M9 debt | The import body limit is 20 MB (`limits.MAX_IMPORT_BODY_BYTES`). M9's conservative ceiling is about 279 turns. The real 100-turn campaign was 2.74 MB, about 13 kB per action, or roughly 1,600 actions. A larger campaign exports and is then refused on import, and nothing says so at export. | **Medium** impact, low frequency | **v1.1, WP-D**: warn at export. Raising the limit or a streaming import is **v1.2** |
| 8 | Identity-confusion diagnostic follow-up | §P risk 8, §S.5 | Not reproduced. The diagnostic cannot see a stale scene, has never run with memory on, and has never run at a 16,384 window. It found no product deficiency. | **Low**: no occurrence since the original, whose campaign is gone | **No v1.1 package.** The diagnostic is re-run as a WP-A2 regression and in the v1.1 release gate, with memory on. A stale-scene detector is **v1.2**. The next real occurrence is classified with `tools/m11_identity.py` |
| 9 | WCAG 1.4.11 control-boundary contrast | §P risk 10, §L | Measured at 1.33:1 resting and 1.75:1 on hover (`--border` and `--border-bright` against `--bg-panel`). M11's argument that the label identifies the control is weakest for text inputs, where the boundary shows where to type. | **Low-medium** | **v1.1, WP-E** |
| 10a | Backup integrity: `integrity_check` | §P risk 14; M9 | `backup.py` runs `PRAGMA quick_check`, which skips index-content verification. | **Low** | **v1.1, WP-D** |
| 10b | Scheduled backups | §P risk 14; M9 | None exist; backups are manual. They need owner policy on interval, retention, location and disk use, and must not contend with a turn's write lock (§O.7). | **Medium** impact, but a new feature | **v1.2** |
| 11 | Real media-provider adapters | §P risk 13; M10; K05 and K06 (FUTURE) | The seam has no consumer. Adding one brings dependencies, a stricter endpoint rule (`DEVELOPMENT.md`) and a new UI. | n/a: a feature, not a risk to existing stories | **Future / optional**, not v1.1 |
### 6.2 Other recorded v1 residual risks and debt
| # | Item | Source | Disposition |
| --- | --- | --- | --- |
| 12 | The state lags the narration at 4,096 with the 3B narrator (1 proposal in 10 applied) | §P risk 17, §S.5 | Model behaviour, not a defect. WP-A2's real-model runs record the proposal outcome counts. No package |
| 13 | Every realistic observation is a 3B model's | §P risk 2 | No package. v1.1 keeps the reference narrator for comparability. A comparative model recommendation stays open (`planning/README.md`, *Still open*) |
| 14 | A campaign that never chose a narration length keeps the pre-M11 hint | §P risk 15 | Deliberate. No change. WP-C drives the control |
| 15 | The GPU host dropped off the PCIe bus after the evidence run | §P risk 7, §E.1 | Operational. Rule 6 in §5 applies to every long run |
| 16 | Release timings are two hosts' | §P risk 1 | No action. No performance requirement exists or is invented |
| 17 | Cross-layer duplication: one fact in state, memory, history and a passage at once | CONTEXT-AND-MEMORY §22; open since M6 and M7 | **v1.2.** It costs tokens (item 1) and is related to item 2. WP-B may touch ranking only where its diagnostic requires |
| 18 | Memory-ranking factors beyond similarity and pin not implemented | CONTEXT-AND-MEMORY §20 | Inside WP-B's bounded fix menu, only if WP-B's diagnostic places the failure at ranking |
| 19 | M8 carried debt: no whole-transcript copy or story search (§77, §78), no discarded-history recovery screen (§63), tablet untuned, RPG world state read-only | `BUILD-MILESTONES.md` M8 and M9 | **Future backlog.** Features, not reliability |
| 20 | M10 debt: `ambience` is an empty shape; a deleted visual profile is unrecoverable in the app; profiles are API-only | M11 §D | **Future**, with media |
| 21 | A hostile host on the trusted LAN; the DNS-rebinding interval | `SECURITY-THREAT-MODEL.md` §10A | Accepted for v1. **Future.** No v1.1 package |
| 22 | A stale `chunk_id` in a restored snapshot | M9 | No action. It is a React key, not a live pointer |
| 23 | Developer-venv residue (`quickjs`, `psycopg`) | §M | No action. It does not ship |
| 24 | A Settings warning comparing the real window to the budget | M8 debt, "unowned" | **Closed by M11**: Settings reports the window and warns |
| 25 | Root `README.md` has stale facts: "1,191 backend tests" (1,421 at closeout), and the Screenshots paragraph calls screenshots "a job for the UI pass in M8" | found in this review | Documentation debt, not release status, so it is not changed in v4.1. It is fixed in WP-A1's documentation update |
---
## 7. Grouping decisions
**The suggested "narrator boundary hardening" package is split into A1 and A2.**
They share a theme but not a mechanism or a test method.
- **A1**, the context reserve, is budget arithmetic in `context/builder.py` and
`contextwindow.py`. It is proved deterministically with a scripted provider
that reports token counts.
- **A2**, the protocol echo, is the contract between the prompt's wording
(`extract.EMIT_RULE`, `events.vocabulary_for_prompt`, the length hint) and the
extractor. It is proved with a replay corpus of real narration and
real-model runs.
Bundling them would make one review carry two unrelated risk profiles.
A2's three items belong together. The example slugs and the call notation are
the *source* of the copied text, and the extractor is its *sink*. Changing one
without the other means proving the change twice. Both also alter the stored
prompt, so they invalidate the same evidence.
**B stands alone.** It is the only package whose success depends on model output
quality. Its test design, a fact that nothing but memory carries, is the hard
part. It follows A2, so that it measures the prompt v1.1 will ship.
**D is narrowed to recovery honesty**, meaning `integrity_check` and the export
warning. Both are small and deterministic, and both are recovery-path changes
with no policy question. Scheduled backups need owner decisions and have real
design risk, including write-lock contention. Raising or streaming the import
limit changes a DoS guard. Both move to v1.2.
**E stays focused** on control boundaries. The same measurement found no other
accessibility defect (M11 §L).
**F is not a package.** Nothing concrete in the product is deficient. The
diagnostic's known blind spots are a stale scene, memory off, and one window
size. The first two can be addressed by running it differently: the v1.1 gate
runs it with memory on. The stale-scene detector is new diagnostic work with no
occurrence to justify it yet, so it is v1.2.
**G is not in v1.1.** It adds no reliability or quality to existing stories, it
brings dependencies and a new endpoint surface, and its tests (K05, K06) are
FUTURE. `MEDIA-EXTENSION-CONTRACT.md` and the M10 seam remain authoritative for
whenever it is taken up.
---
## 8. Work packages
### WP-A1 — Context-window safety reserve
**Objective.** An assembled prompt plus its reply reserve stays at least a
documented safety reserve below the effective window, whatever tokenizer the
narrator uses. Where the server reports its own prompt count, a turn that
exceeded the window, or was evidently truncated, is recorded and shown, never
silent.
**Rationale.** §15.2 closed "budget larger than the window". It did not close
"our count is not the server's count". Measured headroom was 23-42 tokens
(item 1), and the fixed 64-token margin absorbs separators as well as tokenizer
drift. A narrator whose tokenizer runs a few percent heavier than
`cl100k_base` overflows a full 16,384 window. The failure deletes the canon at
the front of the prompt with every request returning 200. The data to detect it
is already stored with each turn and unread.
**Scope.**
- A safety reserve that scales with the effective budget and has a floor. It
replaces or supplements `OUTPUT_SAFETY_MARGIN`, is defined once, and applies
to verified windows, declared overrides and an unverified configured budget
alike. The brief states the tokenizer divergence it is sized to absorb.
- Reading the server-reported prompt-token count from the turn's usage. It is
recorded beside the application's count in the turn's window provenance.
Recorded states:
- `fits`;
- `exceeded`: reported count plus reply reserve is over the window;
- `truncation_suspected`: the reported count is materially below the
application's count;
- `unknown`: no usage was reported.
`exceeded` and `truncation_suspected` are surfaced in the context inspector,
and in the turn's response so the reader sees a notice.
- **Optional, decided in the brief:** calibrating the reserve from observed
server-to-application ratios per endpoint and model. It must be bounded and
quantised so that it does not re-price the prompt prefix every turn (§15.3).
Detection is required either way.
- `ContextOverflow` remains the explicit failure when protected context plus
both reserves exceeds the budget.
- `TECHNICAL-DESIGN.md` §15.2 and `DEVELOPMENT.md`'s context-window section, as
implemented. The root `README.md` stale facts from item 25.
**Non-scope.** The history block trim mechanism, section order, the knowledge
budget share, a per-model tokenizer or any tokenizer download, a hard-coded
window, raising any window, a provider abstraction, a Settings redesign, and
failing or discarding a turn after the fact.
**Likely affected.** `backend/app/context/builder.py`,
`backend/app/contextwindow.py`, `backend/app/routers/adventures/turns.py` and
`insights.py`, `backend/app/providers/openai_compatible.py` (usage, read-only),
the context inspector panel in `frontend/src/pages/Play/`. Tests:
`test_m11_context_window.py`, `test_m11_declared_window.py`,
`test_history_block_trim.py`.
**Acceptance criteria.**
1. **Per configuration.** The application's count plus the output reserve plus
the safety reserve is no more than the effective budget, and the safety
reserve is at least its documented value. This holds for each of: a
verified 4,096 window, a verified 16,384 window, a declared override with no
probe, and an unverified configured budget. There is a test per
configuration.
2. **Divergence.** A scripted provider reports prompt counts at 1.00×, and at
the brief's stated divergence, of the application's count, across a campaign
that fills the window. At both, no turn's reported count plus reply reserve
exceeds the window. At 1.25×, either no turn exceeds (calibration built) or
every exceeding turn is recorded `exceeded`. **An unflagged overflow fails.**
3. **Truncation.** A scripted provider reports 8,194 tokens against an
application count near 15,700. The turn is recorded `truncation_suspected`
in stored provenance, shown in the inspector, and reported in the turn's
response. A provider that reports no usage records `unknown`, never `fits`.
4. **Explicit overflow.** Protected context that fits without the safety
reserve but not with it raises `ContextOverflow` with an actionable message.
The story and state are unchanged (A05, L01).
5. **Canon survives.** `test_the_canon_at_the_front_survives_a_window_far_too_small`
and `test_without_the_cap_the_same_prompt_would_have_overflowed` both pass
with the reserve in place.
6. **Cache stability.** At steady state the history floor still moves in blocks
(`test_history_block_trim.py`). If calibration is built, a test shows the
budget changes only when the observed ratio crosses a stated step.
7. **Real model.** A long run of at least 50 turns runs at a 16,384 window with
memory on, using the reference narrator on the GPU host. Its ten largest
stored prompts are re-counted by the server, as §N did. Every one leaves at
least the documented reserve, and every turn's provenance reads `fits`. The
report carries the headroom table beside v1's.
**Regression requirements.** The full backend and frontend suites; F03, F04 and
M03 coverage; the I-series (a new provenance key must import and export); the
offline container; the browser harness's F05 inspector checks.
**Dependencies.** None. First.
**Compatibility.**
| Area | Effect |
| --- | --- |
| Databases | None expected. Provenance lives in the turn snapshot's JSON, so no migration |
| Bundles | None. Snapshots carry new keys, and absence means "not recorded" |
| Save Points, branches, state, memories, knowledge | None |
| Model settings | None preferred. If a setting is added, it takes a default that reproduces the documented reserve, plus an additive migration |
| Behaviour | At the largest prompts the history window holds slightly fewer actions. Nothing stored changes |
| Docker and local-only | None |
**Test modes.** Deterministic: yes. Real model: yes. Browser: inspector display,
via the harness. Offline: regression only. Long run: yes, at least 50 turns.
**Risk.** Low implementation risk and high value. Its blast radius is the
budget arithmetic of every turn. It changes prompts at full windows, so it
invalidates the v1 long-run evidence for v1.1 (§10).
---
### WP-A2 — Protocol echo hardening and genre-neutral instructions
**Objective.** Stored narration no longer keeps the application-owned protocol
shapes a narrator copies, and the fixed instructions stop supplying
genre-specific identifiers. Recognition uses only strings and syntax the
application itself owns, so story prose is not removed.
**Rationale.** Items 3 and 4. The call notation in the prompt is copied, the
length hint's echo escapes `_is_echoed_instruction`, and a lone trailing
`Scene:` survives. Each leak is replayed as history, so it teaches the next turn.
The example slugs were copied into an office scene.
**Scope.**
- A genre-neutral `EMIT_RULE` example and slug examples. No identifier in the
fixed instructions names a fixture entity or a genre noun.
- A decision, by measurement, on whether the vocabulary is shown in a
notation that does not look callable. The extractor handles call lines
regardless.
- The length hint's wording becomes a named constant shared by the builder and
the extractor, in the way `render.SECTION_HEADINGS` is shared. The extractor
then recognises:
- a line consisting only of a call to an event named in `events.SPECS`
(case-insensitive), optionally `>`-quoted;
- the length hint echoed as a trailing bracket, closed or cut off;
- a renderer heading that is the reply's final line with nothing after it.
- A replay tool that runs the old and new extractors over a corpus of stored AI
turns and writes a per-turn diff report.
- Measurement only: per-turn counts of state fact-line restatements,
model-invented headings, and proposal outcomes (applied, empty, unparseable,
refused), added to `tools/m11_long_run.py`.
**Non-scope.** Removing a fact restated inside a sentence; removing
model-invented headings that are not the renderer's; the event vocabulary itself,
the validator, the fence protocol, or history replay; any model-based cleanup
pass; rewriting stored narration of existing campaigns.
**Likely affected.** `backend/app/narrative/extract.py`, `events.py` (prompt
rendering only), `render.py` (constants), `backend/app/context/builder.py`
(length-hint constant), `tests/test_narrative_state.py`, `tests/test_m11_scifi.py`
(the J03 vocabulary check), `tools/m11_long_run.py`. Also
`TECHNICAL-DESIGN.md` §15.4 and ADR 013's implementation note.
**Acceptance criteria.**
1. **Known shapes are removed.** There is a committed case per shape, cut down
from the four §S.6 turns and anonymised:
- `> set_possession(silver-key, "alice")`;
- `> Create_entity(...)`;
- the length hint echoed as `[Hard limit: …append the state block well
inside the limit.]`, closed and cut off;
- a lone trailing `Scene:`.
The four stored §S.6 texts, replayed, lose exactly those lines. Every other
sentence is intact, asserted as set equality of the remaining sentences.
2. **Adversarial prose is untouched.** Each of these is a test:
- dialogue that mentions `set_possession` mid-sentence;
- `create_entity(ship)` inside a story's own ```python fence;
- a call-shaped line whose name is not in the vocabulary, such as
`> open_door(north)`;
- an in-world bracket, `[Hard limit of the reactor: three hours]`;
- `Scene:` followed by prose;
- `Memory: she remembered the bells`;
- a fact restated inside a sentence;
- a trailing bracket about a state block that lacks the hint's wording.
§O.8's existing negative controls all still pass.
3. **Replay corpus.** Every stored AI turn available from the v1 evidence runs
is replayed through both extractors: §O.8's 443 and the identity run's 10,
kept under `$HOME`, not committed. The package passes when all of these
hold:
- no turn changes except by removing a criterion-1 shape;
- every changed turn appears in the diff report, and the WP report reviews
it;
- the review classifies zero changes as removed story.
The report gives the counts.
4. **False-positive detection.** The replay tool flags any removal that is not
the reply's final segment or a whole line matching a criterion-1 shape, and
any removal of more than a stated share of a turn's prose. An unreviewed
flag fails the package.
5. **Genre-neutral instructions.** A test asserts that `EMIT_RULE`,
`EMIT_REMINDER`, the length hints and the vocabulary text contain no
identifier from either acceptance fixture (Westhaven, Persephone) and none of
J03's genre nouns.
6. **Real model.**
- **Identity diagnostic.** The office fixture at 4,096 on the CPU/HTTPS host
gives 0 proposals naming an example identifier (baseline 1 of 10),
protocol shapes in 0 of 10 stored turns (baseline 4 of 10), and 0 identity
signals. Its `--scripted --inject` self-test still fires.
- **Long run.** At least 50 turns at 16,384 with memory on stores protocol in
0 turns (baseline 0 of 104). Restatement counts are recorded against the
74 of 104 baseline, as a measurement, not a gate.
**Regression requirements.** The full backend suite, including the 25 §O.8
cases; C06 and H05 (refusals are still recorded and shown to the model); J01-J03;
the I-series; the browser harness's hostile-narration checks (H04, H06, H07).
**Dependencies.** After A1. Both change the prompt, and one real-model run can
then cover both. The brief can be written while A1 is under review.
**Compatibility.**
| Area | Effect |
| --- | --- |
| Databases | None. No migration |
| Stored narration of v1.0.0 campaigns | Not rewritten |
| Bundles | None |
| Save Points, branches | None |
| State | None; the validator is unchanged |
| Memories and summaries | Future ones summarise cleaner prose. Existing ones are untouched |
| Docker and local-only | None |
**Test modes.** Deterministic: yes. Replay corpus: yes. Real model: yes. Browser:
regression only. Offline: regression only. Long run: yes, at least 50 turns, and
this run may be the same one as A1's criterion 7 if A1 is already merged.
**Risk.** Medium. The failure to fear is silent removal of prose. Constants-only
recognition, the corpus and the detector are the mitigations.
---
### WP-B — Independent long-term memory retention
**Objective.** An important fact planted early is recoverable at depth 100 or
more from the memory bank itself, with provenance to a memory whose source range
covers the planting turn. This must hold when neither authoritative state, later
narration, the summary nor the history window carries the fact, and lineage
safety must be unchanged.
**Rationale.** Item 5. F02 and M04 pass on state-based recovery, which the owner
accepted. Memory's own retention is unproven, and in a long campaign whose facts
are not all state-shaped it is the only continuity there is.
**Scope.** Diagnosis first, then the smallest sufficient fix.
- **B.1, the diagnostic.** A retention harness that reports, for a planted fact,
each stage with ids and depths:
- **created:** does a memory whose `source_start`..`source_end` covers the
planting turn contain the fact?
- **retained:** is it not `forgotten`?
- **ranked:** where is it in similarity order for the recall query?
- **injected:** is it in the memories section?
The harness runs in two modes. The deterministic mode uses a scripted
narrator, summariser and embedder. The real-model mode uses the reference
narrator and embedder. It also adds a `recovered_through_memory_independent`
verdict to `tools/m11_long_run.py`, keeping the existing verdicts and their
meanings.
- **B.2, the fix,** only at the stages B.1 shows failing, from this bounded menu:
- how the summariser's excerpt is chosen, instead of keeping only the last
2,000 tokens;
- the memory prompt's instruction to keep named facts and objects;
- an eviction rule that does not throw out a never-retrieved early memory
first;
- one additional ranking term (lexical or entity overlap, CONTEXT-AND-MEMORY
§20).
Anything outside the menu needs the owner's agreement in the brief.
**Non-scope.**
- Memory becoming authoritative, or outranking state (F07).
- Any change to lineage attachment or filtering (`tree.attach_memory`,
`forget_node`, E02) or summary lineage (E03).
- Merging memory with imported knowledge, or new embedding models or
dependencies.
- Automatic re-summarisation of existing memories.
- Redesigning cross-layer duplication (§22).
**Likely affected.** `backend/app/memorybank.py` (creation, eviction, ranking);
`backend/app/context/builder.py` (memories section, only if the query changes);
`backend/app/models.py`, only if a field is unavoidable. Tools: `tools/memory_ab.py`,
`tools/m11_long_run.py`. Tests: `test_context_memory.py`, `test_memory_nodes.py`,
`test_m11_leakage.py`, `test_m11_long_run_memory.py`.
**Acceptance criteria.**
1. **Deterministic retention test**, committed. Fact F is planted at depth 3 or
less through narration only, with no state correction and no knowledge source.
These are asserted for the whole run:
- F is never in authoritative state;
- no turn after the planting block contains F's sentinel tokens;
- the summary section never contains F;
- the planting turn is outside the history window at recall.
At a recall depth of 100 or more, the memories section holds a memory that
contains F and whose source range includes the planting depth, and the
context report lists its id and similarity. **The test must fail on the
`432f041` tree**, and the report must name the stage at which it fails.
2. **Past capacity.** The same test runs with more memories written than
`memory_bank_capacity`. The planting-era memory is still active and
retrieved, or its eviction follows a documented, tested rule that the WP
report justifies. No pinned memory is evicted. The "frozen bank" regression
(a new memory evicted at once) still passes.
3. **Lineage.** Fact G is planted only on a line later abandoned by Undo and
divergence. G's memory stays on disk and is absent from every active-line
prompt and from `memories.used`. All 14 tests in `test_m11_leakage.py` and the
E02 tests pass unchanged.
4. **Authority.** A memory contradicting state loses, and memory is still framed
as non-canon (F07).
5. **Real model.** A run of at least 100 turns at 16,384 with memory on, on the
GPU host with logging. The fact is chosen so that the harness can verify the
criterion-1 preconditions on real output. **Pass:** one run meets every
precondition and returns `recovered_through_memory_independent`, with the
memory's provenance. A run whose precondition fails reports which one, and
counts as neither pass nor failure. The report gives the stage-by-stage
diagnostic for every run.
6. **Budget.** The memories section stays inside its budget (F03), and the
steady-state prompt is no larger than A1's arithmetic allows.
7. `CONTEXT-AND-MEMORY.md` §15, §20 and §21 are updated as implemented.
**Regression requirements.** F01-F08, E01-E04, M04 verdicts (the new verdict is
added and the old ones are unchanged), the I-series (memories travel in the
bundle), L04 (`test_memory_rewrite.py`), §O.7's no-write-lock and
failure-recording tests.
**Dependencies.** After A2, so that it measures the prompt v1.1 ships and a
narrator no longer handed protocol to restate. B.1's deterministic harness may
be built earlier.
**Compatibility.**
| Area | Effect |
| --- | --- |
| Databases | No migration preferred. If a memory field is unavoidable, it is additive with a default, tested from a real v1.0.0 database, with schema parity |
| Bundles | Stay v3. Any new memory field is optional on import, and its absence means "unknown". v1.0.0 bundles import |
| Existing memories | Valid and used as they are. A rebuild stays opt-in (`tools/rewrite_memories.py`) |
| Branch history, Save Points, state, knowledge | None |
| Settings | Any change to capacity semantics keeps v1 defaults |
| Docker and local-only | None |
**Test modes.** Deterministic: yes. Real model: yes. Browser: no. Offline:
regression only. Long run: yes, at least 100 turns, possibly several.
**Risk.** High uncertainty, because the outcome depends on the model, and medium
implementation risk. The blast radius is the memory subsystem.
---
### WP-C — Browser release coverage
**Objective.** Every reader-facing behaviour that v1 proved only through the API
or the component suite is driven in a real browser, including an export that
actually leaves the browser as a file.
**Rationale.** Item 6. §T carries two qualifications that exist only because the
harness stops short.
**Scope.** New scenarios in `tools/m11_browser.py`, or a sibling that reuses
`tools/m11_webdriver.py`: Retry; Save Point create and restore; state correction;
narration length; failed generation; export download. Also the Firefox profile
preferences that direct a download to a harness-owned directory under `$HOME`.
The package is harness-only, unless it finds a product defect or a control with
no accessible name. Either is fixed with a regression test and reported as a
product change.
**Non-scope.** New UI, frontend refactors, Selenium or any new dependency,
screenshot diffing, tablet layout, CI.
**Likely affected.** `backend/tools/m11_browser.py`, `backend/tools/m11_webdriver.py`,
the `DEVELOPMENT.md` harness section.
**Acceptance criteria.** The run ends with 0 failed and 0 skipped. It runs against
the built SPA served by FastAPI, with turns from the reference narrator over
trusted-LAN HTTPS. Every assertion reads the rendered DOM or a file on disk.
1. **Retry.** Retry on the newest turn yields a second take, and the indicator
reads 2/2. Stepping to 1/2 shows the original narration unchanged. The takes
persist after a reload.
2. **Save Point.** Create a named Save Point through the UI, play two more
turns, then restore it through the UI. The transcript ends at the named
moment, the position indicator says later story is ahead, and Redo walks
into the later turns. After a reload the Save Point is still listed.
3. **State correction.** A correction submitted through the State panel applies
and is still shown after a reload. A partly refused correction shows its
refusal and reason to the reader.
4. **Narration length.** After changing the control to brief and playing a turn,
then to long and playing a turn, each turn's context inspector shows its
band's word range.
5. **Failed generation.** With the model set through Settings to a name the
server does not serve, a submitted turn shows an error, adds no narration and
keeps the typed input. Setting the model back, the next turn succeeds and the
earlier story is intact.
6. **Export download.** A real click on Export, from both the campaign library
and campaign settings, writes a file with no manual step. The file:
- exists and is not empty;
- parses as `ai-dnd-adventure-v3`;
- imports into a fresh data directory with the same action count, head
position and Save Points.
If the snap Firefox cannot be made to download, the check runs on a non-snap
Firefox under `$HOME`, and `DEVELOPMENT.md` says so. **Exercising only the
click does not pass.**
7. The existing 38 checks pass in the same run.
**Regression requirements.** The existing browser checks; the frontend suite
where a product fix is made.
**Dependencies.** None. It can run at any point. Its final run is repeated on the
v1.1 release candidate.
**Compatibility.** None, unless a product fix is made, and then per §5.
**Test modes.** Browser: yes. Real model: yes, over the HTTPS reference host.
Deterministic: no. Offline: no. Long run: no.
**Risk.** Low product risk. The medium risk is harness flakiness, so every wait is
on a DOM condition, never a sleep used as an assertion.
---
### WP-D — Recovery honesty
**Objective.** A backup is kept only after a full integrity check. A campaign
whose export exceeds what import accepts is exported with a warning saying so,
not silently.
**Rationale.** Items 7 and 10a. Both are recovery gaps that are known, measured
and cheap to close. Neither needs a policy decision.
**Scope.**
- `backup.py` runs `PRAGMA integrity_check` on the finished copy, and the time
it takes is measured.
- Export compares the serialised bundle's size with
`limits.MAX_IMPORT_BODY_BYTES`. When it is over, the file is still delivered
and the reader sees a warning naming the limit and what it means.
- The import refusal for an oversized bundle names the limit.
- `DEVELOPMENT.md` documents the ceiling as measured: about 13 kB per action on
a real campaign, and M9's conservative figure of about 279 turns.
**Non-scope.** Scheduled backups; raising the import limit or streaming import; a
bundle format change or further compression; a restore button.
**Likely affected.** `backend/app/backup.py`, `backend/app/routers/adventures/bundle_io.py`,
`frontend/src/pages/Campaigns.jsx`,
`frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx`,
`frontend/src/pages/backup.test.jsx`, and the backup and bundle tests.
**Acceptance criteria.**
1. A backup of a healthy database reports `integrity_check` ok and is kept.
2. A copy with damage that `integrity_check` detects and `quick_check` does not,
such as an index inconsistent with its table, is rejected and not kept.
Existing backups are still never overwritten.
3. The time for `integrity_check` is recorded on the 100-turn evidence database
and on a synthetic database of 100 MB or more, and the backup completes
through the UI on both.
4. Exporting a fixture campaign over 20 MB succeeds, delivers the file, and
shows a warning naming the import limit, asserted in the API response and in
a component test. A campaign under the limit shows no warning.
5. Importing that bundle is refused with a message naming the limit.
6. A normal campaign's export is byte-identical before and after the package,
apart from timestamp fields.
**Regression requirements.** I01-I07, L01-L04, the backup case in
`test_m11_migration.py`, `backup.test.jsx`, and the offline container's
export and import.
**Dependencies.** None.
**Compatibility.** Databases and bundles: no change. Backups: a stricter check,
with the same file.
**Test modes.** Deterministic: yes. Browser: the warning, optionally in WP-C's
harness. Offline: regression only. Real model: no. Long run: no.
**Risk.** Low.
---
### WP-E — Control-boundary contrast
**Objective.** Control boundaries meet WCAG 1.4.11's 3:1 against their panel, at
rest and on hover.
**Rationale.** Item 9.
**Scope.**
- The boundary token values in `frontend/src/styles/tokens.css`, and any
component that overrides them.
- `tools/contrast_audit.py` treats boundary pairs below 3:1 as failures, not
advisories.
- The browser harness measures rendered boundary contrast.
- Before-and-after screenshots for the owner's approval.
**Non-scope.** A palette redesign, typography, layout, tablet work, a
screen-reader audit, and other WCAG criteria. A defect the same measurement finds
is recorded, not taken on.
**Likely affected.** `frontend/src/styles/tokens.css`, `backend/tools/contrast_audit.py`,
`backend/tools/m11_browser.py`, `frontend/src/a11y.test.jsx`.
**Acceptance criteria.**
1. `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and
exits 0 on the package tree.
2. Every text pair still clears 4.5:1 (1.4.3). The rendered text contrasts
measured by the harness do not fall below v1's (14.57, 5.48, 13.57 and
5.88:1) without a stated reason.
3. The rendered boundary of the story input and of a primary control, at rest
and on hover, is at least 3:1. The focus indicator is still visible.
4. The owner approves the before-and-after screenshots, and the WP report records
the approval.
**Regression requirements.** The frontend suite and lint; the browser harness's
accessibility checks.
**Dependencies.** None. Its browser checks go into WP-C's harness if WP-C has
landed, and into `m11_browser.py` otherwise.
**Compatibility.** None.
**Test modes.** Deterministic: yes. Browser: yes. Everything else: no.
**Risk.** Low.
---
## 9. Order and dependencies
```text
A1 context safety reserve ──► A2 protocol echo + neutral instructions ──► B memory retention
(B.1 harness may start early)
C browser coverage ── independent
D recovery honesty ── independent
E boundary contrast ── independent (checks land in C's harness if C is first)
│
▼
v1.1 release validation (§10)
```
**Recommended order: A1, A2, B, C, D, E.**
- **A1 first.** It is the only item that can silently remove canon from a
prompt. It is deterministic to test, small in blast radius, and it
establishes the provenance that A2's and B's real-model runs will read.
- **A2 second.** Its leak is silent and compounds through replay. It must
precede B, because it changes what the memory pass summarises.
- **B third.** It carries the highest continuity value and the most uncertainty.
Measuring it before A1 and A2 settle would measure a prompt v1.1 does not ship.
- **C, D and E** close verification, recovery and accessibility gaps with no
known story risk. They depend on nothing, so the owner may move any of them
earlier. While B waits on long runs is a natural slot. Each still has its own
brief and review.
| Property | A1 | A2 | B | C | D | E |
| --- | --- | --- | --- | --- | --- | --- |
| Can begin independently | yes | after A1 | after A2 | yes | yes | yes |
| Schema or migration | no | no | avoid; additive if unavoidable | no | no | no |
| Export format | no (additive evidence key) | no | no; optional field at most | no | no | no |
| Changes acceptance tests (`V1-ACCEPTANCE-TESTS.md`) | no | no | no | no | no | no |
| Changes the stored prompt | yes | yes | possibly | no | no | no |
| Real-model validation | yes | yes | yes | yes, for turns | no | no |
| Browser testing | regression | regression | no | yes | optional | yes |
| Offline / no-network testing | regression | regression | regression | no | regression | no |
| Long-run testing | yes, 50 turns or more | yes, 50 turns or more | yes, 100 turns or more | no | no | no |
| Risk | low | medium | high uncertainty | low | low | low |
## 10. Compatibility summary
| Area | A1 | A2 | B | C | D | E |
| --- | --- | --- | --- | --- | --- | --- |
| v1.0.0 campaign databases | open unchanged | unchanged | unchanged; an additive migration only if unavoidable | — | unchanged | — |
| Export and import bundles | v3; new evidence key | — | v3; an optional field at most | — | v3; export warns | — |
| Save Points | — | — | — | — | — | — |
| Branch history | — | — | lineage rules unchanged | — | — | — |
| Narrative state | — | validator unchanged | memory never outranks state | — | — | — |
| Memories and summaries | — | new ones from cleaner prose | creation, retention and ranking change; existing rows kept | — | — | — |
| Knowledge sources | — | — | — | — | — | — |
| Model settings | none preferred | — | capacity defaults kept | — | — | — |
| Docker and local-only | — | — | — | — | — | — |
## 11. v1.1 release criteria
v1.1 is not called v1.1.0 until all of the following hold on one release
candidate tree:
1. **Every package in scope is accepted**, each with its report, and with its
acceptance criteria passing on the candidate or on a tree whose product code
the candidate carries unchanged.
2. **The v1 contract still passes.** All 82 REQUIRED FOR V1 tests hold, with H09
not applicable on the same condition. None is relaxed.
3. **Suites and builds:** the backend suite, the frontend suite and lint, the
production build, and `docker build --no-cache`, with the image's SPA
identical to the local build.
4. **Offline:** `tools/m11_offline.py` passes all checks with no network and a
fresh volume.
5. **Browser:** the existing 38 checks plus WP-C's and WP-E's pass with 0 failed
and 0 skipped, on the candidate, over trusted-LAN HTTPS.
6. **One v1.1 long run** on the candidate's product code: 100 turns or more at a
16,384 window with memory on, on the GPU host with logging. It passes M01-M04.
A1's headroom table shows the documented reserve on every re-counted prompt,
and every turn is `fits`. A2's leak count is 0. B's independent-retention
verdict is recorded.
12. **WP-B's memory limitation is reported, not summarised away.** The v1.1
release report states each of these, and never shortens them to "WP-B
passed":
- deterministic independent-memory recovery: **PASS**;
- reference-model independent-memory recovery: **FAIL** on the
precondition-valid attempt;
- the failing stage: **memory creation**, the summariser's content
selection;
- the owner's decision to accept that limitation for v1.1.
The release long run's independent-retention verdict (item 6) is read
against it. A recovery there is reported as evidence, not as a reversal of
the limitation, unless it meets every isolation precondition.
13. **Carried residuals are listed with their status:**
- the mid-reply narrator instruction echo that A2's trailing cleanup does
not remove (WP-B.1 §K);
- the doubled full stop in the memory-search scene text (WP-B.2 §R 10).
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
protocol shapes in stored narration.
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all
checks.
9. **Upgrade from a real v1.0.0 database.** A database is created by the
`v1.0.0` tree and played. It has retained history, an undone head, Save
Points, memories, summaries, imported knowledge and a narration-length
choice, and uses loopback or placeholder settings with no real hostnames.
Opened by the candidate, its transcript, head, Redo availability, state, Save
Points, memories, knowledge and settings compare identical, and any migration
is forward-only with schema parity. A v1.0.0 export imports into v1.1.
Because the format stays v3, a v1.1 export of that campaign is checked for
import into v1.0.0, and the result is reported.
10. **Documentation:** `README.md`, `DEVELOPMENT.md`, the as-implemented sections
named by each package, and `VERSION.md`.
11. **Owner events**, each separate: the signed release commit, `main`, and a
`v1.1.0` tag.
## 12. Scope recommendation
**Recommended: Option 1, a focused v1.1.**
| Ship in v1.1 | Defer to v1.2 | Future / optional |
| --- | --- | --- |
| WP-A1 context safety reserve | Scheduled backups (10b) | Real media-provider adapters (11; K05, K06) |
| WP-A2 protocol echo and neutral instructions | Raising or streaming the import limit (7) | Whole-transcript copy, story search (§77, §78) |
| WP-B independent memory retention | Reducing in-sentence restatement (2) | Discarded-history recovery screen (§63) |
| WP-C browser release coverage | Stale-scene and derived-contamination detectors in the identity diagnostic (8) | Tablet layout |
| WP-D recovery honesty | Cross-layer duplication suppression (17) | Editable RPG world state |
| WP-E control-boundary contrast | | Media debt: `ambience`, visual-profile recovery and UI (20) |
| | | Trusted-LAN residual limits: address pinning against DNS rebinding (21) |
| | | Comparative narrator-model recommendation (13) |
**Why focused.** Four of the six packages close silent failure modes or
verification gaps that the v1 evidence itself named. The other two are small and
deterministic. Everything deferred is either a new feature, or needs a policy
decision the owner has not been asked for, or rests on an occurrence that has not
happened. A broader v1.1 that took scheduled backups and media would add the two
packages with the most new surface. It would also push the long-run
re-validation, which every prompt change already requires, further from the
changes it validates.
**If B's real-model criterion cannot be met.** Precondition-valid runs may recover
nothing even after B.2. In that case the owner chooses between shipping v1.1 with
B's deterministic criteria met and the real-model result recorded as a residual
risk, or holding v1.1 for B. The plan does not pre-decide this.
## 13. Coding briefs needed next
These briefs are not written here, and none is to be executed from this document.
1. **WP-A1 — context-window safety reserve. Write this one first.** It carries
the highest user risk, has no dependencies, is deterministically testable,
and produces the provenance that later real-model runs read. The brief must
settle three things: the tokenizer divergence the reserve is sized for, whether
calibration is built or detection alone, and whether any setting is added
(recommended: none).
2. WP-A2 — protocol echo hardening and genre-neutral instructions. It can be
drafted while A1 is under review.
3. WP-B — independent memory retention: B.1 diagnostic, then B.2 fix. Two briefs
are an option if B.1's findings should be reviewed before a fix is chosen.
4. WP-C — browser release coverage.
5. WP-D — recovery honesty.
6. WP-E — control-boundary contrast.
7. v1.1 release validation, written only after every in-scope package is
accepted.
## 14. Decisions for the owner
- Confirm Option 1, the focused v1.1.
- A1: the divergence the reserve must absorb; calibration or detection only; that
a detected truncation is flagged, not turned into a failed turn.
- B: whether B.1 and B.2 are one brief or two; the choice in §12 if the
real-model criterion is not met.
- D: that the 20 MB import limit stays in v1.1.
- E: approval of the visual change.
- Whether v1.1 real-model validation stays on the reference narrator, as this plan
recommends.
- The report naming and rotation in `planning/reports/` (§5 rule 5).
+146 -3
View File
@@ -1,8 +1,151 @@
# Planning Package Version
- **Package:** Adventure Storyteller Planning Package v4.0
- **Revision date:** 2026-09-14
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M11 complete**; **M11 accepted at its closeout (2026-09-14), and the v1 release gate passed** on the release-candidate tree. All 82 REQUIRED FOR V1 tests pass, with H09 not applicable. There is no further planned milestone. The signed closeout commit and any `v1.0.0` tag are the repository owner's, and neither exists as of this revision.
- **Package:** Adventure Storyteller Planning Package v4.6
- **Revision date:** 2026-09-16
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C is signed as `59b5ebc`, and **WP-D (recovery honesty) and WP-E (control-boundary contrast) are signed as `87a4032`**. **Integrated v1.1 release validation has run on candidate `87a4032` and PASSED** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): 82 REQUIRED v1 tests hold (81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a `--no-cache` image whose SPA is file-for-file identical to the local build, offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window run passing M01-M04 with every turn `fits` and 0 protocol leaks, identity 0/0, recovery 16/16, a real v1.0.0 upgrade identical on all 15 fields with both bundle directions importing, and a release smoke of 15/15. WP-B's reference-model memory limitation remains an accepted, documented residual. **v1.0.0 is still the released version: no release commit, no `main` update and no `v1.1.0` tag exist** — those are the owner's events.
## v4.6 — v1.1 integrated release validation (2026-09-16)
Release validation of candidate `87a4032`, not a work package: no requirement,
acceptance test, schema, bundle format or product code changed. The evidence is
`reports/v1.1/V1.1-RELEASE-REPORT.md`, sections A-W.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/v1.1/V1.1-RELEASE-REPORT.md` | **New.** The frozen candidate, the 82-row v1 acceptance matrix, every gate's result, the A1 headroom table, the A2 leak count, the WP-B verdict kept in both halves, the residual classification, and the final decision | release report |
| `reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL` **PENDING → APPROVED**, sourced and dated to the owner's release-validation brief; the signed commit predated the review | correction of record |
| `V1.1-PLAN.md`, `planning/README.md`, `VERSION.md` | Status: WP-D/WP-E signed `87a4032`; validation passed; the three owner events still outstanding | status |
| `README.md` | v1.0.0 **remains** released; v1.1 implemented and validated but untagged; schema figure corrected to 94 | product docs |
| `backend/tools/v11_upgrade_check.py`, `backend/tools/v11_release_smoke.py` | **New**, harness only: the real-v1.0.0 upgrade gate and the release-shaped smoke test | tooling |
| `backend/tools/m11_identity.py` | Reads `AIDND_TEST_EMBED_MODEL` and enables the memory bank, so the diagnostic can run with memory on as the gate requires | tooling |
**Requirement changes: zero. Product-code changes: zero.**
**Outcome:** `V1.1 RELEASE VALIDATION: PASS`. Carried residuals: WP-B's
reference-model memory limitation, the mid-reply instruction echo (still
reproducible on the stored fixture, absent from release evidence), and the
doubled full stop. K1 is classified v1.2 backlog. **No release commit, no `main`
update, no `v1.1.0` tag.**
## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15)
The last two planned v1.1 packages, implemented and reported separately. No
requirement, acceptance test, schema or bundle format changed. WP-D's evidence is
in `reports/v1.1/V1.1-WP-D-REPORT.md`, WP-E's in `V1.1-WP-E-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/v1.1/V1.1-WP-D-REPORT.md` | **New.** Full `integrity_check` on the finished backup copy, proved against a fixture `quick_check` calls healthy; an export that says when this version could not import it back, carried in headers because the response body *is* the bundle | work-package report |
| `reports/v1.1/V1.1-WP-E-REPORT.md` | **New.** Control boundaries raised to clear WCAG 1.4.11 (3:1), the contrast audit turned from a report into a gate, rendered before/after boundary measurements, and the `.slice-7` finding the change itself created | work-package report |
| `V1.1-PLAN.md` | Status: WP-C signed `59b5ebc`; WP-D and WP-E complete and staged | status |
| `planning/README.md` | Current state | index |
| `DEVELOPMENT.md` | How large an export can get: the 20 MB import ceiling, ~13 kB per action, M9's ~279-turn figure and why it is not a turn limit, and the four export headers | developer docs |
**Requirement changes: zero.**
Two owner decisions are outstanding, both recorded rather than assumed: WP-E's
before/after screenshots await approval (`OWNER SCREENSHOT APPROVAL: PENDING`),
and WP-D records that no real-browser click was made on *Back up now* — the
endpoint that button calls was driven instead.
## v4.4 — WP-C browser release coverage (2026-09-15)
Harness work, with one narrow product fix it found. No requirement, acceptance
test, schema or bundle format changed. The evidence is in
`reports/v1.1/V1.1-WP-C-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `reports/v1.1/V1.1-WP-C-REPORT.md` | **New.** The six reader workflows driven in a real browser, the existing 38 checks, the download environment, harness defects J1-J7, and product defects K1 (open) and K2 (fixed) | work-package report |
| `V1.1-PLAN.md` | Status: WP-B.2 committed; WP-C complete and staged | status |
| `planning/README.md` | Current state | index |
| `DEVELOPMENT.md` | The browser harness command and flags; Firefox download preferences and the `$HOME` rule; what counts as a finished download; waiting on conditions, never sleeping | developer docs |
**Requirement changes: zero.**
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
WP-B.1's diagnostic (`beb17ad`) placed three memory deficiencies; WP-B.2 corrects
exactly those, one at a time, each verified before the next. No requirement or
acceptance test changed, and no schema, bundle format or setting default
changed. Evidence and the WP-B decision are in
`reports/v1.1/V1.1-WP-B2-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1 WP-B.2)**: a block longer than 2,000 tokens is shown to the summariser as its opening and its end with an omission marker, inside the same budget; a block that fits is unchanged; the marker is never stored. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` §18 | **As implemented**: the retrieval query is the player's input plus a bounded scene context, recorded per turn. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` §20 | **As implemented**: `final = semantic + 0.15 × lexical`, the rarity-weighted lexical term over the input, the sweep that chose the weight, pins and ties. | as-implemented record |
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1)**: the memory prompt is unchanged from v1.0.0, and the accepted limitation: the reference summariser can omit or misattribute a fact from a block it was given whole. | as-implemented record |
| `DEVELOPMENT.md` | The GPU-host kernel/Ollama watch command corrected: `-k -u ollama` matched nothing; the OR form records both. | developer docs |
| `CONTEXT-AND-MEMORY.md` §21 | **As implemented**: selection unchanged in shape; eviction ordered by coverage first (boundaries kept, smallest hole first), recency second, v1.0.0 order as fallback. | as-implemented record |
| `V1.1-PLAN.md` | Status: WP-B accepted with a documented real-model limitation. §11 release criteria 12 and 13: the v1.1 release report must state the deterministic PASS and reference-model FAIL at memory creation, and list the carried residuals. | status, release gate |
| `planning/README.md` | Current state. | index |
| `reports/v1.1/V1.1-WP-B2-REPORT.md` | **New.** B2.1-B2.3 designs and evidence, full deterministic acceptance, real-model attempts, the rejected B2.4 prompt experiment, compatibility, offline, and the WP-B disposition. | work-package report |
**Requirement changes: zero.**
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
Two v1.1 work packages, implemented in sequence. No requirement or acceptance
test changed, and no schema or bundle format changed. Evidence is in
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`.
| Document | Change | Kind |
| --- | --- | --- |
| `TECHNICAL-DESIGN.md` §15.2 | **As implemented (v1.1 WP-A1)**, covering: the safety reserve, `max(256, ceil(5%))`; what replaced M6's 64-token margin; the server's own count; the four accounting states; and keeping the turn. | as-implemented record |
| `TECHNICAL-DESIGN.md` §15.4 | **As implemented (v1.1 WP-A2)**: the vocabulary shown in the wire format, genre-neutral placeholders, and extractor rules R1-R4, each anchored to application-owned text. | as-implemented record |
| `DECISIONS/013-authoritative-narrative-state-document.md` | An implementation note for v1.1. No event type, field, validation rule or proposal record changed. | as-implemented note |
| `V1.1-PLAN.md` | Status: A1 and A2 implemented and staged. | status |
| `planning/README.md` | Current state. | index |
| `reports/v1.1/V1.1-WP-A1-A2-REPORT.md` | **New.** The combined review package, with A1 and A2 kept separate. | work-package report |
| `README.md`, `DEVELOPMENT.md` | The reserve and the accounting. The stale backend test count and the Screenshots paragraph are corrected. | developer docs |
**Corrective work before commit (owner review, 2026-09-14).** Report §R.
| Document | Change |
| --- | --- |
| `TECHNICAL-DESIGN.md` §15.2 | A cold model is loaded once before its turn is built (`contextwindow.ensure_window`). |
| `TECHNICAL-DESIGN.md` §15.4 | R5: the echoed continue hint is recognised by its own sentence, and the application-opened tail above it is removed. |
| `DEVELOPMENT.md`, `README.md` | The cold-model load, in operator terms. |
**Requirement changes: zero.**
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
Documentation only. No product code, no requirement, no acceptance test and no
schema changed.
**Part 1 — post-release correction.** v4.0 was written before three owner
events that have since happened: the closeout commit was signed (`432f041`),
`main` was fast-forwarded to it, and the signed tag `v1.0.0` was created on it
and pushed. The current-state wording that said otherwise is corrected. The M11
report is not edited: its §T records the state at closeout, which was true when
written.
| Document | Change | Kind |
| --- | --- | --- |
| `README.md` | Status: v1.0.0 released; tag and `main` at `432f041`; v1.1 on `v1.1-development`. | developer docs |
| `planning/README.md` | Current state, the release events, the map (v1.0.0 released, v1.1 planning), the stop rule restated for work packages, `V1.1-PLAN.md` in the document table and reading order. | index |
| `planning/BUILD-MILESTONES.md` | Header status, **stale since M8** ("M9 is next"), now says complete and closed. A dated post-release note under M11's status, leaving the closeout paragraph as it stood. The *Post-v1 backlog* points at its triage. | milestone status |
| `planning/VERSION.md` | This header and entry. | package version |
**Part 2 — the v1.1 plan.**
| Document | Change | Kind |
| --- | --- | --- |
| `planning/V1.1-PLAN.md` | **New.** Every post-v1 backlog item and every other recorded v1 residual risk, triaged. Six v1.1 work packages in order (A1, A2, B, C, D, E), each with objective, rationale, scope, non-scope, affected subsystems, acceptance criteria, regression requirements, dependencies, compatibility and risk. The v1.1 release criteria. A v1.2 list and future work. | planning |
**Recommended scope: a focused v1.1.** Context-window safety, protocol-echo
hardening, independent memory retention, browser coverage, recovery honesty and
control-boundary contrast. Scheduled backups, identity-diagnostic extensions and
media adapters are deferred, with the reason for each.
**First brief to write:** WP-A1, the context-window safety reserve.
**Requirement changes: zero.** Every v1.1 package improves the implementation of
an existing requirement, so `SPECIFICATION.md` and `V1-ACCEPTANCE-TESTS.md` are
unchanged.
## v4.0 — M11 closeout: v1 release validation accepted (2026-09-14)
@@ -0,0 +1,881 @@
# Adventure Storyteller v1.1 — Integrated Release Validation
**Status:** COMPLETE — **V1.1 RELEASE VALIDATION: PASS**. The decision, and what
it deliberately does not cover, is in §W.
This report answers one question: **does this exact candidate preserve the
complete v1 contract and satisfy every accepted v1.1 package on one integrated
release tree?** It is release validation, not a work package. Nothing here adds
a feature, and no release tag is created by it.
---
## A. Repository / provenance
| | |
| --- | --- |
| **Candidate SHA** | **`87a40326a29533c8d52c9f9f41022e7b499b1de7`** |
| Branch | `v1.1-development`, up to date with `origin/v1.1-development` |
| Working tree at freeze | **clean** — nothing modified, nothing staged |
| Commit | *v1.1: harden recovery and control boundaries* (WP-D + WP-E) |
| **Owner signature** | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, made 2026-09-16 05:37:13 EDT |
| Tag at HEAD | **none** — no `v1.1.0` tag exists |
| v1.0.0 baseline | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, **an ancestor** |
| Package ancestry | `d63804f` (WP-A1/A2), `beb17ad` (WP-B.1), `0c1ba83` (WP-B.2), `59b5ebc` (WP-C) — **all ancestors** |
| Diff v1.0.0..HEAD | 61 files, +14,833 / −366 |
| LICENSE / PROVENANCE | **unchanged since v1.0.0** (empty diff) |
### A.1 Frozen candidate identity
| | |
| --- | --- |
| Dependency locks | `backend/requirements.txt` `sha256:ed28bc0f8970cf4e…`, `frontend/package-lock.json` `sha256:355cb370837ade01…`, `frontend/package.json` `sha256:2016580ddfa176a9…`, `backend/requirements-dev.txt` `sha256:06d7695816b201e9…` |
| Schema | `LATEST_VERSION` **94**, 93 migrations (`PRAGMA user_version`) |
| Bundle format | **`ai-dnd-adventure-v3`** |
| Import ceiling | 20 MB (`MAX_IMPORT_BODY_BYTES`), unchanged |
| Frontend build | `dist` built 2026-09-16T05:39:55, 16 files, `sha256(dist) = ea2753ad24f61959fe084f4674911acc` |
| Docker image | `storyteller:release-87a4032`, `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398`, 312 MB |
| Firefox / geckodriver | 155.0.1 / 0.37.1 (2026-09-04) |
| Docker | 29.8.0, build 88096ef |
| CPU HTTPS reference host | Ollama **0.33.0**, serving `qwen2.5:3b-instruct`, `qwen2.5:3b-instruct-16k`, `nomic-embed-text:latest`; certificate verifies through the machine's CA store with no bypass |
| GPU inference host | Ollama **0.34.0**; `qwen2.5:3b-instruct-16k` digest `21ff8cc52f375f19`, `nomic-embed-text:latest` digest `0a109f422b47e3a3`; no model resident at start |
| Evidence root | `$HOME/v11-evidence/release-87a4032/` — never `/tmp`, and no real hostname appears in any committed file |
---
## B. Package acceptance inventory
| Package | Status | Source |
| --- | --- | --- |
| **WP-A1** context-window safety reserve | **ACCEPTED** | `V1.1-WP-A1-A2-REPORT.md`, signed `d63804f` |
| **WP-A2** protocol-echo cleanup, genre-neutral prompting | **ACCEPTED** | same report and commit |
| **WP-B** independent memory | **ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION** | `V1.1-WP-B1-REPORT.md`, `V1.1-WP-B2-REPORT.md` §S, signed `beb17ad` / `0c1ba83` |
| **WP-C** browser release coverage | **ACCEPTED** | `V1.1-WP-C-REPORT.md`, signed `59b5ebc` |
| **WP-D** recovery honesty | **ACCEPTED** | `V1.1-WP-D-REPORT.md`, signed `87a4032` |
| **WP-E** control-boundary contrast | **ACCEPTED** | `V1.1-WP-E-REPORT.md`, signed `87a4032` |
### B.1 WP-B's qualification, carried whole
The WP-B disposition is **not** shortened to "WP-B passed" anywhere in this
report. Its own §S records:
```text
B2.1 RANKING: PASS
B2.2 EVICTION: PASS
B2.3 EXCERPT CREATION: PASS
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED
DETERMINISTIC WP-B: PASS
REAL-MODEL WP-B: FAIL
WP-B OVERALL:
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION
```
In the release contract's own words (§11 items 12), that is:
```text
DETERMINISTIC WP-B: PASS
REFERENCE-MODEL INDEPENDENT MEMORY: FAIL
OWNER ACCEPTED THE LIMITATION FOR v1.1
```
The failing stage is **memory creation** — the summariser's content selection —
not ranking, eviction or injection, each of which passes deterministically.
### B.2 WP-E screenshot approval — a correction of record
The committed WP-E report read `OWNER SCREENSHOT APPROVAL: PENDING`, because it
was written before the owner reviewed the images. The owner's release-validation
brief (2026-09-16) states the before/after screenshots are approved and
instructs this validation to record it. The WP-E report is updated to
`APPROVED` as part of this closeout (§V), sourced to that brief and dated. No
visual code changed during release validation, so the approval stands (§Q).
---
## C. v1 acceptance matrix
Every test marked **REQUIRED FOR V1** — there are **82** — against evidence taken
on **this candidate**. Evidence types follow M11's: `browser` (the 101-check run,
§G), `campaign` (the 102-turn integrated run, §H), `container` (the offline run
on the candidate image, §F), `process` (spawned server processes — recovery §M,
upgrade §N), `suite` (the 1,723-test backend suite, §D). No historical result
from different product code is used where the contract asks for candidate
evidence.
**Result: 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified.**
### A — Local-first operation
| ID | Result | Evidence on this candidate |
| --- | --- | --- |
| A01 Start application offline | **PASS** | container: first page load, fresh volume, no route and no DNS |
| A02 Storyteller loopback default | **PASS** | suite; every harness reached it on `127.0.0.1`; `docker-compose.yml` publishes `127.0.0.1:8000:8000` |
| A03 No cloud API key | **PASS** | suite; container: no secret in an export |
| A04 Campaign survives restart | **PASS** | campaign: **3 process restarts, 4 process starts**, state compared across each; container: campaigns survive a container restart |
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real `failed_call` at turn 69 against an unserved model, play resumed; container: same with no model reachable |
| A06 Trusted-LAN Ollama inference | **PASS** | browser: the whole 101-check run over **trusted-LAN HTTPS with a private CA**, verification on, no bypass, storyteller loopback-bound |
### B — Core play
| ID | Result | Evidence |
| --- | --- | --- |
| B01 Natural language action | **PASS** | browser (real turns through the UI) + campaign (102 accepted) |
| B02 Dialogue input | **PASS** | campaign: dialogue beats in the turn list |
| B03 Continue | **PASS** | suite; browser: the Continue control present and enabled |
### C — Story authority and state
| ID | Result | Evidence |
| --- | --- | --- |
| C01 Campaign canon is preserved | **PASS** | campaign: canon present in the prompt on **102 of 102** turns |
| C02 Possession state | **PASS** | campaign (the silver key) + suite |
| C03 Character knowledge is not invented | **PASS** | suite |
| C04 Manual state correction | **PASS** | campaign: **2 state corrections**; browser: C3's accepted and refused corrections; suite |
| C05 Canon beats reference | **PASS** | suite |
| C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real extraction across 102 turns, every event validated or refused; suite |
### D — Non-destructive history
| ID | Result | Evidence |
| --- | --- | --- |
| D01 Undo one turn | **PASS** | browser + campaign (`undo`) |
| D02 Minimum five undos | **PASS** | suite; campaign (`undo_redo`) |
| D04 Redo | **PASS** | browser + campaign |
| D05 Redo invalidated by new continuation | **PASS** | campaign: `diverged`, after which Redo is gone |
| D06 Retry narrator response | **PASS** | campaign: **2 retries** |
| D07 Select prior retry take | **PASS** | campaign: `take_selected` |
| D08 Retry does not delete prior take | **PASS** | campaign + suite |
| D09 Edit earlier user input | **PASS** | suite |
| D10 Edit narrator output | **PASS** | suite; browser (hostile-Markdown plants through the narrator-edit path) |
| D11 Named checkpoint | **PASS** | campaign: **2 Save Points**; browser; suite |
| D12 Restore checkpoint | **PASS** | campaign: `save_point_restored`; browser |
| D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; recovery §M |
| D14 Delete checkpoint | **PASS** | suite; browser: the delete confirmation dialog |
### E — Branch and derived-data isolation
| ID | Result | Evidence |
| --- | --- | --- |
| E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py`, with a positive control |
| E02 Abandoned memory cannot leak | **PASS** | as above |
| E03 Abandoned summary cannot leak | **PASS** | as above |
| E04 Scene state is lineage-safe | **PASS** | as above |
### F — Long-term memory and context
| ID | Result | Evidence |
| --- | --- | --- |
| F01 Recent turns remain coherent | **PASS** | campaign: history populated every turn, newest always included |
| F02 Old important event retrieval | **PASS** | campaign §K: the planting turn outside the window and the fact recovered — **through authoritative state**, not independent memory (§K states which) |
| F03 Prompt remains bounded | **PASS** | campaign: 1,602–14,982 tokens against a 16,384 budget across 102 turns |
| F04 Output token reserve | **PASS** | campaign: `output_reserve` 500 present and subtracted on every turn |
| F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt |
| F06 Retrieval provenance | **PASS** | campaign: knowledge and memory provenance per turn; suite |
| F07 Heuristic memory is not canon | **PASS** | suite |
| F08 Memory failure is non-fatal | **PASS** | suite; container: derived work fails with no model and turns still commit; campaign: **0 post-turn failures**, no database-lock errors |
### G — Imported knowledge
| ID | Result | Evidence |
| --- | --- | --- |
| G01 Import local text | **PASS** | browser: through the real file input; container: offline |
| G02 Import local Markdown | **PASS** | campaign: **3 sources imported**; browser |
| G03 Classification | **PASS** | campaign: all three classes; recovery §M confirms them after a move |
| G04 Disable knowledge source | **PASS** | suite |
| G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite |
| G06 Reference retrieval | **PASS** | suite |
| G07 Inspiration is low authority | **PASS** | suite |
| G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all, import still works |
| G09 Remote Markdown image does not auto-load | **PASS** | browser: no remote image src in the rendered story |
| G10 Prompt injection in source is treated as data | **PASS** | suite; browser: injection text rendered as text |
### H — Security
| ID | Result | Evidence |
| --- | --- | --- |
| H01 No unexpected outbound connections | **PASS** | container (no network at all) + suite `test_egress.py` |
| H02 No telemetry | **PASS** | suite |
| H03 No cloud provider required | **PASS** | container: a full campaign offline |
| H04 Model output cannot execute shell | **PASS** | browser + suite |
| H05 Invalid state event rejected | **PASS** | suite; campaign: refusals recorded |
| H06 Stored XSS protection | **PASS** | browser: `onerror` and `<script>` in accepted narration, neither executed |
| H07 JavaScript URL protection | **PASS** | browser: no `javascript:` href in the DOM |
| H08 Path traversal import rejected | **PASS** | suite |
| **H09 ZIP Slip protection** | **NOT APPLICABLE** | the product extracts no archives, and a test enforces it — the same condition v1 recorded |
| H10 Restrictive CORS and local API behaviour | **PASS** | browser: an unknown API path is a 404 with a non-HTML body; suite |
| H11 No first-use runtime asset download | **PASS** | container: every asset local with no network; suite: the tokenizer table is vendored |
| H12 Inference endpoint enforcement | **PASS** | suite; smoke §R: a public endpoint refused by the running image; the container's refusal of an unresolvable name is the same policy (§R) |
### I — Export, import and recovery
| ID | Result | Evidence |
| --- | --- | --- |
| I01 Export campaign | **PASS** | campaign: the 102-turn campaign exported (3,071,683 bytes) |
| I02 Import exported campaign | **PASS** | process §M: imported into a database and directory that never existed |
| I03 Branch/disposable history export | **PASS** | §M: **88 actions retained beyond the active line** after the move |
| I04 Checkpoint export | **PASS** | §M: both Save Points restore after the move |
| I05 Knowledge provenance export | **PASS** | §M: all three classes with their content |
| I06 Database/export contains no API secrets | **PASS** | §M: the bundle carries no secret; suite; container |
| I07 Export/import preserves an undone active head | **PASS** | §N: the v1.0.0 campaign's undone head and `can_redo` survive upgrade and both bundle directions; suite `test_m9_portability.py`. *(The long-run bundle ended head-at-tip, so §M exercises the other case — stated in §M rather than implied)* |
### J — Genre neutrality
| ID | Result | Evidence |
| --- | --- | --- |
| J01 Science-fiction campaign | **PASS** | suite `test_m11_scifi.py` |
| J02 Generic entity support | **PASS** | as above |
| J03 Genre profiles are configuration | **PASS** | as above |
### K — Future media architecture
| ID | Result | Evidence |
| --- | --- | --- |
| K01 Scene snapshot exists | **PASS** | suite; container: a scene packet builds offline |
| K02 Visual character profile | **PASS** | suite |
| K03 Visual location profile | **PASS** | suite |
### L — Data integrity
| ID | Result | Evidence |
| --- | --- | --- |
| L01 Atomic turn commit | **PASS** | container + campaign: a real induced failure, no narration accepted, no half-written state |
| L02 State reconstruction | **PASS** | suite; campaign: state compared across 3 restarts |
| L03 Checkpoint reconstruction after restart | **PASS** | campaign + §M |
### M — Long-run
| ID | Result | Evidence |
| --- | --- | --- |
| **M01** 100-turn campaign | **PASS** | **102 accepted turns**, every scheduled operation exercised, **0 post-turn failures** (§H) |
| **M02** Restart during long campaign | **PASS** | **3 genuine process restarts** (4 process starts); everything crossed as bytes on disk |
| **M03** Long-run context stability | **PASS** | the prompt held **13,492–14,982** tokens over the last 70 turns against a 16,384 budget; window verified **102/102**; canon present **102/102** |
| **M04** Long-run memory recall | **PASS**, qualified | the planted clue was outside the history window (planted depth 1, floor 72) and reached the prompt: `m04_verdict: **recovered_through_state_only**`. Recovery was **through authoritative state**, not independent memory — §K states this distinction and does not relabel it |
### SHOULD and FUTURE
Not counted as REQUIRED. **SHOULD:** B04, D03, K04, L04 — all still pass on the
candidate (suite; browser for B04's direction toggle). **FUTURE:** K05, K06 —
deliberately not run; both need a media provider this release does not build.
## D. Backend / frontend suites
| Suite | Result |
| --- | --- |
| **Backend** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,210.6 s) |
| **Frontend** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0 — 0 errors**, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
**Every skip explained — one category, and it is the expected one.** All 17 are
environment-gated real-model tests, skipped because `AIDND_TEST_*` is
deliberately unset for the deterministic suite:
| File | Skipped | Gate |
| --- | --- | --- |
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` + `AIDND_TEST_MODEL` |
| `test_narrative_realistic.py` | 3 | same |
| `test_m11_real_window.py` | 3 | same; one needs `AIDND_TEST_WIDE_MODEL` |
| `test_provider_wiring.py` | 1 | same |
**0 xfailed.** No deficiency WP-B fixed remains parked as an expected failure —
B.1's two strict xfails became ordinary passes in B.2 and stayed that way.
**Lint warnings are the documented, unchanged set:** the pre-existing
`only-export-components` and unused-import warnings recorded at WP-C, WP-D and
WP-E. None is in a file this candidate changed relative to those packages.
## E. Production build and Docker image
| | |
| --- | --- |
| Command | the repository's documented production build (`DEVELOPMENT.md`), with `--no-cache` |
| Image | `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398` (312 MB) |
| Dependencies | **installed, not reused** — `npm ci` and `pip install --no-cache-dir` both executed in the log |
| `CACHED` steps | **2**, and both are `WORKDIR` metadata (`/build`, `/app`) — no dependency or source layer was cached |
| **Image SPA vs local build** | **file-for-file identical**: 16 files each, `diff -r` clean, combined `sha256 = ea2753ad24f61959fe084f4674911acc` on both sides |
The image is built from the candidate tree for this validation. No earlier
work-package image was reused.
## F. Offline / no-network
`tools/m11_offline.py` against **the candidate image**, `--network none`, fresh
volume. Evidence: `…/release-87a4032/offline/offline-report.json`.
**23 checks, 23 passed, 0 failed.** Including: no route to the public Internet;
no external DNS; first page load; no remote origin named; CSP served; every
referenced asset local; campaign creation; state extraction; local file import;
prompt assembly; knowledge search; a turn with no model reachable reported as a
failure with no narration accepted, the player's words kept and state unchanged;
export; import with state; no secret in the export; the media module inert with
no provider; campaigns surviving a container restart.
**No first-use download occurred**, which is what the container's absent network
makes unfalsifiable rather than merely unobserved.
## G. Browser release validation
`tools/m11_browser.py` on the candidate, production build served by FastAPI,
Firefox 155.0.1 / geckodriver 0.37.1, narrator over **trusted-LAN HTTPS with the
private CA** (`endpoint_class: trusted-LAN HTTPS`), storyteller on loopback.
`kind: release regression` — not partial, not `--only`. 693 s.
| Suite | Passed | Failed | Skipped |
| --- | --- | --- | --- |
| **M11** (the v1 release regression) | **38** | **0** | **0** |
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
| **WP-E** (control boundaries) | **10** | **0** | **0** |
| **Total** | **101** | **0** | **0** |
**A1 accounting:** 8 narrator turns, **`fits` on every one**; `turns_not_clean`
**0**; `protocol_shapes_in_narration` **0**. The verified window on this host was
4,096 (`source: loaded`); the 16,384 evidence is the long run's (§I).
WP-C's proofs are all present in the 53: Retry and alternate takes, Save Point
create/restore/Redo, state correction including a refusal shown as a refusal,
narration length reaching the prompt, failed generation and recovery, and **real
export downloads** from both the library and campaign settings, each file landing
on disk and importing into a fresh application. WP-E's ten are the rendered
boundary measurements of §Q.
## H. Integrated 100+ turn long run
One new campaign on the candidate's product code, GPU inference host, with the
owner's power/link/kernel logging running before the first turn.
Evidence: `…/release-87a4032/long-run/`.
| | |
| --- | --- |
| Narrator | **`qwen2.5:3b-instruct-16k`**, digest `21ff8cc52f375f19` |
| Embeddings | **`nomic-embed-text`**, digest `0a109f422b47e3a3` |
| Window | **16,384**, `window_verified` on **102 of 102** turns |
| Memory / summaries | **on** — 19 memories in the bank, **12 summaries** written |
| **Accepted turns** | **102** (target 100) |
| Restarts | **3** genuine process restarts, 4 process starts |
| Elapsed | 1,330 s |
| Export | 3,071,683 bytes; database 2,613,248 bytes |
| Status | `complete`; `aborted_reason` null, `failed_reason` null |
**Not a repeated-turn benchmark.** Every scheduled operation fired and is in the
timeline: 3 restarts, 1 undo, 1 undo→redo, 2 retries, 1 take selection, 2 Save
Points, 1 Save Point restore, 1 divergence, 2 state corrections, 3 knowledge
imports, memory activation, 1 deliberate failed call, 1 export, the planted clue
and the planted independent fact, and the recall probe.
### H.1 M01–M04
| | Verdict | What decides it |
| --- | --- | --- |
| **M01** 100-turn campaign | **PASS** | 102 accepted turns, each with a committed action and state document |
| **M02** Restart during long campaign | **PASS** | 3 genuine `uvicorn` restarts; everything that survived crossed as bytes on disk |
| **M03** Long-run context stability | **PASS** | prompt 13,492–14,982 tokens over the last 70 turns against a 16,384 budget; canon present 102/102; window verified 102/102 |
| **M04** Long-run memory recall | **PASS**, and qualified | the clue was planted at depth 1, the history floor reached depth 72, and it was **not** in the recent window; it reached the prompt through **authoritative state**. `m04_verdict: recovered_through_state_only`. §K keeps the distinction the criterion was written around |
### H.2 State and derived-work integrity
**0 post-turn failures across 102 turns, and no `database is locked` error.** The
only failure-shaped events in the whole timeline are the two the run creates on
purpose: the scheduled `failed_call` at turn 69 (a model name the server does not
serve — A05/L01 evidence) and the independent-fact precondition notes (§K).
Narrative-state proposals were recorded, applied or refused as designed across
the run, and no accepted narration carried an unresolved protocol block (§J).
## I. Context-window / A1 evidence
**Every turn `fits`.** Across all 102 accepted turns the accounting status was
`fits` — **0 `exceeded`, 0 `truncation_suspected`** — and the window was verified
at 16,384 on every one.
**The ten largest stored prompts**, re-counted against what the server itself
reported:
| Turn | App estimate | Server count | Difference | Window | Output reserve | Safety reserve | Observed margin | Status |
| ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | :--- |
| 86 | 14,990 | 15,005 | +15 | 16,384 | 500 | 820 | 879 | fits |
| 83 | 14,978 | 14,993 | +15 | 16,384 | 500 | 820 | 891 | fits |
| 97 | 14,965 | 14,980 | +15 | 16,384 | 500 | 820 | 904 | fits |
| 41 | 14,960 | 14,975 | +15 | 16,384 | 500 | 820 | 909 | fits |
| 66 | 14,951 | 14,966 | +15 | 16,384 | 500 | 820 | 918 | fits |
| 33 | 14,941 | 14,956 | +15 | 16,384 | 500 | 820 | 928 | fits |
| 35 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
| 77 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
| 38 | 14,927 | 14,942 | +15 | 16,384 | 500 | 820 | 942 | fits |
| 45 | 14,922 | 14,937 | +15 | 16,384 | 500 | 820 | 947 | fits |
**The reserve is preserved on every re-counted prompt.** The estimate runs
exactly **15 tokens below** the server's own count on all ten — a constant,
known offset rather than drift — and the smallest observed margin anywhere in the
run is **879 tokens**, against the documented safety reserve of 820.
**Against v1.** M11's closeout recorded a remaining margin of **23–42 tokens**.
The same measurement on this candidate is **879 at its tightest** — roughly
twenty to thirty times the headroom, which is what WP-A1 was for.
## J. Protocol-leak / A2 evidence
**Release-gate result: 0.** Across the run's **105 stored AI actions**,
`protocol_leaks` reports **0 leaking**, `example_ids` empty.
Measured separately, by the detector's own four rules:
| Shape | Count in this run |
| --- | --- |
| State/section heading with an indented entry | **0** |
| Event list (`"events"`) | **0** |
| Event-call syntax (`create_entity(`, `set_scene(`, …) | **0** |
| Hard-limit / continue-hint echo | **0** |
The browser run agrees independently: `protocol_shapes_in_narration` **0** across
its narrated turns (§G).
**The known mid-reply echo did not recur — and is still not fixed.** WP-B.1
recorded one stored reply (action 153, depth 143) where the narrator echoed the
length hint mid-reply and then continued the story, which A2's trailing cleanup
does not remove. Replaying **that stored fixture through this candidate's
detector** still flags it — 1 of 105 AI actions, matched by the hard-limit/hint
rule alone. So the residual is live (§T.2); what this release run shows is that
no equivalent shape occurred in **its** 105 replies. This report does not claim
protocol leakage is solved.
Ordinary fact and state restatement in prose was not counted: it is measurement,
not application-owned protocol, and no rule treats it as a leak.
## K. Memory / WP-B evidence
### K.1 At the final recall point
| | |
| --- | --- |
| **Created** | 19 memories in the bank; 12 summaries |
| **Retained** | the planting turn's era survived to the end of a 102-turn run |
| **Ranked** | 4 memories were selected into the prompt at the recall point |
| **Injected** | the memory section reached the prompt (14,395 tokens that turn) |
| **Independent-memory verdict** | **not demonstrated** — `precondition_failed: absent_from_later_narration` |
### K.2 The independent-fact probe, precondition by precondition
| Precondition | Held? |
| --- | --- |
| The planted turn is outside the history window (planted depth 3, floor 72) | **yes** |
| Absent from authoritative state | **yes** |
| Absent from the summary | **yes** |
| Absent from imported knowledge | **yes** |
| Absent from later narration | **no** — 6 violations, the first at turn 6 |
The narrator restated the planted fact in later narration, so the probe could not
isolate memory as the only path. The run therefore **records no independent
recovery**, and nothing here is relabelled as one.
### K.3 The two statements the contract requires, kept apart
```text
deterministic independent-memory recovery: PASS
reference-model independent-memory limitation: ACCEPTED RESIDUAL
```
- **Deterministic** (`WP-B.2` §I, and the suite on this candidate): the
`independent_full` scenario fails on v1.0.0 at creation and returns
`recovered_through_memory_independent` on the candidate, with isolation
asserted every turn and provenance resolving to the planting turn.
- **Reference model:** failed on the precondition-valid attempt in WP-B.2, and
in this release run the attempt was not precondition-valid at all. The failing
stage remains **memory creation** — the summariser's content selection.
- **No new regression.** B2.1 ranking, B2.2 eviction and B2.3 excerpt creation
all pass deterministically in the 1,723-test suite on this candidate, and the
bank behaved normally through the run (19 memories, ranked and injected). What
this run shows is the known summariser-quality limitation, not a fault in
ranking, eviction or injection.
## L. Identity diagnostic
**Scripted half — complete.** `tools/m11_identity.py --scripted` on the
candidate: 10 identity stresses, **0 signals**, verdict *no objective identity
defect detected*.
**The detector's negative control fires.** With `--inject`, the same harness on
the same candidate raises **7 signals** — `shared_display_name` once and
`duplicate_character_creation` on turns 5–10 — and preserves each turn's
evidence. A clean run therefore means something: the check is capable of
failing.
**Model-backed half — complete, on the candidate, with memory on.**
Evidence: `…/release-87a4032/identity-memory/`.
| | |
| --- | --- |
| Model | **`qwen2.5:3b-instruct-16k`** (the reference narrator) |
| Context window | **16,384** |
| Memory status | **on** — `memoryBankEnabled: true`, embedding model `nomic-embed-text` |
| Summary status | **on** — `autoSummarize: true`; **1 summary written** |
| **Identity signals** | **0** across all 10 stresses |
| **Stored protocol shapes** | **0** of 10 AI actions, by the release-gate detector |
| Fixture | accepted successfully — five entities kept distinct (`bill`, `alice`, `roger`, `john`, `office`) |
| State proposals | recorded and applied with no shared display name and no duplicate creation |
| Scripted detector self-test | still fires (7 signals under `--inject`) |
**A harness correction made during validation, and what it does not invalidate.**
The diagnostic shipped with `embedding_model=""` and the memory bank switched
off, so a release run of it would have reported a clean identity result with
memory never taking part — which is not what the gate asks for. Two lines now
read `AIDND_TEST_EMBED_MODEL` and enable the bank. **Harness-only: no product
code changed**, so no product evidence became stale; the corrected harness
repeated its own check, which is the run reported above.
**One honest observation:** with memory enabled the bank still wrote **0
memories** in this 10-beat campaign — 21 actions is enough to pass
`MEMORY_START`, but the summary pass is what ran and produced the single
summary. Memory was configured and active; it was not meaningfully *exercised*
here. The bank's real exercise is the 102-turn run (§K), which wrote 19.
**No claim about the historical root cause.** The post-M8 identity finding's
campaign was destroyed and its cause cannot be established. A clean run here is
evidence that the product does not do the things it can be blamed for on this
fixture — not a discovery of what happened then.
## M. Recovery
`tools/m11_recovery.py` against **the integrated run's own bundle**
(3,071,683 bytes), imported into a database file that never existed, in a
directory that never existed, by a second server process — so migrations ran
from nothing and this is the fresh-install path as well as the import path.
**16 checks, 16 passed, 0 failed.**
| Claim | Result |
| --- | --- |
| The destination database did not exist beforehand | PASS |
| The bundle imports into a clean directory | PASS — 209 actions in the file, 121 on the active line |
| The active transcript is not empty | PASS |
| Authoritative state came across | PASS — 6 entities, 1 fact |
| The campaign's own canon came across | PASS |
| The narration-length choice came across | PASS |
| Redo availability matches what the file said | PASS |
| Retained (undone) history came across | PASS — **88 actions retained beyond the active line** |
| Both Save Points restore | PASS — *On the ridge*, *Before the ridge* |
| Every imported class came across, with content | PASS — 3 sources |
| The moved campaign accepts a new change, unrefused | PASS |
| The bundle carries no secret | PASS |
| The moved campaign exports again, same story length | PASS |
**One thing this does not prove, stated rather than implied.** The long run
ended with its head at the tip, so the file's head was at the tip and Redo was
correctly unavailable after import. The **undone-head** case (I07) is proved by
§N's upgrade campaign, which ends on an Undo with Redo available and survives
both bundle directions, and by `test_m9_portability.py` in the suite — not by
this bundle.
## N. v1.0.0 upgrade compatibility
A campaign **built and played by the `432f041` application** in its own
worktree, then opened by the candidate. Evidence:
`…/release-87a4032/upgrade/upgrade-report.json`.
**Phase 1 — v1.0.0 builds it.** 20 turns accepted, 39 actions, canon knowledge
imported, a Save Point taken, one Undo so the head is not at the tip, narration
length chosen, memory bank and auto-summary on. It contains what §11 item 9
names: retained history, an undone head with Redo available, a Save Point,
**6 memories**, **2 summaries**, imported knowledge, a narration-length choice
and narrative state. Settings hold a **loopback placeholder**
(`http://127.0.0.1:11434/v1`) — no real hostname is in the evidence database.
**Phase 2 — the candidate opens the same file.** Nothing was copied; the
candidate's migrations ran against it.
| Field | Before | After |
| --- | --- | --- |
| transcript | 39 actions | **identical** |
| newest action (the head) | — | **identical** |
| total / can_undo / can_redo | 39 / true / **true** | **identical** |
| checkpoints | 1 | **identical** |
| narrative state | 7 keys | **identical** |
| memories | **6** | **identical** |
| summaries | **2** | **identical** |
| knowledge sources | 1 | **identical** |
| narration length | `brief` | **identical** |
| memory_bank_enabled / auto_summarize | true / true | **identical** |
| settings | loopback placeholder | **identical** |
| **schema `user_version`** | **94** | **94** |
**All 15 census fields compared identical; 0 differ.** Schema parity is exact:
v1.0.0 and the candidate both stamp `user_version` 94 with 93 migrations, so the
upgrade required no migration at all, and nothing was rewritten in passing. A
fresh-install database from the candidate carries the same 94 (§M's import ran
migrations from nothing).
**A first pass at this gate was discarded.** It played 5 turns, which is below
the memory and summary thresholds, so it compared **0 memories against 0
memories** and proved nothing about two of the criterion's required contents;
its settings also carried the live hostname rather than a placeholder. Both were
corrected and the gate was rerun — the run reported above.
## O. Bundle compatibility
| Direction | Result |
| --- | --- |
| **v1.0.0 export → imported by v1.1** | **PASS** — imported, id 2 |
| **v1.1 export → offered to v1.0.0** | **ACCEPTED** — v1.0.0 imported it, id 1 |
Both directions were **executed**, not inferred from the unchanged format
string. The format is `ai-dnd-adventure-v3` on both sides, and it did not change
during release validation.
Backward import succeeding means the compatibility question the brief raised —
whether optional or additive v1.1 evidence data would break a v1.0.0 importer —
is answered in the negative for this campaign's contents: v1.0.0 accepted the
candidate's bundle whole. No bundle-format change was made or needed.
## P. WP-D regression
Reconfirmed on the candidate; no backup-affecting product code changed after
WP-D, so its 117 MB browser measurement is not repeated (the owner's brief
permits this).
| Claim | Result |
| --- | --- |
| The completed backup copy is verified with `PRAGMA integrity_check` | **PASS** — `app/backup.py:193`, docstring at 198–201 records why the full check replaced `quick_check` |
| The corruption fixture still separates the two pragmas | **PASS** — `quick_check` → `ok`, `integrity_check` → `row 145 missing from index i_t_k` |
| Oversized export still succeeds and is delivered | **PASS** |
| Importability metadata names the effective ceiling | **PASS** — "This export is larger than this version's **20 MB** import limit (… bytes). The file was exported successfully, but this version cannot import it." |
| Normal export unchanged in content | **PASS** |
| Oversized import still refused | **PASS** — 413 naming the limit |
| Suite | **12 passed** (`tests/test_v11_d_recovery.py`) |
## Q. WP-E regression
| Claim | Result |
| --- | --- |
| Contrast audit exit code | **0** |
| All applicable control boundaries ≥ 3:1 | **PASS** — every boundary pair clears 3:1 (1.4.11) |
| Applicable text contrast still compliant | **PASS** — every text pair clears 4.5:1 (1.4.3); baselines 14.57 / 13.57 / 5.48 / 5.88 unchanged |
| Browser boundary checks | **PASS** — the 10 WP-E rows in §G |
| Focus visibility | **PASS** — M11's visible-focus check inside the 38, plus WP-E's focused-edge measurement |
| Gate tests | **11 passed** (`tests/test_v11_e_contrast.py`), including 2.99:1 failing and 3.00:1 passing |
```text
OWNER SCREENSHOT APPROVAL: APPROVED
```
Approved by the owner in the release-validation brief of 2026-09-16. **No visual
code changed during release validation**, so that approval remains valid; had any
changed, it would have been void and new screenshots would have been required.
## R. Release-shaped smoke test
The **final no-cache candidate image**, a fresh volume, published on loopback,
with the private CA installed into the container's own trust store. Evidence:
`…/release-87a4032/smoke/`.
**15 checks, 15 passed, 0 failed.**
| Claim | Result |
| --- | --- |
| The container starts | PASS |
| The application answers on loopback | PASS |
| The port is published on **loopback only** | PASS — `8000/tcp -> 127.0.0.1:…` |
| This machine's **LAN address does not serve** the application | PASS |
| The first page loads | PASS — HTTP 200 |
| The shell references no remote origin | PASS — none found |
| A CSP is served | PASS |
| **The approved HTTPS narrator verifies through its private CA** | PASS — HTTP 200 through `tlstrust.ssl_context()`, **no bypass** |
| **A public endpoint is refused** | PASS — HTTP 400 |
| A campaign is created | PASS |
| **One real narrator turn is accepted** | PASS |
| The container restarts and serves again | PASS |
| The transcript survived the restart | PASS |
| The narrative state survived the restart | PASS |
| **Firefox renders the reopened campaign** | PASS — 470 characters of story |
**A finding worth recording, and it is not a product defect.** The first attempt
failed at `PUT /api/settings` with **HTTP 400**. The cause: a `.local` name is
mDNS, a Docker container has no mDNS resolver, and `endpoints.py` correctly
refuses an endpoint whose address it cannot classify — the policy behaving
exactly as designed. The fix is to resolve the name inside the container
(`--add-host`), **not** to substitute the IP address, because the certificate is
issued for the hostname and substituting the address would have quietly bypassed
the hostname verification this test exists to prove. Harness-only; no product
code changed.
This is supplemental evidence, not a substitute for the gates above.
## S. Security / local-only review
| Claim | Evidence on the candidate |
| --- | --- |
| Served on loopback only | every harness reached the application on `127.0.0.1`; the documented container run publishes loopback |
| Endpoint policy | `app/endpoints.py` admits loopback (v4 and v6), the three RFC1918 ranges, link-local, IPv6 unique-local and CGNAT, and refuses the public Internet; a name resolving to both a private and a public address is refused |
| Inference actually used | trusted-LAN **HTTPS** with a private CA for the browser gate (§G); plain HTTP to a LAN GPU host for the long run, which `SECURITY-THREAT-MODEL.md` §83 permits and which is not A06 evidence |
| No secret in exports | offline gate, WP-D tests and the M11 suite |
| CSP served | `default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self'; frame-ancestors 'none'` |
| Other headers | `x-content-type-options: nosniff`, `referrer-policy: same-origin`, `x-frame-options: DENY` |
| No remote origin in the shell | offline gate: the page names none, and every asset is local |
`'unsafe-inline'` remains on `style-src` only, because React writes inline
`style` attributes; it is deliberately absent from `script-src`.
## T. Known residual risks
Each is classified, and none is collapsed into another category.
### T.1 WP-B reference-model memory limitation — ACCEPTED RESIDUAL
```text
deterministic independent memory: PASS
reference-model independent memory: FAIL
failing stage: memory creation — the summariser's content selection
owner decision: accepted for v1.1
```
**Status on this candidate:** unchanged, and no broader regression. The release
long run did not demonstrate independent recovery, but it also could not: its
probe failed the `absent_from_later_narration` precondition because the narrator
restated the fact (§K.2). Ranking, eviction and excerpt creation all pass
deterministically in the 1,723-test suite, and the bank worked normally through
102 turns (19 memories, ranked and injected). **Accepted residual**, carried
visibly, not a blocker.
### T.2 A2 mid-reply application-instruction echo — ACCEPTED RESIDUAL, still live
The known occurrence is WP-B.1's action 153 (depth 143): the narrator echoed the
length hint **mid-reply** and then continued the story, which A2's trailing
cleanup does not remove.
**Reproduced on this candidate.** Replaying that stored fixture through the
release-gate detector still flags it — **1 of 105** AI actions, matched by the
hard-limit/hint rule alone. The extractor still leaves it. This report therefore
does **not** claim protocol leakage is solved.
**But it did not recur in release evidence.** This run's own 105 stored replies
leak **0** (§J), and the browser run's narrated turns leak 0. The brief's
stop-and-report condition — *an equivalent shape occurring in this final run* —
was **not** triggered, so validation continues. No broader sanitizer was written;
that remains for owner review.
### T.3 Doubled full stop in memory-search scene text — ACCEPTED RESIDUAL
`"rain outside.."` when the state's scene summary already ends in punctuation.
Only the embedding query sees it; the effect is one stray token. Not changed
during release validation, deliberately: cleanliness is not a reason to alter
product behaviour after evidence is taken.
### T.4 K1 — "Correct" on an Important Facts row is always refused
**Reproduced, unchanged on this candidate**, deterministically and without a
browser: Correct on a Characters row (`subject='mara'`) applies, **201**; Correct
on an Important Facts row (`subject='f1'`) is refused, **400 — "add_fact names
subject='f1', which does not exist."** The cause is frontend-side: the panel
sends the row key as `add_fact.subject`, and the validator checks `subject` as an
entity reference.
**Classification: v1.2 backlog, not a release blocker.** It blocks no v1
REQUIRED test — C04 passes through the working correction paths (§C) — and WP-C's
State-panel correction coverage passes. It is a narrow bug with an obvious fix
(offer Correct only against entities, or send facts without a subject), but
fixing it during release validation would change product code after the evidence
above was taken, which the brief forbids without a product brief. **Left for the
owner.**
### T.5 Import ceiling and scheduled backups — INTENTIONAL, not unfinished work
The 20 MB import ceiling is deliberate and unchanged; WP-D made it honest rather
than raising it. Scheduled backups remain unbuilt by design. Neither is a
residual defect.
### T.6 Harness corrections made during validation — no product evidence invalidated
Three, all harness-only, each named where it occurred: the identity diagnostic
did not enable memory (§L); the smoke test needed hostname resolution inside the
container (§R); and the first upgrade campaign was too short to write memories
and carried a live hostname (§N). **No product code changed at any point during
release validation**, so no black-box or long-run evidence became stale. Each
corrected harness repeated its own affected check.
## U. Deferred v1.2 / future work
| Item | Why it is deferred |
| --- | --- |
| Raising the import ceiling, or a streaming import | The plan assigns it to v1.2; WP-D's scope was honesty about the limit, not the limit |
| Scheduled backups; a restore button | Explicitly out of WP-D's scope |
| K1's Correct-on-a-fact-row fix (§T.4) | A narrow frontend bug needing a product brief |
| A broader mid-reply protocol sanitizer (§T.2) | Needs owner review; A2's cleanup is deliberately trailing-only |
| The doubled full stop (§T.3) | Cosmetic, embedding-query only |
| K05 generate local image, K06 multi-turn video | FUTURE tests; both need a media provider this release does not build |
| Reference-model independent memory (§T.1) | Needs a stronger summariser or a different creation strategy — a v1.2 investigation, not a v1.1 fix |
## V. Documentation changes
Current documents were brought up to date. **No historical milestone report was
rewritten, and no failed WP-B real-model evidence was turned into success.**
| Document | Change |
| --- | --- |
| `README.md` | Status now says v1.0.0 **remains** the released version, that all six v1.1 packages are complete and accepted, that release validation passed on candidate `87a4032`, that WP-B ships with a documented limitation, and that **no `v1.1.0` tag exists and `main` is unchanged`**. Also corrected a stale figure: the schema is versioned at **94**, not "92 and counting" |
| `planning/V1.1-PLAN.md` | WP-D/WP-E recorded as signed `87a4032`; the release-validation outcome summarised with its residuals; the three owner events named as still outstanding |
| `planning/VERSION.md` | Same status correction, plus a new revision entry for the closeout |
| `planning/README.md` | Current-state paragraph rewritten for the same facts |
| `planning/reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL: PENDING` → **`APPROVED`**, with the source (the release-validation brief) and date recorded, and a note that the signed commit predated the review |
| `planning/reports/v1.1/V1.1-RELEASE-REPORT.md` | **New** — this document |
| `DEVELOPMENT.md` | Unchanged: its WP-C/WP-D sections already describe the candidate as built |
**New tools committed with this closeout** (harness only, no product code):
`backend/tools/v11_upgrade_check.py` (Gate 9) and
`backend/tools/v11_release_smoke.py` (§R), plus two narrow corrections to
`backend/tools/m11_identity.py` (read `AIDND_TEST_EMBED_MODEL`; enable the memory
bank) so the diagnostic can run with memory on.
## W. Final release decision
**The question this validation set out to answer:** does candidate `87a4032`
preserve the complete v1 contract and satisfy every accepted v1.1 package on one
integrated release tree?
| Gate | Result |
| --- | --- |
| 1 Package acceptance | **PASS** — six packages accepted; WP-B's qualification carried whole (§B.1) |
| 2 v1 acceptance contract | **PASS** — 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified (§C) |
| 3 Suites, lint, build, Docker | **PASS** — 1,723 / 175 / 0 errors; image SPA file-for-file identical (§D, §E) |
| 4 Offline / no-network | **PASS** — 23/23 on the candidate image (§F) |
| 5 Browser release run | **PASS** — 101/0/0 over trusted-LAN HTTPS, every turn `fits` (§G) |
| 6 Integrated long run | **PASS** — 102 turns, M01–M04, 0 post-turn failures (§H) |
| 7 Identity diagnostic | **PASS** — 0 signals, 0 protocol shapes, memory on (§L) |
| 8 Recovery | **PASS** — 16/16 on the run's own bundle (§M) |
| 9 v1.0.0 upgrade | **PASS** — 15/15 identical, schema parity at 94 (§N) |
| 10 WP-D regression | **PASS** (§P) |
| 11 WP-E regression | **PASS**, screenshots approved (§Q) |
| Bundle compatibility | **Both directions execute and import** (§O) |
| Release smoke | **PASS** — 15/15 from the shipped image (§R) |
**A1** holds its reserve on every re-counted prompt, with a smallest margin of
**879 tokens** against v1's 23–42. **A2**'s release-gate leak count is **0**.
**WP-B**'s deterministic independent-memory recovery passes and its
reference-model limitation remains an accepted, documented residual — stated in
§K.3 in both halves, never shortened to "WP-B passed".
No product code was changed at any point during release validation, so no
evidence was invalidated. Three harness corrections were made and each corrected
harness repeated its own check (§T.6).
```text
V1.1 RELEASE VALIDATION:
PASS
```
### What this decision is not
These are separate, and only the first is done:
```text
WP-A-E accepted: YES
release validation passed: YES
release candidate prepared: YES (87a4032, with this report)
release commit signed: NO
main updated to v1.1: NO
v1.1.0 tagged: NO
```
The closeout changes are **staged and uncommitted**. Nothing was committed,
pushed, merged or tagged by this validation. The owner's next decision is to
review this evidence, resolve anything they disagree with, then sign the v1.1
release commit and publish `v1.1.0`.
File diff suppressed because it is too large Load Diff
+623
View File
@@ -0,0 +1,623 @@
# v1.1 WP-B.1 — Independent Long-Term Memory Retention Diagnostic
**Status:** COMPLETE. Diagnostic only: no memory behaviour was changed. The final decision is in §S.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `d63804f22ecbaed80741241a154cdaef82f7b2ed` — *v1.1: harden context window and narrator protocol boundary* (WP-A1/A2), signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its parents | `ac465ed` (planning v4.1, signed) and `432f041` (the signed v1.0.0 release commit; tag `v1.0.0`; `main`) |
| Working tree at start | clean |
| `git diff --stat v1.0.0..HEAD` | 27 files, +4,880 / −88: the planning commit and WP-A1/A2 |
| Comparison baseline | `432f041` (v1.0.0), run from a throwaway worktree (§D) |
`app/memorybank.py`, `app/summaries.py`, `app/context/lineage.py`, `app/tree.py` and
`app/vectors.py` are unchanged between `v1.0.0` and HEAD. A1 and A2 did not touch
the memory pipeline, apart from adding `accounting` to `attempts.ATTEMPT_KEYS`.
---
## B. Existing memory pipeline
Answered from the code at HEAD. Nothing was changed.
| # | Question | Answer |
| --- | --- | --- |
| 1 | How are memories generated? | `memorybank.run_post_turn`, a fire-and-forget task after each accepted turn (`schedule_post_turn`), runs `_create_due_memories` when the campaign has `auto_summarize`. It writes one memory per block of `MEMORY_INTERVAL` = 6 story actions past the memory cursor. It starts once the story has `MEMORY_START` = 12 actions, and only when `SETTLE_SLACK` = 1 action sits past the block. At most `MAX_MEMORIES_PER_RUN` = 5 memories are written per run. |
| 2 | What range does a memory cover? | The block's first and last action depths: `Memory.source_start`, `Memory.source_end`. |
| 3 | How are source depth and lineage stored? | `tree.attach_memory` sets `Memory.branch_id` and `Memory.depth` from the block's **last** node, so a memory is visible exactly on paths that contain that node. |
| 4 | Memory text length limit | Prompt-only: `MEMORY_MAX_WORDS` = 50 in `MEMORY_SYSTEM_PROMPT` ("1-2 plain sentences"). Nothing truncates the stored text. |
| 5 | What input does the summariser receive? | `summarize_block`: a cast brief (`cast_brief`, which reads story cards, persona and plot essentials), then `"Story excerpt:\n\n{excerpt}\n\nMemory:"`. The excerpt is the block's action texts joined by blank lines. |
| 6 | Where does the 2,000-token truncation happen? | `summarize_block`: `excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)`, with `MEMORY_EXCERPT_TOKENS` = 2,000. It keeps the block's **last** 2,000 `cl100k_base` tokens. The cast brief is matched against the untruncated block. |
| 7 | How does capacity and eviction work? | `_evict_over_capacity`, at the end of every `run_post_turn`. It counts non-forgotten memories in the whole adventure (not lineage-scoped). Above `settings.memory_bank_capacity` (default 80), it marks `forgotten = True` on the overflow unpinned memories, ordered by `coalesce(last_used_at, created_at)` ascending, then `use_count`. Forgotten rows are kept. |
| 8 | How is `last_accessed` updated? | The field is `Memory.last_used_at`. `record_use` sets it, and increments `use_count`, for every memory in the turn's `memories.used`. That write is part of the turn's single commit (§O.7 of the M11 report). Dry runs never count. |
| 9 | How are memories ranked? | `retrieve_memories`. The query is the text of the newest `RETRIEVAL_WINDOW_ACTIONS` = 4 actions, truncated to the last `RETRIEVAL_WINDOW_TOKENS` = 600 tokens. It is embedded with the configured embedding model. Every eligible memory is scored by `vectors.cosine`. Pinned memories are taken first, the rest fill `memory_top_k` (default 5) in score order, and `_drop_redundant` skips a candidate at cosine ≥ 0.93 to one already chosen, within the same authority. **There is no lexical term, recency term, importance term or similarity floor** (CONTEXT-AND-MEMORY §20, as implemented). |
| 10 | How do pins affect ranking and eviction? | Ranking: a pinned memory is always selected, and counts toward `memory_top_k`. Eviction: pinned memories are never evicted, and if every active memory is pinned, capacity is exceeded. |
| 11 | How does the active-lineage clause filter memories? | `lineage.path_of(db, adventure).clause(models.Memory)` matches `(branch_id, depth)` against the head's path entries, capped at each fork depth. Retrieval also requires `forgotten = False` and `embedded = True`. |
| 12 | How do retrieved memories enter `build_context`? | Through the `memory_bank` argument. The builder renders `"Memories from earlier in the story. Lines marked [inferred] are interpretation, not established fact — do not treat them as settled truth:"` plus one `- [inferred]? text` line per memory, as section `used_memories`. It is a live section priced into protected context, placed after history and before `narrative_state`. |
| 13 | Where is the provenance recorded? | `context_snapshot["memories"]`: the whole retrieval result, meaning `used` (id, text, similarity, pinned, authority, source range), `considered` and `suppressed`. It is stored per turn, so a past turn's selection is inspectable. |
| 14 | How does memory survive export and import? | `bundle._exported_memory` carries text, pinned, forgotten, `sourceStart`, `sourceEnd`, `useCount`, authority, branch and depth. **It does not carry the vector, `last_used_at` or `created_at`.** On import the memories are re-embedded by the post-turn pass, and their recency restarts from the import. |
**A finding from the inspection itself.** The v1 long-run harness (`tools/m11_long_run.py`)
planted M04's clue as an accepted **state correction** (`CLUE_FACT`, via
`add_fact`). The player turn mentioning it uses the sentinel code, but the fact the
recall looks for was established in state. Memories are written from **story
text**. So every v1 M04 recovery could only ever run through state, or through the
narrator restating state, and none could have tested memory on its own. That is why
§P risk 5 of the M11 report could say "no run showed memory keeping a planted fact"
without any run having given memory the chance.
---
## C. Diagnostic design
### C.1 What is built
| File | Kind | Purpose |
| --- | --- | --- |
| `backend/tools/memory_diagnostic.py` | tool, new | See below: the fact spec, isolation checks, four-stage diagnosis, deterministic stubs and scenario runner. |
| `backend/tools/v11_b1_memory.py` | tool, new | CLI. `scenarios` runs the deterministic campaigns against an isolated database and writes JSON. `diagnose` runs the four stages against a **copy** of a finished real campaign's database. |
| `backend/tools/m11_long_run.py` | tool, extended | `--independent-fact`: see §C.3. |
| `backend/tests/test_v11_b1_memory_diagnostic.py` | tests, new | The deterministic diagnostic. |
| `backend/tests/test_v11_b1_long_run_verdict.py` | tests, new | The new long-run verdict. |
What `tools/memory_diagnostic.py` holds:
- **`Fact`:** a planted fact, with whole-word carry and leak matching.
- **`isolation()`:** every non-memory layer checked.
- **`diagnose()`:** the four stages and the verdict.
- **`rank_bank()`:** production ranking recomputed for every eligible memory.
- **The stubs:** `BestCaseSummariser`, `ConceptEmbedder` and `ScriptNarrator`.
- **`run_scenario()`:** a campaign played through the real turn route.
**No application file is changed.** No column, table, migration or setting is
added. The diagnostic's extra fields are computed at report time.
### C.2 How each stage is judged
| Stage | Judged from | Output (real field names where they exist) |
| --- | --- | --- |
| **created** | Memories on the active lineage with `source_start ≤ plant_depth ≤ source_end` whose text carries F (both "sundial" and "teapot"). For every covering memory, the block is re-read (`memorybank.source_block`) and cut exactly as `summarize_block` does, to report whether F was in the block and whether it was in **the excerpt the summariser saw**. | `memory_id`, `source_start`, `source_end`, `memory_text`, `covering_memories[].{block_tokens, fact_in_block, fact_in_summariser_excerpt}` |
| **retained** | That memory's row, plus the eviction order production would use (the same `ORDER BY`, read-only) | `forgotten`, `pinned`, `embedded`, `on_active_lineage`, `use_count`, `last_used_at`, `created_at`, `active_memories`, `memory_bank_capacity`, `eviction_position`, `reason` |
| **ranked** | `rank_bank`: the same catalogue clause, cosine, pin rule and `memorybank._drop_redundant`, over **every** eligible memory, with the recall turn's own query (`history.tail(4, exclude=recall AI node)`, cut to 600 tokens). It is checked against the recall turn's stored `memories.used` (`replica_matches_stored_selection`). | `semantic_score` (`similarity`), `lexical_score` (always `None`: no such term exists), `final_score`, `rank`, `of`, `top_k_cutoff`, `selected`, `suppressed_as_duplicate_of`, `query` |
| **injected** | The recall turn's stored `context_snapshot`: `memories.used` names the memory, and its text is in the `used_memories` section | `context_component`, `in_stored_memories_used`, `text_in_section`, `token_count` |
Verdicts, in order: `not_created`, `created_but_evicted`, `retained_but_not_ranked`,
`ranked_but_not_selected`, `selected_but_not_injected`, `injected`. The fifth is
added to the brief's list, so that "the retrieval picked it" and "the narrator was
shown it" stay distinguishable.
### C.3 Isolation (precondition) checks
A result counts only if every check holds.
| Check | How |
| --- | --- |
| `state_document` | `adventure.narrative_state`: entities, facts, relationships, threads, scene and possessions, plus the whole document |
| `state_snapshots` | `narrative_state_after` of every node on the active lineage |
| `later_narration` | every AI turn deeper than the planting block's end, and before the recall turn |
| `summary` | `summaries.current` and the recall prompt's `story_summary` section |
| `knowledge` | every `KnowledgeSource.content`, and the recall prompt's imported-knowledge sections |
| `recent_history` | the recall prompt's `history.floor_depth` is greater than the planting depth, and F is not in the `history` or `recent_history` sections |
| `state_section` | the recall prompt's `narrative_state` section |
The negative control for the precondition itself is
`test_the_isolation_check_fails_when_another_layer_carries_the_fact`.
### C.4 The deterministic stubs, and what they model
- **`BestCaseSummariser`.** An ideal memory writer. A memory keeps every sentence of
its excerpt that carries a planted fact, plus one sentence naming the block's own
place so that memories differ. Summary updates never mention a planted fact. **A
creation failure under this stub is the application's, not a model's.**
- **`ConceptEmbedder`.** A 96-dimension deterministic embedding. Words in a small
concept table ("sundial", "dial", "hour", "clock", …) share a dimension, other
words are hashed, and the vector is normalised. It models a paraphrase landing
near the original. **It says nothing about `nomic-embed-text`.**
- **`ScriptNarrator`.** Narration that names only filler places and never a planted
fact, with an empty state block, so state never records F.
Scenarios are played through the real `POST /actions` route. Automatic post-turn
scheduling is replaced by an explicit `run_post_turn` settle after every turn, so
eviction happens at a known turn. Depths: the opening is 0, turn *n*'s player
action is 2n−1 and its reply 2n. The planted fact is a `story` action.
### C.4.1 Scenarios
| Scenario | Turns | Capacity | top_k | History budget | Prose per reply | Planted at |
| --- | --- | --- | --- | --- | --- | --- |
| `independent_default` | 52 (recall at depth 106) | 80 | 5 | 4,096 | ~60 words | depth 1 |
| `past_capacity` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1 |
| `past_capacity_pinned` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1; first other memory pinned |
| `past_capacity_low_top_k` | 52 | 8 | 2 | 4,096 | ~60 words | depth 1 |
| `long_block_fact_early` | 10 | 80 | 5 | 16,384 | ~850 words | depth 1, early in a long block |
| `long_block_fact_late` | 10 | 80 | 5 | 16,384 | ~850 words | depth 5, late in the same-sized block |
| `lineage_control` | 52 | 80 | 5 | 4,096 | ~60 words | F at depth 1; G on line A, then Undo × 9 and divergence at turn 30; Save Points before and after G |
### C.5 The long-run verdict
`tools/m11_long_run.py --independent-fact` plants a second, **story-only** fact
at depth 3, right after M04's own plant, and never corrects it into state. Every
existing M04 behaviour and verdict is unchanged.
**Every accepted turn** records the first failure of:
- `absent_from_state`;
- `absent_from_summary`;
- `absent_from_later_narration`, meaning narration deeper than the planting depth
plus 6.
The plant and those first failures survive `--resume`.
**At recall** a dedicated question is played. `_independent_recall` reads the
recall turn's stored context and the campaign database (read-only), and
`_independent_memory_verdict` returns one of:
- `recovered_through_memory_independent`, only when every one of
`planted_turn_outside_history`, `absent_from_state`, `absent_from_summary`,
`absent_from_knowledge` and `absent_from_later_narration` holds, and a memory
covering the planting turn carries the fact and was injected;
- `precondition_failed:<name>` or `precondition_unknown:<name>`, never a recovery;
- `not_recovered:not_created`, `not_recovered:evicted` or
`not_recovered:not_injected`.
Ranking is recomputed afterwards by `tools/v11_b1_memory.py diagnose`, against a
copy of the run's database.
---
## D. v1.0.0 baseline
**How it was run.**
- A throwaway worktree was checked out at `432f041` (`git describe`: `v1.0.0`).
- Only the three B.1 files were copied in: `tools/memory_diagnostic.py`,
`tools/v11_b1_memory.py` and `tests/test_v11_b1_memory_diagnostic.py`.
- The imported `app` was confirmed to come from the worktree.
- The worktree was removed afterwards, and the `v1.0.0` tag and commit were not
touched.
- `app/memorybank.py` is byte-identical between `v1.0.0` and HEAD.
- Evidence: `$HOME/v11-evidence/b1/v100/`, with HEAD's in
`$HOME/v11-evidence/b1/head/`.
| | v1.0.0 (`432f041`) | HEAD (`d63804f` + B.1 files) |
| --- | --- | --- |
| `test_v11_b1_memory_diagnostic.py` | **24 passed, 2 xfailed (strict)** | 24 passed, 2 xfailed (strict) |
| `independent_default` | `injected` | `injected` |
| `past_capacity` (capacity 6, top_k 5) | **`created_but_evicted`** | `created_but_evicted` |
| `past_capacity_pinned` | `created_but_evicted` (the pinned memory kept) | same |
| `past_capacity_low_top_k` (capacity 8, top_k 2) | **`created_but_evicted`** | `created_but_evicted` |
| `long_block_fact_early` | **`not_created`** | `not_created` |
| `long_block_fact_late` | `injected` | `injected` |
| `lineage_control` | `injected`; G never injected after the divergence | same |
**The two independent-retention acceptance criteria fail on v1.0.0, and each names
the stage.**
```text
criterion: an early fact is recalled from memory past capacity
created: yes
retained: no
FAILURE STAGE: retention (capacity eviction)
criterion: a fact early in a long block is remembered
created: no (the fact was in the block, not in the summariser's excerpt)
FAILURE STAGE: creation (input truncation)
```
Under the best-case summariser, with blocks shorter than 2,000 tokens and a bank
under capacity, v1.0.0 carries the fact all the way to injection. The two failure
stages above are therefore **application mechanisms**, reached under conditions
a long campaign meets:
- a block of long narration;
- more memories than `memory_bank_capacity`.
Which of them a real campaign meets first is §K's question.
---
## E. Creation results
`long_block_fact_early` and `long_block_fact_late` use the same block geometry: the
opening (depth 0) and turns 1-3. Each reply is about 850 words, and the block
`source_start` 0 … `source_end` 5 is **2,079 tokens**, 79 over
`MEMORY_EXCERPT_TOKENS`.
| Shape | Planted at | Fact in block | Fact in summariser excerpt | Memory written | Verdict |
| --- | --- | --- | --- | --- | --- |
| F early in the block | depth 1 | yes | **no** | "The travellers spent time at the ferry landing." | **`not_created`** |
| F late in the same-sized block | depth 5 | yes | yes | "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." | `injected` |
`summarize_block` keeps the **last** 2,000 tokens. An early-block fact is cut off
before the summariser reads it, even by an overflow of only 79 tokens. No summariser
quality can recover what it was never given.
`independent_default`: 60-word replies, a 4,096 budget, and a block under 2,000
tokens. Memory 1 covers depths 0-5, and its text carries F. Both the block and the
excerpt contain F.
---
## F. Retention / capacity results
Default: the bank holds 17 active memories at depth 106, against capacity 80. F's
memory is retained, has `use_count` 16, and sits at eviction position 14 of 17.
**Past capacity** (the brief's capacity test). 17 memories are written over 52 turns.
| Scenario | F's uses before eviction | F's last retrieval | First eviction | F evicted | F first evicted? | Created and evicted in the same pass | Pinned kept |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `past_capacity` (6 / top_k 5) | 12 | turn 18 | turn 21, memory 1 | **turn 21** | **yes** | none | — |
| `past_capacity_pinned` (6 / top_k 5, memory 2 pinned) | 12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | **memory 2 never evicted** |
| `past_capacity_low_top_k` (8 / top_k 2) | 4 | turn 24 | turn 27, memory 3 | **turn 36** (4th eviction) | no | none | — |
**What the trace shows about the rule.**
- **"Never retrieved" is not the mechanism.** F was retrieved, 12 or 4 times, while
it still ranked in the top-k for the recent-narration query.
- **Eviction orders by `coalesce(last_used_at, created_at)`.** An early fact that
recent narration never mentions stops being retrieved once newer memories fill
the top-k. It then ages out:
- at capacity 6, it is gone 3 turns after its last use, as the first eviction;
- at capacity 8 with `top_k` 2, 12 turns after its last use, as the fourth.
- **Recall itself cannot rescue it.** Retrieval is driven by recent narration, and
a fact nobody mentions is exactly the one that loses recency.
- **Pinned memories stay protected** (`past_capacity_pinned`).
- **No memory is evicted by the pass that created it.** The frozen-bank regression
fix still holds in all three scenarios.
---
## G. Ranking results
The `independent_default` recall turn is at depth 106. Its query is the newest 4
actions, cut to 600 tokens: three narration turns and the one-line question.
| | Value |
| --- | --- |
| F's `similarity` (semantic score) | **0.241** |
| Lexical score | none: memory ranking has no lexical term |
| Pin effect | none (not pinned) |
| Final rank | **2 of 17** |
| `memory_top_k` cutoff | 5 |
| Selected | yes |
| Replica agrees with the turn's stored `memories.used` | **yes** |
The same memory against three stand-alone queries, `ConceptEmbedder`:
| Query | Rank | Similarity | Selected |
| --- | --- | --- | --- |
| direct: "I ask Mara where she hid the amber sundial." | 1 of 17 | 0.708 | yes |
| paraphrase: "… the little brass dial that tells the hour." | 1 of 17 | 0.636 | yes |
| unrelated: "… what rope costs at the landing this season." | 5 of 17 | **0.064** | **yes** |
Two diagnostic observations. Neither is a failure in this scenario.
1. **The production query dilutes the question.** The question alone scores 0.708;
inside the four-action window it scores 0.241. F survives at rank 2 of 17. In a
bank where more memories share the recent narration's vocabulary, the same
dilution would push it past `top_k`.
2. **There is no relevance floor.** With more memories than `memory_top_k`, five are
injected whatever their similarity. An unrelated query still injects F at 0.064.
This matters to F only in the other direction: it can ride along even when it is
not relevant.
---
## H. Injection results
`independent_default` injects F: `memories.used` in the recall turn's stored snapshot
names memory 1, its text is in the `used_memories` section, and that section is 97
tokens. The same holds in `long_block_fact_late` and `lineage_control`.
**`selected_but_not_injected` never occurred**: whatever retrieval selected, the
builder rendered.
---
## I. Lineage negative control
`lineage_control`:
- G is planted on line A at depth 41, turn 21.
- G's memory 7 is written.
- A Save Point is placed on line A after G.
- Undo ×9, then divergent writing at turn 30. The last action before the
divergence has id 59.
| Assertion | Result |
| --- | --- |
| G's memory stays stored | **yes** (memory 7 present) |
| G is not eligible on the active lineage | **yes** (the path clause returns nothing) |
| G is never injected after the divergence | **yes** (no turn with id > 59 names it or carries its text) |
| `memories.used` does not report it after the divergence | **yes** |
| Returning to line A (Save Point restore) makes it eligible again | **yes** (memory 7 eligible) |
| F on the active line is unaffected | `injected`, isolation holds |
Before the divergence, G's memory was legitimately used on line A (turn 22). The
first version of this check counted that as a leak. That was a defect in the
diagnostic, and the scan now starts after the divergence.
No lineage code was touched. `test_m11_leakage.py`, 14 tests, passes unchanged (§P).
---
## J. Authority negative control
`test_a_memory_that_contradicts_state_loses_and_changes_nothing` sets up the
conflict like this:
- a state correction adds the fact "the tavern lamp is lit";
- a hand-written memory says "The tavern lamp was never lit that night.";
- the memory is pinned, so it is injected;
- a turn is played with an empty proposal.
| Assertion | Result |
| --- | --- |
| The narrative state document is unchanged by retrieval and the turn | **yes** (identical before and after) |
| The state fact is in the prompt's `narrative_state` section | yes |
| The memory is in `used_memories`, under "Memories from earlier in the story …" | yes, framed as historical and non-canon context |
| `narrative_state` comes after `used_memories`, so state is read last and settles the conflict | yes |
F07 semantics are unchanged: memory never writes state.
---
## K. Real-model attempts
**Setup.**
- **Command:** `tools/m11_long_run.py --turns 100 --independent-fact`.
- **Host and models:** the GPU inference host (Ollama 0.34.0), `qwen2.5:3b-instruct-16k` at a verified 16,384 window, embeddings by `nomic-embed-text:latest`.
- **Memory:** the memory bank and auto-summarise on.
- **Harness settings:** `memory_top_k` 4 and `context_token_budget` 16,384.
- **Logging:** the owner's power, link and kernel logging was running on the host before the run started.
- **Permission:** inference was used only with the owner's explicit approval.
**Evidence:**
- `$HOME/v11-evidence/b1/real-1/`: `summary.json`, `recall-independent.json`, `timeline.jsonl` and `campaign.db`;
- `$HOME/v11-evidence/b1/real-1-diagnosis/diagnosis.json`: the four stages, recomputed on a copy of the database with the same embedding model.
### K.1 Attempt 1 — PRECONDITION FAILED
| | |
| --- | --- |
| Run status | `complete`: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 active memories (19 eligible on the recall line), 1,308 s elapsed |
| Window / accounting | verified 16,384 on every turn; `fits` |
| M04 verdict (unchanged) | `recovered_through_state_only` |
| Independent fact | planted at depth 3 by the player turn *"I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one."* It was never corrected into state |
| **Verdict** | **`precondition_failed:absent_from_summary`** (the verdict names the first failed precondition in its fixed order) |
| Precondition | Result | First failure |
| --- | --- | --- |
| planted turn outside recent history | **held**: the history window started at depth 66 | — |
| absent from authoritative state | **held**: the document, every node's snapshot and the recall prompt's state section | — |
| absent from imported knowledge | **held**: 3 sources | — |
| **absent from later narration** | **failed**: the narrator mentions the fact at depths 6, 8, 10, 14, 16, 18, 20, 22, 24, 26 and later | accepted turn **5** |
| **absent from the active summary** | **failed** | accepted turn **9** |
| fact text in the recall prompt's history sections | **present** (restated narration inside the window) | — |
**Why it did not qualify.** The 3B narrator took the planted detail up as a motif
and restated it for the rest of the campaign ("Mara's amber sundial flickered
softly, a silent reminder of their shared history"). The summariser, whose prompt
asks it to "preserve important established facts", folded it into the running
summary. Neither is a defect in the harness. They are the other layers doing what
they do with a salient fact, which is exactly what makes a clean memory-only
measurement hard to obtain with a real narrator.
**The four stages, diagnosed anyway.** They are not evidence for independent
retention, but they are evidence for the mechanism.
| Stage | Result |
| --- | --- |
| **created** | **yes.** Memory 1 covers depths 0-5. The block is 688 tokens, so there was no truncation: the fact was in the block and in the summariser's excerpt. The memory text: *"Aldric taps the silver key against his chest … Mara slipped the amber sundial into the cracked teapot, her fingers tightening on the silver key. …"* |
| **retained** | **yes.** Not forgotten, `use_count` 53, 33 active memories against capacity 80, eviction position 28 |
| **ranked** | **no.** Eligible and embedded. For the recall turn's own query it scored `similarity` **0.866** and ranked **10 of 19**, against `top_k_cutoff` 4. It was not selected and was not suppressed as a duplicate. The replica matched the stored selection |
| **injected** | not reached. The recall turn's `memories.used` = [30, 29, 31, 5] |
**What the narrator was given instead.** Five **later** memories also carry the
fact, all written from narration that restated it: memories 12 (depths 66-71), 28
(81-86), 30 (93-98), 17 (96-101) and 18 (102-107). Memory 30 was injected at
recall. The fact reached the narrator through memory, but through a
**restatement's** memory, not the planting-era one.
**The same memory 1 against reference queries** (`nomic-embed-text`):
| Query | Rank | Similarity | Selected |
| --- | --- | --- | --- |
| the recall turn's production query | 10 / 19 | 0.866 | no |
| paraphrase: "… the little brass dial that tells the hour" | 6 / 19 | 0.593 | no |
| unrelated: "… what rope costs at the landing this season" | 12 / 19 | 0.435 | no |
Two properties of the real embedder matter to B.2:
1. **A high floor.** An unrelated query still scores 0.44.
2. **A crowded top.** The bank is full of near-identical "Aldric and Mara step out,
the silver key's weight in his pocket" memories, so 0.866 was not enough to reach
the top 4.
*A side finding outside B.1's scope.* The export's protocol-leak counter flags
**1 of 105** stored turns: action 153, depth 143. Mid-reply, the narrator echoed the
length hint and the state reminder, with an `Events: [...]` line, and then continued
the story. WP-A2's cleanup rules act only at the end of a reply, so an instruction
echo with story after it stays. It is recorded here for the v1.1 backlog. B.1 did not
touch it.
---
## L. First failing stage
**Deterministic, with a best-case summariser and a concept embedder, identical on
v1.0.0 and HEAD:**
| Condition | First failing stage |
| --- | --- |
| a bank under capacity, blocks under 2,000 tokens | **none**: created, retained, ranked (2 of 17) and injected |
| a bank past `memory_bank_capacity` | **retention**: evicted by least-recently-used order after recent narration stops retrieving it (§F) |
| the fact early in a block over 2,000 tokens | **creation**: the fact is in the block, but not in the summariser's excerpt (§E) |
**Real model, attempt 1** (isolation not met, mechanism only): created yes, retained
yes, **ranking failed first**. The planting-era memory scored 0.866 but ranked 10th of
19 behind later, near-identical memories, and outside `top_k` 4.
**Across all the evidence, the first stage an early fact fails at a real campaign's
length is ranking.** In both the deterministic default and the real run, the memory
exists and is retained at 100 turns, where the bank is below capacity. What decides
whether the narrator is shown it is its rank against a recent-narration query in a
bank of similar memories.
- **Retention (eviction)** is a second, later failure. It is certain once a campaign
outgrows capacity: about 480 actions at the defaults.
- **Creation (truncation)** is a third, conditional one. It needs blocks longer than
2,000 tokens, which the 16,384 window with 500-token replies did not produce
(688 tokens).
---
## M. Evidence for likely root cause
1. **The retrieval query is recent narration, not the question** (§B.9, §G). The
query is the last 4 actions cut to 600 tokens, so a one-line recall question is
outweighed by three turns of prose. The effect is deterministic: the question's
own similarity of 0.708 fell to 0.241 in the production query. In the real run,
recent prose about the same tavern, key and people made every memory look
similar, and ten ranked above the planting-era one.
2. **Ranking has no term that favours the planting-era record** (§B.9;
CONTEXT-AND-MEMORY §20). There is cosine only: no lexical match on the question's
rare terms ("sundial", "teapot"), no importance, and no preference for the earliest
or a coverage-distinct source. Later restatement memories carry the same words in
more familiar company, and outrank the original.
3. **Eviction is purely least-recently-used** (§B.7, §F). Retrieval is driven by
recent narration, so exactly the facts nothing recent mentions lose recency and go
first. Recall itself cannot rescue them, because they are no longer retrieved.
4. **The summariser reads only the last 2,000 tokens of a block** (§B.6, §E). This is
proven deterministically. It did not bite at the real run's block sizes.
5. **v1's M04 never tested memory** (§B finding). The clue was planted as state, so
the long-standing "memory does not keep the fact" observation was never a
measurement of memory.
---
## N. What B.2 is allowed to change
B.2 is allowed only the smallest changes the evidence supports, one mechanism at a
time, each with a failing test first. In order of the evidence:
1. **Ranking** (the first failing stage). Candidates:
- build the retrieval query so the player's newest input is not drowned out, for
example by giving the newest player action its own weight or its own query;
- and/or add one inspectable ranking term from CONTEXT-AND-MEMORY §20, most
directly a lexical match on the query's rare terms.
Either must keep `replica_matches_stored_selection` meaningful: a
diagnostic-visible score, recorded in `memories.used`.
2. **Eviction** (the certain second failure). Stop least-recently-used eviction from
discarding a never-again-retrieved early memory first. For example, weight eviction
by coverage, keeping the only memory of a story range, or by age, instead of recency
alone. The frozen-bank protection must be kept.
3. **Creation** (conditional). Choose the summariser's excerpt so that a fact early in
a long block is not cut. For example, the head and tail, or the whole block up to a
larger bound.
Each change turns one of B.1's diagnostics into a passing result:
- `past_capacity` and `long_block_fact_early` flip their strict xfails;
- a real-model re-run shows `ranked: yes` for the planting-era memory.
## O. What B.2 must not change
- **Lineage safety.** `tree.attach_memory`, the path clause, `forget_node` and E02
stay as they are. An abandoned line's memory stays stored and ineligible (§I).
- **Summary lineage** (E03) and summary content policy.
- **Authority.** Memory never writes state and is never framed as canon (F07, §J).
- **Imported-knowledge authority and retrieval.**
- **Pins.** Pinned memories stay always-selected and never evicted.
- **The frozen-bank fix.** A memory is never evicted by the pass that created it.
- **The single-commit use counter** (M11 §O.7).
- **F01-F08, E01-E04, the M04 verdicts, the bundle format and the schema**, unless a
migration is separately justified.
- **The deterministic diagnostic itself.** B.2 flips the strict xfails. It does not
weaken the scenarios or the isolation checks.
---
## P. Tests / regression
| Run | Result |
| --- | --- |
| `test_v11_b1_memory_diagnostic.py` on HEAD | **24 passed, 2 xfailed (strict)** |
| `test_v11_b1_memory_diagnostic.py` on v1.0.0 (§D) | **24 passed, 2 xfailed (strict)**, identical |
| `test_v11_b1_long_run_verdict.py`, `test_m11_long_run_memory.py`, `test_m11_long_run_resume.py` | **52 passed** |
| **Full backend suite**, HEAD plus the B.1 files | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,323 s). The 17 skips are the tests that need a real model, the same 17 as before. The 2 strict xfails are the two diagnosed retention criteria (§D). |
| **Full backend suite, final re-run at staging** (after the real-model attempt; the staged tree) | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,912 s), identical |
No frontend file was changed, so the frontend suite, lint and build are not affected.
---
## Q. Compatibility
| | Result |
| --- | --- |
| Application code changed | **none.** `git diff --stat HEAD -- backend/app frontend` is empty |
| Schema migration | none |
| Bundle format | unchanged |
| Stored campaign behaviour | unchanged |
| Memory behaviour | **unchanged.** Creation, eviction, ranking, pins and lineage are all as in v1.0.0. `memorybank.py` is byte-identical to the tag |
| What changed | Diagnostic tooling (`tools/memory_diagnostic.py`, `tools/v11_b1_memory.py`), an opt-in harness mode (`m11_long_run.py --independent-fact`, with every existing behaviour and M04 verdict unchanged when the flag is off) and tests |
| New strict xfails | 2 (§D). They document the two diagnosed defects, and the suite stays green. B.2 must remove them deliberately when it fixes the mechanisms |
## R. Security / local-only
Checked against the B.1 diff: the `m11_long_run.py` changes plus the four new files.
| | Result |
| --- | --- |
| New network client, endpoint, URL or TLS setting | **none.** The only address in the new code is `http://127.0.0.1:9/v1`, a refused loopback port the deterministic scenarios configure so nothing is contacted |
| External embedding service or remote vector store | **none.** The deterministic runs use `ConceptEmbedder` in-process. The real-model run uses the configured Ollama embedding model through the application's existing provider |
| Endpoint policy (`endpoints.py`, ADR 011) and TLS (`tlstrust.py`) | unchanged; not in the diff |
| Network calls in the real-model run | only the configured Ollama host on the trusted LAN, through the application's own provider and probe paths |
| Real identifiers in committed files | none. Checked again at staging |
| Offline container regression | **not run.** The harness builds and runs a Docker container, and this session's permission policy refused it. B.1 changes no runtime code and nothing in the image (the image carries `backend/app` and `frontend/dist` only), so the container would be byte-identical to the A1/A2 image. That image passed 23/23 on the corrective tree (`V1.1-WP-A1-A2-REPORT.md` §R.3) |
---
## S. Final decision
**What was established.** Every B.1 requirement was carried out except the clean
real-model run, which was attempted:
- the current mechanism (§B);
- a deterministic, isolated harness with four stage outputs and a verdict at recall
depth ≥ 100 (§C);
- the v1.0.0 baseline (§D);
- capacity and eviction (§F), the creation window (§E) and ranking (§G);
- the lineage and authority negative controls (§I, §J);
- the `recovered_through_memory_independent` long-run verdict, with the M04 verdicts
unchanged;
- tests and regression (§P), compatibility (§Q) and security (§R).
**The real-model gap.** The one real-model attempt ran cleanly, but failed isolation
(`precondition_failed:absent_from_summary`). The narrator and summariser restated the
fact, so a clean memory-only result on a real model was **not obtained**. The owner
decided not to make a second attempt, because the same narrator behaviour would very
likely repeat. The attempt is reported in full, and its stage diagnosis is used as
mechanism evidence only (§K).
**B.1 DIAGNOSTIC: COMPLETE**
**FIRST FAILING STAGE: RANKING.** In a real 100-turn campaign the early fact's memory
was created (it carries the fact) and retained (active, 33 of 80), but ranked 10th of
19 (similarity 0.866) against `top_k` 4. Later memories with near-identical wording
outranked it, under a query made of recent narration (§K, §L). Two further failures
are proven deterministically, identically on v1.0.0:
- **retention**, past `memory_bank_capacity`: least-recently-used eviction removes an
early memory first;
- **creation**, for a fact early in a block over 2,000 tokens: the summariser's
last-2,000-token excerpt drops it.
**WP-B.2 RECOMMENDED CHANGE: retrieval ranking first.** Build the retrieval query so
the newest player input is not diluted by three turns of recent narration, and add one
inspectable lexical term for the query's rare words to cosine ranking. The score must
be recorded in `memories.used`. Acceptance: a real-model re-run shows `ranked: yes`
for the planting-era memory.
Then, in separate test-first steps:
1. make eviction coverage-aware instead of purely least-recently-used, so the only
memory of an early range is not discarded first (flips
`test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity`);
2. choose the summariser excerpt so a fact early in a long block is kept (flips
`test_acceptance_a_fact_early_in_a_long_block_is_remembered`).
Everything in §O stays unchanged. B.2 has not been started.
File diff suppressed because it is too large Load Diff
+475
View File
@@ -0,0 +1,475 @@
# v1.1 WP-C — Browser Release Coverage
**Status:** COMPLETE, staged for owner review. The final run passed 91 checks with 0 failed and 0 skipped. The decision is in §R.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `0c1ba836babe1447ad3693b4d95189a325b7b3b6` — *v1.1 WP-B.2: independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its ancestry | `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `ac465ed` (plan v4.1), `432f041` (v1.0.0) |
| Working tree at start | clean; nothing staged |
| WP-D, WP-E | not started |
---
## B. Existing browser harness
`backend/tools/m11_browser.py` drives Firefox through `backend/tools/m11_webdriver.py`, a
dependency-free W3C WebDriver client. It builds its fixture campaign through the API,
plays two narrator turns through the streaming endpoint, and then asserts on the
rendered DOM. The M11 closeout run (`3652dc6`, run 3) passed **38/38**, 0 failed,
0 skipped, on snap Firefox 155.0.1 with geckodriver 0.37.1 and `qwen2.5:3b-instruct`
over trusted-LAN HTTPS.
What it did not do, and WP-C closes: drive Retry and takes, Save Point create and
restore, state correction, narration length or failed generation through the UI,
and prove an export leaves the browser as a file.
Before WP-C it also slept, for a fixed time, before several assertions: after Undo
and Redo, after planting hostile narration, after choosing a knowledge file, around
the delete dialog, and around opening a panel. §J records that as a harness defect.
---
## C. Browser and download environment
| | |
| --- | --- |
| Firefox | **155.0.1, the snap** (`/snap/bin/firefox`). No second Firefox was installed |
| geckodriver | 0.37.1 (the snap) |
| Headless | yes |
| Frontend | the production build (`vite build`) served by FastAPI on loopback; no Vite dev server |
| Certificates | `acceptInsecureCerts: false`, unchanged |
**The download profile.** `m11_webdriver.firefox_download_prefs` is passed as
`moz:firefoxOptions.prefs`:
- `browser.download.folderList` 2, `browser.download.dir` `<--out>/downloads`,
`browser.download.useDownloadDir` true;
- `browser.download.start_downloads_in_tmp_dir` false;
- `browser.download.always_ask_before_handling_new_types` false;
- `browser.helperApps.neverAsk.saveToDisk` `application/json,application/octet-stream`;
- the download panel suppressed.
**The folder.** `--out/downloads`, which must be under `$HOME`
(`require_under_home`). The harness deletes it at the start of a run and creates it
fresh. Evidence lives under `$HOME/v11-evidence/wp-c/`.
**Does the snap Firefox download?** Yes. Measured first, with a probe
(`$HOME/v11-evidence/wp-c/probe/`): a loopback page runs the product's own download
pattern (a JSON blob, an `<a download>` click, an immediate `revokeObjectURL`).
- Download folder under `~/v11-evidence`: 33-byte file written.
- Download folder under `~/Downloads`: 33-byte file written.
The first probe wrote nothing, and that was a probe defect (J1), not the snap. The
non-snap fallback the plan allows was therefore not needed.
**When a download counts as finished** (`m11_webdriver.wait_for_download`, tested
without a browser in `test_v11_c_browser_helpers.py`, 7 tests). All of these at once:
- a name absent from the listing taken before the click;
- no `*.part` file in the folder;
- more than zero bytes;
- the same size across 3 consecutive polls.
A zero-byte, partial, pre-existing or still-growing file never counts, and neither
does the "Campaign exported." toast.
---
## D. Retry scenario
Real narration: **yes**. All checks use the reader-facing controls on the play page.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| Type in "What you do next", press **Send** | a new narration renders and the page is idle; its exact text is recorded | PASS |
| — | **Retry** is offered (enabled) on the newest narration | PASS |
| Press **Retry** | the take indicator on the newest narration reads **2/2** | PASS |
| — | the second take's text differs from the first (otherwise 1/2 could not be told from 2/2) | PASS |
| Press **‹** (Previous take) | the indicator reads **1/2**, and the narration is identical to the first recorded text | PASS |
| — | the second take's text is not shown anywhere in the transcript | PASS |
| Press **›** (Next take) | **2/2**, showing the second take | PASS |
| Reload the page | the indicator still reads **2/2** on the live take, showing the second take | PASS |
| Press **‹** after the reload | **1/2** still shows the first text, unchanged | PASS |
All text comparisons are of the rendered `.turn-text`. No database was read.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## E. Save Point scenario
Real narration: **yes** (two turns after the Save Point).
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| Press **Save Point**, type a name in "Save this moment", submit | a row with that exact name appears | PASS |
| — | the row's moment is the moment being read ("Moment 7") | PASS |
| **Send** two turns | two new narrations render; position "Moment 11" | PASS |
| In Save Points, press **Restore**, then confirm **Restore** | the position reads "Moment 7 · later story ahead" | PASS |
| — | the transcript ends at the Save Point: its last narration is the one read there, and neither later narration is shown | PASS |
| — | the position says later story is ahead | PASS |
| — | **Redo** is enabled | PASS |
| Press **Redo** until it is disabled (each press waits for the position to change) | the last two narrations are the two later turns, text-identical, at "Moment 11" | PASS |
| Reload the page | the position is still "Moment 11" | PASS |
| — | the named Save Point is still listed | PASS |
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## F. State-correction scenario
**Owner decision (2026-09-15).** The State panel cannot produce a *partly* refused
correction:
- "Save correction" sends exactly one `add_fact`, and "That's wrong" sends exactly
one `invalidate_fact`;
- so a correction is applied whole or refused whole (HTTP 400);
- the route's partial application (`refused` alongside applied changes) is
reachable only through the API.
WP-C drives what the reader can reach: an accepted correction that persists, and a
refused correction whose refusal and reason are visible, with the refused change
not applied. It records the partial refusal as unreachable from the reader UI. No
product change was made for it.
Real narration: **yes** for the refusal (it needs history to step back over); the
accepted correction needs none and also passed in the no-narrator smoke runs.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| In **State**, press **Correct something**, type a fact, press **Save correction** | the form closes and the State panel shows the fact | PASS |
| Reload, open **State** | the fact is still shown | PASS |
| In a second tab press **Undo** (the correction belongs to the moment it was made at); in the first tab, whose panel still shows the fact, press **That's wrong** on it | a failure notice carrying the correction's refusal ("can't be applied") appears | PASS |
| — | it is labelled **"That correction was not applied"**, with no "Try that turn again" and no claim that typed text was kept (§K2) | PASS |
| Open **Show technical details** | the reason is visible: "That correction can't be applied — no fact 'f3' to invalidate." | PASS |
| Second tab **Redo**, close it; reload the first tab at the corrected moment | the fact still stands, with its **That's wrong** control: the refused withdrawal was not applied | PASS |
**Partial refusal.** As the owner decided, a *partly* refused correction is not
reachable from the reader UI, and was not driven. The route's partial application
remains covered by the backend suite.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## G. Narration-length scenario
Real narration: **yes** (one turn per band). The model's actual length is not
asserted, because nothing in the product contract requires it. What is asserted is
that the chosen band reached the turn's own prompt.
The expected sentence comes from the product's own `builder.length_hint` at the
run's 400-token output cap:
- **brief:** "must not exceed 180 words, and it should not stop short of about 70";
- **long:** "must not exceed 236 words, and it should not stop short of about 118".
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| **Settings** panel: choose **Brief** in "Narration length", press **Save changes** | the button reads "Saved" | PASS |
| **Send** a turn; choose **Long**, **Save changes**; **Send** a second turn | both turns render | PASS (both) |
| On the brief turn press **Inspect context** | the prompt sections of "The exact text the narrator was sent" contain brief's range, and do **not** contain that turn's own reply (so this is the turn's record, not the dry run of the next) | PASS |
| On the long turn press **Inspect context** | the same, with long's range | PASS |
| Reload, open **Settings** | "Narration length" still reads **long** | PASS |
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## H. Failed-generation scenario
**Owner decision (2026-09-15).** The literal sequence (save an unserved model, then
submit a turn) cannot be driven. The model check (`modelStatus.jsx`) marks a
configured model absent from the endpoint's list as `missing-model`, and
`blocksPlay` disables Send, Continue and Retry up front. That is M8's intended
behaviour, not a defect. So WP-C drives both of these:
1. **The unserved model** saved through Settings: the reader is told and cannot
send, and the story is unchanged.
2. **A submitted failure:** a model the endpoint lists but that cannot narrate,
`nomic-embed-text:latest`. Send stays enabled, and the turn fails in the open.
Then recovery with the reference model. The report records that path 2 uses a
listed, non-narrating model rather than "a name the server does not serve".
Real narration: **yes**.
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| **Settings**: "Type a model name instead", type an unserved name, **Save** | the page says "Saved" | PASS |
| Open the campaign | the header model status is `missing-model` and the setup notice is shown | PASS |
| — | **Send** and **Continue** are disabled | PASS |
| — | the story is unchanged | PASS |
| **Settings**: choose `nomic-embed-text:latest` from the installed-model picker, **Save** | "Saved" | PASS |
| Open the campaign (model status `ready`); type a turn; **Send** | a failure notice is shown: "Generation failed", with the server's reason under the details, `"nomic-embed-text:latest" does not support chat` (HTTP 400) | PASS |
| — | no narration was added | PASS |
| — | the typed text is still in the input box | PASS |
| — | the earlier story is text-identical | PASS |
| Open **State** | the rendered state is identical to before the failure | PASS |
| **Settings**: choose `qwen2.5:3b-instruct`, **Save** | "Saved" | PASS |
| Open the campaign; type a turn; **Send** | exactly one new narration | PASS |
| — | the earlier story is intact | PASS |
| Reload | the successful turn is still the last narration | PASS |
The failed turn's player moment stays in the transcript, as A05 intends
(`player_moment_kept_in_transcript` in the report).
**Deviation, owner-approved.** Path 2 uses a model the endpoint *lists* but that
cannot narrate, not "a name the server does not serve", because an unserved name
is caught before a turn can be submitted. Endpoint policy was not bypassed: both
paths use the same trusted-LAN HTTPS endpoint, and only the model name changed.
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
`server.log`, `geckodriver.log`, and the downloaded files.
## I. Export-download scenario
Real narration: not needed for the download itself. In the final run the exported
campaign holds 19 moments of real narration, two takes and a Save Point.
**C6a — the campaign library**
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| On the library page, press **Export** on the "Release Regression" card | a new file is written to `downloads/`; it is finished (no `.part`, stable size) | PASS |
| — | it is not empty: **136,739 bytes** | PASS |
| Parse the file | `format` is **`ai-dnd-adventure-v3`** | PASS |
| Start a **fresh application** (new database); press **Import campaign**; give its file input the downloaded path | the browser lands on the imported campaign's play page | PASS |
| Compare what the reader sees of the import with what the reader saw of the original (library card, position, Save Points panel) | same title ("Release Regression") | PASS |
| — | same number of moments (19) | PASS |
| — | same position ("Moment 18") | PASS |
| — | same Save Points, which also match the file | PASS |
The file's `headDepth` is 17, which the play page shows as "Moment 18", and its
`actions` count is 19.
**C6b — campaign settings**
| Browser action | Observable assertion | Result |
| --- | --- | --- |
| In the campaign's **Settings** panel, press **Export campaign** | a second, separate file is written and finished | PASS |
| — | not empty: 136,739 bytes | PASS |
| Parse the file | `format` is `ai-dnd-adventure-v3` | PASS |
Both files: `$HOME/v11-evidence/wp-c/final/``downloads/library-export.json` and `downloads/settings-export.json`.
The imported application's log is `import-server.log`.
---
## J. Harness defects found
| # | Defect | How found | Fix |
| --- | --- | --- | --- |
| J1 | The download probe's page wrote `URL.createObjectURL` inside an inline `onclick`, where `URL` is `document.URL`, a string. No blob was made, so it looked exactly like "the snap cannot download" | `gecko.log`: `TypeError: URL.createObjectURL is not a function` | The probe's script uses `window.URL`, the way the product's module does. Both download folders then worked |
| J2 | C3's first refusal withdrew a fact in a second tab and withdrew it again in the stale tab. Withdrawing keeps the fact, marked `invalidated` (C04's audit record), so the second withdrawal was valid and accepted. Nothing was refused, and "the refused change was not applied" passed without meaning anything | Smoke run: two C3 failures; `server.log` shows 201 for every correction; `narrative/apply.py` | The second tab steps the story back past the correction with Undo, so the stale tab withdraws a fact the story at that position does not have. The validator then refuses it: `no fact … to invalidate` |
| J3 | Clicks on a control just under the play page's fixed composer were intercepted ("Show technical details" in a failure notice, a turn's "Inspect context") | Dev run 1: `element click intercepted` | `Browser.click` scrolls the element to the centre of the view, then uses the real WebDriver click |
| J4 | The Settings model field is a text box until the endpoint's model list arrives, then a picker. Choosing before the check finished raced that swap | Dev run 1: `no such element: input#model` | Wait for the header's model status to leave `checking` first |
| J5 | C3's "a refused correction is shown" waited for *any* failure notice, so a notice about something else would have passed | Dev run 1: it passed on a notice titled "Generation failed" (§K2) | It now requires the notice to carry this correction's refusal ("can't be applied"), and asserts how it is labelled |
| J7 | C4 required the band's sentence in the inspector and the turn's own reply to be absent, to tell a turn's record from the dry run. But the inspector also renders what came back (`raw_output`, "What came back, before the state block was removed") inside the same section, so the absence could never hold | Dev run 2: both C4 inspector checks failed. The stored records show the brief turn's `length_hint` section carrying "must not exceed 180 words … about 70", the long turn's carrying "…236 … 118", and every turn's reply present in `raw_output` | The sentence and the absence are read from the prompt sections only, excluding the "What came back" block |
| J6 | M11's checks slept before assertions: after Undo and Redo, after planting hostile narration (1 s), after choosing a knowledge file (0.5 s), around the delete dialog (0.8 s and 0.6 s), and around opening a panel (0.8 s). A sleep is not evidence of what it waited for | Reading the harness against the brief's rule | Each is now a wait on the condition the check needs: the position changed or returned, the planted text rendered, Import enabled, a dialog present or gone, a panel-specific element present. §38's absence check now first waits for the knowledge library to render |
**Checked and not a defect.** In dev run 2, the original take at depth 6 (action 9)
had no `length_hint` in its stored record, while its Retry (action 10) did. The
original take's record holds only the per-attempt fields
(`attempts.ATTEMPT_KEYS`: world state, narrative state, raw output, usage,
accounting), with no sections and no settings. That is how a take that is no
longer live is stored, not a prompt built without the range. Every live turn's
prompt carried its band.
Also added: a panel opens only if it is not already open, because a tab toggles
its panel closed, and it is recognised by an element only that panel renders, not
by its title, which the tab itself already shows.
---
## K. Product defects found
### K1 — "Correct" on an Important Facts row is always refused (not fixed)
The State panel offers **Correct** on every row with a key. On an Important Facts
row the key is the fact's id, and on the scene-summary row it is `"summary"`.
`saveCorrection` sends that key as `add_fact.subject`, and the validator checks
`subject` as an entity reference. So every such correction is refused.
**Reproduced deterministically** against the real application (scratch `TestClient`,
no browser, no model):
| Correction the panel sends | Result |
| --- | --- |
| "Correct" on the Characters row (`subject='mara'`) | **201**, applied |
| "Correct" on the Important Facts row (`subject='f1'`) | **400** "That correction can't be applied — add_fact names subject='f1', which does not exist." |
**Not fixed in WP-C.** It does not prevent any WP-C behaviour: "Correct something"
and entity-row corrections work, and C3 uses them. A fix (offer Correct only
against entities, or send facts without a subject) is a small UX choice, left to
the owner as a v1.1 follow-up.
### K2 — a refused correction was presented as a failed turn (fixed)
**Found in the browser** (dev run 1, §J5). When the story refused a correction, the
failure notice said:
- the title **"Generation failed"**;
- the hint "Nothing was added to your story. You can try that turn again.";
- the button **Try that turn again**;
- the line "what you typed is still in the box below".
None of that is true of a State-panel correction. `classifyError` had no rule for
the server's refusal ("That correction can't be applied — …"), so it fell through
to the generation default. That falsified exactly what C3 checks: that a refusal
is shown to the reader as a refusal.
**Fix** (frontend only, no backend change):
- `errors.js`: one rule, checked first, for "correction can't be applied". It gives
`kind: state`, the title "That correction was not applied", the hint "Nothing in
the story or its state was changed. The reason is in the technical details.",
`retryable: false` and `keptInput: false`.
- `FailureNotice.jsx`: the "what you typed" line is shown unless a failure says
`keptInput: false`. Every other kind still shows it, so M8's A05 contract is
unchanged.
**Regression** (`failurePaths.test.jsx`, 3 tests):
- a refused correction classifies as a state refusal, not retryable, with the
reason kept;
- its notice shows the reason and no "Try that turn again" or typed-input claim;
- a failed turn still claims the typed words were kept.
In the browser, C3 now asserts the label, the absence of "Try that turn again" and
the absence of the typed-input line.
Frontend after the fix: **168/168** tests, lint exit 0 (15 pre-existing warnings, 0
errors, none in changed files), production build passes.
---
## L. Existing 38-check regression
**38/38 passed, 0 failed, 0 skipped** in the same run, tagged `M11`. The names are
identical to the M11 closeout run, so there is a one-to-one mapping and no check was
split, merged or dropped:
- B01 a turn is accepted (×2);
- A/UX: the tab title (×3);
- B position indicator (×3), D01 Undo, D04 Redo (×2);
- H06 (×3), H07, G09, H04;
- G01 knowledge import (×2);
- A11y dialog focus (×4);
- §38 narrator-only text absent;
- F05 context inspector;
- H10 (×2), H11/CSP (×2);
- A11y names, focus, tabindex, hover, contrast (×4), input focus.
What changed in them is how they wait (J6), not what they assert.
## M. New WP-C checks
**53/53 passed, 0 failed, 0 skipped**, tagged `WP-C`:
| Scenario | Checks | Result |
| --- | --- | --- |
| C1 Retry | 8 | 8 PASS |
| C2 Save Point | 9 | 9 PASS |
| C3 State correction | 6 | 6 PASS |
| C4 Narration length | 5 | 5 PASS |
| C5 Failed generation | 14 | 14 PASS |
| C6a Library export and import | 8 | 8 PASS |
| C6b Settings export | 3 | 3 PASS |
```text
existing M11 checks: 38/38
WP-C new checks: 53/53
failed: 0
skipped: 0
```
**Runs that are not the final evidence**, kept under `$HOME/v11-evidence/wp-c/`:
| Run | Where | Result | Why it is not evidence |
| --- | --- | --- | --- |
| `probe/` | loopback page | first: no download (J1); second: both folders written | environment probe |
| `smoke-1` | no narrator | 44 passed, 2 failed (J2), 6 skipped | partial |
| `smoke-2` | no narrator | 43 passed, 0 failed, 7 skipped | partial |
| `dev-gpu-1` | GPU host, plain HTTP | 71 passed, 3 failed (J3, J4; J5 found) | development, and not HTTPS |
| `dev-gpu-2` | GPU host, plain HTTP | 89 passed, 2 failed (J7) | development, and not HTTPS |
| `dev-gpu-3` | GPU host, `--only length` | 7 passed | development, and a subset |
## N. Production build / frontend verification
| | Result |
| --- | --- |
| Frontend suite (`npm test`) | **168/168**, 14 files. It was 165 before; the 3 new tests are K2's regressions |
| Lint (`npm run lint`, oxlint) | **exit 0: 0 errors**, 15 warnings. All are the pre-existing `only-export-components` kind, and none is in a file WP-C changed |
| Production build (`npm run build`) | passes; `dist/index.html` sha256 `62b6ea5eb4ce02a09f23cc1d48c335c2ada36b208b23c237338b39bf63a26cc5`, built 2026-09-15 15:11 from the final WP-C tree |
| What the browser ran | that build, served by FastAPI (`uvicorn app.main:app` on `127.0.0.1`); no Vite server |
| Backend product code | **unchanged**: nothing under `backend/app` is in the diff. The full backend suite was not rerun. The changed harness is tested by `test_v11_c_browser_helpers.py` (7 passed) |
| Offline regression | **23 passed, 0 failed** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`, on the final WP-C tree), rerun because the frontend bundle changed (K2). Evidence: `$HOME/v11-evidence/wp-c/offline/` |
## O. Trusted-LAN / security
| | |
| --- | --- |
| Narrator | `qwen2.5:3b-instruct` on the CPU reference host, over **trusted-LAN HTTPS**. The certificate is from the private CA in this machine's trust store, and verified; there is no bypass |
| Window and A1 | every narrator turn (8): window **verified at 4,096**, accounting **`fits`**. None `exceeded` or `truncation_suspected` |
| Protocol echoes (A2) | none of the protocol shapes the harness looks for (a state fence, a hard-limit or reminder bracket, `Events: [`) appeared in any stored narration in this run. This is not a v1.1 protocol-leak result: the release gate owns that, and the mid-reply echo from WP-B.1 remains a separate residual |
| Storyteller | loopback only, both applications (the original and the fresh import) |
| Browser | `acceptInsecureCerts: false`; the CSP checks (H11) passed |
| Endpoint policy | unchanged; C5 changed only the model name, never the endpoint |
| Downloads | written only under `$HOME` (enforced); the harness makes no network request of its own beyond loopback and the configured endpoint |
| New dependencies | none: the harness still uses only `urllib`, no Selenium or Playwright |
| Real identifiers in committed files | none (scanned at staging) |
## P. Compatibility
| Area | Effect |
| --- | --- |
| Database schema, migrations | none |
| Bundle format | none: the downloads are `ai-dnd-adventure-v3` and import unchanged |
| History, Save Points, state semantics, memory, knowledge | none. WP-C drove them and changed nothing in them |
| Backend | no application code changed |
| Frontend | one behaviour change (K2): the refusal of a State-panel correction is labelled as a refusal rather than as a failed turn. Every other failure's classification, retry offer and typed-input claim is unchanged (M8's tests pass) |
| WP-A, WP-B | untouched |
## Q. Residual risks
| # | Risk |
| --- | --- |
| 1 | **K1**: "Correct" on an Important Facts or scene-summary row is always refused. Reproduced, not fixed: a small UX choice for the owner |
| 2 | **Partial refusal is API-only.** The reader UI cannot produce a partly refused correction (owner decision). The display for one exists but is unreachable from the panel's own controls |
| 3 | **C5 path 2 depends on the reference host listing an embedding model.** On a host without one, the submitted-failure path has no model to use |
| 4 | **Model nondeterminism in C1.** "The second take is a different narration" would fail if the model returned identical text for a retry. It did not in any run |
| 5 | **The harness runs on this machine's snap Firefox.** A different Firefox or a Chromium would need the download preferences re-checked |
| 6 | **The heuristic protocol-shape scan is not A2 evidence.** It is recorded only |
| 7 | The WP-B real-model memory limitation and the doubled-full-stop scene text are unchanged, and are not WP-C's |
## R. Final decision
```text
RETRY: PASS
SAVE POINT: PASS
STATE CORRECTION: PASS
NARRATION LENGTH: PASS
FAILED GENERATION: PASS
EXPORT DOWNLOAD — LIBRARY: PASS
EXPORT DOWNLOAD — SETTINGS: PASS
EXISTING BROWSER REGRESSION: PASS
WP-C NEW BROWSER COVERAGE: PASS
WP-C OVERALL:
PASS
```
**PASS**, on these grounds:
- the final run ended with **failed: 0, skipped: 0**;
- both export controls produced a real, finished, non-empty `ai-dnd-adventure-v3`
file on disk;
- the library file imported into a fresh application with the same title, moment
count, position and Save Points.
Two criteria were met in the form the owner approved (2026-09-15):
- **State correction:** an accepted correction, and a refused correction with its
reason. Partial refusal is recorded as unreachable from the reader UI.
- **Failed generation:** the up-front block for an unserved model, plus a submitted
failure with a listed model that cannot narrate.
**Product change:** K2, a mislabelled refusal, fixed narrowly with regression tests.
**Product defect left open:** K1.
Nothing is committed, pushed or tagged. WP-D and WP-E have not started.
+526
View File
@@ -0,0 +1,526 @@
# v1.1 WP-D — Recovery Honesty
**Status:** COMPLETE — **PASS**. The decision, and what was deliberately not
claimed, is in §P.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD at start | `59b5ebc2d85bdf53d38b6bcf347c496dd3432be1` — *v1.1 WP-C: browser release coverage*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
| Its ancestry | `0c1ba83` (WP-B.2), `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `432f041` (v1.0.0) |
| Working tree at start | clean; nothing staged |
| WP-E | not started when WP-D was implemented |
---
## B. Existing backup behaviour
`backend/app/backup.py`, as v1 shipped it. The six questions the brief asks:
| # | Question | Answer (before WP-D) |
| --- | --- | --- |
| 1 | How is the backup file created? | SQLite's **online backup API** (`sqlite3.Connection.backup`, `pages=-1`), from the live database opened read-only through a `mode=ro` URI. The destination is a temporary file `…​.db.partial` **in the destination directory**, so the rename below is atomic |
| 2 | Where does validation happen? | `_verify()`, on the finished copy, opened as its own read-only connection — not on the source, and not through the connection that wrote it |
| 3 | When does the final filename appear? | Only after verification: `os.replace(working, target)`. A failed or interrupted run never leaves a file wearing a backup's name |
| 4 | How are failures cleaned up? | `_discard()` removes the partial file, and `BackupError` is raised with what went wrong. The source is untouched |
| 5 | Could an existing good backup be overwritten? | **No.** `_unused_name()` stamps each backup with the time and adds a counter if that name (or its `.partial`) exists |
| 6 | Was the check on the source or the copy? | The copy |
So the mechanism was already right. The one gap was the **strength** of the check:
`PRAGMA quick_check`, which reads every page and every record but skips the
cross-check between a table and its indexes.
---
## C. `integrity_check` implementation
One function changed:
```python
# app/backup.py, _verify()
rows = connection.execute("PRAGMA integrity_check").fetchall() # was quick_check
```
- It still runs on the **finished copy**, opened as its own read-only
connection, before the rename.
- A failure still raises `BackupError` naming what was wrong, still discards the
partial file, and still leaves the source and every earlier backup untouched.
- The returned `integrity` field, the filename, the directory, the reported
fields (`filename`, `bytes`, `pages`, `seconds`, `integrity`) and the API are
unchanged.
- The module docstring and `_verify`'s docstring now record why the trade M9
made (speed over the index cross-check) was not needed at these sizes, with
the measurements in §E.
**No scheduled backups were added.** There is still no restore endpoint, no
retention policy and no timer: a backup happens when the reader asks for one.
---
## D. Corruption fixture
`tests/test_v11_d_recovery.py: build_corrupt_copy()`.
A small database with `t(id, k, filler)` and an index `i_t_k ON t(k)`, 400 rows
with keys `k000000`…`k000399`. `dbstat` names the index's own leaf pages, and one
digit inside one indexed key on the first of them is changed (`k000144` →
`k000944`). Every page stays structurally sound and every record still parses:
what is broken is only the agreement between the index and its table.
**Proved before it is used as evidence**
(`test_the_fixture_is_the_difference_between_the_two_checks`):
| Pragma | Result |
| --- | --- |
| `PRAGMA quick_check` | **`ok`** |
| `PRAGMA integrity_check` | **`row 145 missing from index i_t_k`** |
That is the difference the package rests on: not a database both checks reject,
but one the old check called healthy.
---
## E. Backup performance measurements
Each pragma was run in its **own fresh process**, alternating, three times, because
a first measurement warms the page cache and a naive ordering makes whichever
check runs second look faster. (It did: an early single-pass measurement showed
`integrity_check` at 1.9 ms against `quick_check` at 65.5 ms, purely from cache
warmth.)
**The real campaign database** — the M11 100-turn evidence campaign,
`m04-final/campaign.db`, 2,367,488 bytes (2.26 MB):
| | Run 1 | Run 2 | Run 3 |
| --- | --- | --- | --- |
| `quick_check` | 5.0 ms | 3.4 ms | 7.6 ms |
| `integrity_check` | 3.5 ms | 5.7 ms | 5.3 ms |
At this size the two are indistinguishable. A whole backup through
`backup.create()` — copy, full check and rename — took **70.8 ms** and **27.4 ms**
on two runs (578 pages).
**An index-heavy synthetic database of 105.2 MB** (105,160,704 bytes; 25,674
pages of 4,096 bytes), shaped to give the cross-check real work: `actions` with
152,000 rows and `memories` with 15,200, under three indexes
(`i_actions_adv_depth`, `i_actions_kind`, `i_memories_adv`):
| | Run 1 | Run 2 | Run 3 |
| --- | --- | --- | --- |
| `quick_check` | 172.6 ms | 92.4 ms | 97.1 ms |
| `integrity_check` | 227.8 ms | 196.0 ms | 227.4 ms |
A whole backup of it through `backup.create()`: **0.58 s** for 25,674 pages,
`integrity = ok`.
**A campaign-shaped database of 117.4 MB** (117,403,648 bytes; 2,200 actions),
built through the application's own models so the product path can run against
it (§E.1):
| | Run 1 | Run 2 | Run 3 |
| --- | --- | --- | --- |
| `quick_check` | 89.5 ms | 74.2 ms | 64.4 ms |
| `integrity_check` | 90.0 ms | 82.9 ms | 75.9 ms |
**Reading.** The cost of the cross-check tracks **rows and index entries, not
bytes**. On the campaign schema — 2,200 fat rows — the full check costs about
7 ms more than the quick one at 117 MB. On the index-heavy synthetic — 167,200
rows under three indexes at a similar size — it costs about 96 ms more. Both sit
inside a backup of about half a second. There is no performance requirement in
this project and WP-D does not invent one; the measurements are here because the
M9 trade was made on a speed argument, and at these sizes that argument does not
hold.
### E.1 Through the endpoint the button calls — implementation evidence only
**This does not satisfy acceptance criterion 3, and an earlier draft of this
report wrongly said it did.** The criterion asks for the backup to *complete
through the UI*; a reader does not call an endpoint. What follows is evidence
about the implementation — the procedure, its result and its cost — and it is
kept for that reason. The acceptance evidence is in **§E.2**.
`POST /api/backups` is the only call the **Back up now** button makes, and it
takes no parameters, so it was driven directly against a **copy** of each
database — the evidence database is evidence and is not written to:
| Database | Result |
| --- | --- |
| Evidence campaign, 2,367,488 bytes | **201** in 0.117 s wall — `pages: 578`, `seconds: 0.038`, `integrity: ok`; `GET /api/backups` then lists 1 backup |
| Campaign-shaped, 117,403,648 bytes | **201** in 0.536 s wall — `pages: 28,663`, `seconds: 0.452`, `integrity: ok` |
**Why a second large database exists.** The index-heavy synthetic has no
application schema, and the server runs its migrations at startup, so the app
will not start against it: driving the endpoint there fails in a migration that
renumbers `actions`, before any backup is attempted. It remains valid evidence
for the pragma comparison, which is a property of SQLite and not of this schema,
but criterion 3 needs a database the product can actually open — hence the
campaign-shaped one, which §E.2 then drives through the browser.
Artefacts stay under `$HOME` (`v11-evidence/wp-d/`), not in the repository.
---
### E.2 Through the UI — the acceptance evidence for criterion 3
The reader-facing **Back up now** control, clicked in a real Firefox, on the
production build served by FastAPI: the WP-C path, with no Vite dev server, no
inference, and no sleep used as an assertion. Every wait is on a condition the
page or the filesystem can show.
`backend/tools/wpd_backup_ui.py`, run as
`python -m tools.wpd_backup_ui --out $HOME/v11-evidence/wp-d/ui --case both`.
**34 checks, 34 passed, 0 failed** —
`$HOME/v11-evidence/wp-d/ui/backup-ui-report.json`.
The route a reader takes is the route the tool takes: open the application, click
**Settings** in the navigation, open *Back up everything on this machine* (a
`<details>` that loads what is on disk when it opens), then press the button.
| | **Case 1 — real campaign database** | **Case 2 — ≥100 MB application database** |
| --- | --- | --- |
| Source | the M11 evidence campaign, copied | the campaign-shaped database of §E, copied |
| **Database size** | **2,367,488 bytes** (2.3 MB) | **117,403,648 bytes** (112.0 MB) |
| Schema | the application's own, `user_version` 94 | the application's own, `user_version` 94 |
| **Browser action** | click **Back up now** (twice — see below) | click **Back up now** |
| **Observable UI result** | toast: *"Backup written: adventure-storyteller-20260915-225000.db (2.3 MB)."*, no error toast, the control returns from *Backing up…*, and the file appears in the panel's list | toast: *"Backup written: adventure-storyteller-20260915-225007.db (112.0 MB)."*, no error toast, control returns, file listed |
| **Backup path** | `…/ui/real/data/backups/adventure-storyteller-20260915-225000.db` | `…/ui/large/data/backups/adventure-storyteller-20260915-225007.db` |
| **Backup file size** | **2,367,488 bytes** | **117,403,648 bytes** |
| **`integrity_check` result** | **`ok`** | **`ok`** |
| **Integrity-check elapsed** | **5.8 ms** | **78.5 ms** |
| **Total elapsed backup** | **0.108 s** (click → the UI says it is done) | **0.719 s** |
"Total elapsed" is measured from the click to the rendered result, so it is what
the reader waits, not what the server reports. The `integrity_check` above is an
independent second opinion, run here on the artefact the UI produced — the
application had already verified the copy before keeping it.
**The size the UI reports is the size on disk.** "112.0 MB" in the toast is
117,403,648 bytes shown in the page's own units; the file matches the source
database byte for byte.
**Existing-backup protection, established through the UI itself** (Case 1's
second press, rather than a backup planted by a library call):
| Claim | Result |
| --- | --- |
| A second press writes a *different* file — `…-225000-2.db`, the counter `_unused_name` adds when two backups land in the same second | PASS |
| The first backup still exists afterwards | PASS |
| And is byte-for-byte what it was — 2,367,488 bytes, `sha256:7d45555566…` | PASS |
| And still passes `integrity_check` | PASS |
No corruption was fabricated through the browser; rejection behaviour is already
proved deterministically in §D.
**Supplemental endpoint evidence.** The servers' own logs corroborate that the
button drove each backup: Case 1 logged two `POST /api/backups → 201 Created`,
each followed by `GET /api/backups → 200 OK` as the panel reloaded its list;
Case 2 logged one of each. Screenshots and logs are beside the report JSON.
**A harness defect this run found, in my own check and not in the product.** The
first pass reported six failures. The toast renders a decorative mark before its
message — `<span class="toast-mark" aria-hidden="true">❖</span><span>…</span>` —
so the button's `textContent` begins with `❖`, and my assertion matched
`textContent.startswith("Backup written:")`. In that same run the UI had shown a
non-error toast naming the file, the file was on disk, its name was in the
panel's list, and the copy verified: the product was right and the assertion was
reading the mark. The check now reads the message span. **No product code was
changed** — the expected change for this work was none, and none was needed.
---
## F. Existing export/import limit behaviour
| | |
| --- | --- |
| Import ceiling | `limits.MAX_IMPORT_BODY_BYTES` = 20 MB, **unchanged by WP-D** |
| Where it is enforced | `BodySizeLimitMiddleware`, on the declared `Content-Length`, before the body is read |
| Refusal | HTTP **413**, "Request too large (limit 20 MB)." — it already named the limit |
| Export before WP-D | `GET /adventures/{id}/export` returned the bundle. Nothing compared its size with the ceiling, so a campaign could be exported and then refused by its own importer |
---
## G. Oversized-export warning implementation
**Where the answer comes from.** `limits.import_limit_label()` and
`limits.oversized_export_warning(size)` derive the sentence from
`MAX_IMPORT_BODY_BYTES`. A test changes the constant and asserts the wording
follows, so no number is written twice.
> This export is larger than this version's 20 MB import limit (21,230,000 bytes).
> The file was exported successfully, but this version cannot import it.
**Where it travels: headers, not the body.** The export response *is* the bundle —
the browser saves exactly those bytes as the file — so a warning inside it would
become part of a portable story file and of every checksum taken over one. The
route returns the same body with:
```text
X-Export-Bytes the serialised size
X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
X-Importable-By-This-Version true / false
X-Export-Warning only when false
```
**Which size is measured.** The compact serialisation (`separators=(",", ":")`,
`ensure_ascii=False`), which is what this response sends *and* what the browser
POSTs back on import — the bytes `BodySizeLimitMiddleware` weighs. The
pretty-printed file the reader downloads is larger and is not what import reads.
The route serialises once and returns those bytes, so `X-Export-Bytes` is the
length of the body actually sent.
**The bundle is unchanged.** No key was added to it (§I), and the format stays
`ai-dnd-adventure-v3`.
---
## H. Oversized export evidence
`test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back`
writes real rows into a campaign until its bundle genuinely exceeds the ceiling —
nothing is mocked, and the export serialises all of it.
| Claim | Result |
| --- | --- |
| The response body is larger than 20 MB | PASS |
| It parses, and is `ai-dnd-adventure-v3` with its actions | PASS |
| `X-Importable-By-This-Version: false` | PASS |
| `X-Export-Warning` names the limit ("20 MB"), says the export succeeded, and says this version cannot import it | PASS |
| `X-Export-Bytes` equals the body length | PASS |
| The same bundle is refused by import, with the limit named (§J) | PASS |
| A normal campaign's export carries **no** warning, and `X-Importable-By-This-Version: true` | PASS |
| The warning follows the constant (changed to 50 MB in a test: the sentence says 50 MB, and no longer says 20 MB) | PASS |
---
## I. Normal-export compatibility
The same campaign database (`m04-final/campaign.db`) exported through a worktree
at `59b5ebc` (before WP-D) and through the WP-D tree:
| | |
| --- | --- |
| Before | 2,699,076 bytes |
| After | 2,699,076 bytes |
| Comparison | **identical** — the parsed documents compare equal, key for key |
No timestamp allowance was needed: this campaign's bundle carries no field that
moves between exports. The format is `ai-dnd-adventure-v3` in both.
`test_the_bundle_itself_never_carries_the_warning` additionally asserts that no
key of an oversized bundle mentions the warning, the limit or importability.
---
## J. Import-refusal behaviour
Unchanged in behaviour, and already naming the limit:
| | |
| --- | --- |
| Status | **413**, from the middleware, on `Content-Length`, before the body is parsed |
| Message | "Request too large (limit 20 MB)." |
| WP-D test | `test_that_same_bundle_is_refused_by_import_naming_the_limit` posts the *actual oversized export* and asserts 413 and that the detail contains `limits.import_limit_label()` |
| Not relaxed | the same boundary, the same status, the same middleware. `test_a_bundle_under_the_limit_still_imports` keeps the ordinary path honest (201) |
No wording correction was needed.
---
## K. Frontend behaviour
`api.exportAdventure` now returns the bundle **and** what the server said about
it (`exportBytes`, `importLimitBytes`, `importable`, `warning`), read from the
headers.
**The no-header case.** `importable` is `resp.headers.get(…) !== 'false'`, so
only the literal string `false` is read as a refusal: an older server that sends
no headers — or a header that arrives malformed — yields `importable: true` and
`warning: null`, and the page says nothing it was not told. The two size fields
fall back to `null` unless they parse as a positive number.
This is asserted directly: three tests stub `fetch` with real response headers
and check what `api.exportAdventure` makes of them — a warning with the sizes it
named, an ordinary export marked importable, and an older server sending no
headers at all. They exist because the page-level tests below mock
`api.exportAdventure` itself and so cannot see a header name, which meant a typo
on the frontend side would have left the whole suite green (§O.7).
Both reader-facing entry points keep delivering the file first and then report:
| Entry point | Under the limit | Over the limit |
| --- | --- | --- |
| Campaign library, **Export** | file written; "Campaign exported." | file written; the warning, as an error-styled toast |
| Campaign settings, **Export campaign** | file written; "Campaign exported." | file written; the warning |
No new modal, no redesign: the existing toast carries it.
**Tests** (`pages/exportHonesty.test.jsx`, **7**): three read real response
headers through a stubbed `fetch` (above), and four drive both entry points —
the file is delivered in every case; the warning is shown when the server sends
one, naming 20 MB and saying the export succeeded; the ordinary confirmation is
shown when it does not.
---
## L. Offline regression
`tools/m11_offline.py`, against the production build, in the offline container:
**23 checks, 23 passed, 0 failed** —
`$HOME/v11-evidence/wp-d/offline/offline-report.json`.
The two this package could have broken are in it and passed:
| Check | Result |
| --- | --- |
| a campaign exports offline | ok |
| and imports offline, with its state | ok |
| no secret is present in the export | ok |
| a local file imports offline | ok |
The export route now sets headers and serialises compactly; the offline
container still exports a campaign and imports it back with its state, so
neither the round trip nor the secret-scrubbing changed.
## M. Full regression
| Suite | Result |
| --- | --- |
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,712 passed, 17 skipped, 0 failed, 0 xfailed** (936.6 s) |
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
| **Offline regression** | 23 passed, 0 failed (§L) |
**The backend count reconciles exactly.** WP-B's closing tree was 1,693; WP-C
added 7 (`test_v11_c_browser_helpers.py`) and did not rerun the suite because no
application code changed; WP-D adds the 12 in `test_v11_d_recovery.py`.
1,693 + 7 + 12 = **1,712**.
**The 17 skips are named, not assumed.** Run with `-rs`, every one is an
environment-gated real-model test, and none is new:
| File | Skipped | Gate |
| --- | --- | --- |
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
| `test_narrative_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
| `test_m11_real_window.py` | 3 | two the same; one `AIDND_TEST_WIDE_MODEL` |
| `test_provider_wiring.py` | 1 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
That is the same 17 B.1 and B.2 recorded. WP-D used no model and added no skip.
**The plan's named regression requirements**, run individually rather than
assumed to be inside the total:
| Requirement | Result |
| --- | --- |
| I01-I07, L01-L04 | **25 passed**, 1,704 deselected |
| The backup case in `test_m11_migration.py` | **1 passed** |
| `backup.test.jsx` | **7 passed** |
| The offline container's export and import | passed (§L) |
**Frontend arithmetic.** WP-C's baseline was 168. WP-D adds 4 page-level export
tests and 3 header-parsing tests: 168 + 7 = **175**. The 15 lint warnings are
the same pre-existing `only-export-components` and unused-import kind recorded
at WP-C, and **none is in a file WP-D changed**.
## N. Compatibility
| Area | Effect |
| --- | --- |
| Database schema | unchanged; no migration |
| Bundle format | `ai-dnd-adventure-v3`, unchanged |
| Normal bundle contents | byte-identical (§I) |
| Existing backups | still valid files with the same names; only the check that admits a new one is stricter |
| v1.0.0 databases | open unchanged (no schema or data path changed) |
| Import limit | unchanged at 20 MB |
| Scheduled backups | none, as before |
| History, branches, Save Points, state, memory, knowledge | untouched |
| Endpoint policy, local-only operation | untouched |
## O. Residual risks
1. **Closed: the backup is now driven through the UI.** This risk previously
read "the browser click was not driven for the backup", and the owner
correctly refused criterion 3 on endpoint evidence. §E.2 drives the real
**Back up now** control in Firefox on both databases — 34 checks, 0 failed —
including the ≥100 MB case, which does mean the harness starts a server
against a 117 MB database and takes 0.719 s to do it. What remains is
ordinary coverage scope: this runs as its own tool rather than inside the
release harness, so it is not part of the 101-check run in the WP-E report.
2. **The 20 MB ceiling is unchanged.** A campaign past it still cannot be
imported by this version. WP-D makes that audible at export; raising the
limit or streaming import stays v1.2, as the plan assigns it.
3. **The warning is about *this* version.** It says what this build's importer
will accept. A future build with a higher ceiling could import a file this
one warned about, and the wording ("this version") is chosen so that stays
true rather than becoming a lie.
4. **`integrity_check` is not a guarantee of recoverability.** It proves the
copy's pages, records and indexes agree. A database that was already
logically wrong when it was copied is copied faithfully and passes. The
package makes the check honest, not omniscient.
5. **No restore path.** There is still no restore button and no scheduled
backup — both explicitly out of scope. A reader who needs a backup restores
it by replacing the file themselves, as before.
6. **The header channel depends on the browser reaching headers.** Both export
call sites read them through `fetch`, so a proxy that stripped `X-` headers
would silently return to v1 behaviour: the file still arrives, the warning
does not. Local-only operation makes that unlikely, and the failure is the
old behaviour rather than a wrong claim.
7. **Header parsing was untested — now closed.** Writing this section exposed
it: the page-level tests mock `api.exportAdventure` and the backend tests
assert what the server sends, so nothing read an actual header name, and a
frontend-side typo would have left every test green and the reader silently
uninformed. Three tests that stub `fetch` with real headers now cover it
(§K). What remains is the ordinary version of this risk: the two sides agree
by matching string literals in two files, and only a browser-level test would
catch a mismatch introduced in both at once.
## P. Final decision
**Against the plan's acceptance criteria** (§ WP-D, *Recovery honesty*):
| # | Criterion | Verdict |
| --- | --- | --- |
| 1 | A healthy database's backup reports `integrity_check` ok and is kept | **PASS** — `test_a_healthy_backup_passes_the_full_check_and_is_kept`, and both endpoint runs returned `integrity: ok` (§E.1) |
| 2 | A copy `integrity_check` rejects and `quick_check` does not is rejected and not kept; existing backups still never overwritten | **PASS** — the fixture is proved to be exactly that difference (§D), and three tests cover rejection, the untouched earlier backup, and the unchanged file semantics |
| 3 | The time is recorded on the evidence database and on a synthetic ≥100 MB, and the backup completes through the UI on both | **PASS** — pragma timings in §E on three databases, and the backup completes **through the UI** on both in §E.2: 0.108 s with `integrity_check` `ok` in 5.8 ms on the 2.3 MB campaign, 0.719 s with `ok` in 78.5 ms on the 117 MB application database. An earlier draft claimed this criterion on endpoint evidence (§E.1); that claim was wrong and is corrected |
| 4 | An over-20 MB export succeeds, delivers the file, and warns naming the limit, in the API response and a component test; under the limit, no warning | **PASS** — §H, and 7 frontend tests (§K) |
| 5 | Importing that bundle is refused with a message naming the limit | **PASS** — §J, asserted on the actual oversized export |
| 6 | A normal export is byte-identical before and after, apart from timestamps | **PASS** — fully identical, no timestamp allowance needed (§I) |
**What I corrected rather than reported around.** Two claims of my own failed
checking and were fixed, not softened: §E originally rested criterion 3 on
`backup.create()` calls, and driving the real endpoint showed the index-heavy
synthetic cannot serve that criterion at all (the app's startup migrations abort
against a schema-less database) — hence the campaign-shaped database in §E.1.
And §K asserted the no-header fallback from the code alone, which exposed that
**no test read any header name**; three tests now do (§O.7).
**Not done, and not claimed:** the import ceiling is unchanged at 20 MB, by the
plan's own assignment of the raise to v1.2. The UI backup runs as its own tool
rather than inside the release harness (§O.1).
**No product code changed for the UI verification.** The run exposed one defect,
and it was in my own assertion, not in the application (§E.2). Backend
**1,723 passed / 17 skipped / 0 failed** and offline **23/23** therefore stand
as prior evidence, unchanged and not re-run; the targeted backup and browser
checks were re-run instead.
```text
BACKUP INTEGRITY: PASS
BACKUP THROUGH UI — REAL DB: PASS
BACKUP THROUGH UI — >=100 MB DB: PASS
EXPORT HONESTY: PASS
WP-D OVERALL:
PASS
```
All WP-D changes are **staged and uncommitted**. No commit, no push, no tag.
WP-E has not begun.
+380
View File
@@ -0,0 +1,380 @@
# v1.1 WP-E — Control-Boundary Contrast
**Status:** COMPLETE — **PASS**, with owner approval of the screenshots
outstanding. The decision, and what was deliberately not claimed, is in §O.
---
## A. Repository baseline
| | |
| --- | --- |
| Branch | `v1.1-development` |
| HEAD | `59b5ebc` — *v1.1 WP-C: browser release coverage*, signed by the owner |
| Working tree at start | WP-D staged (10 files), nothing committed |
| Criterion | WCAG 2.1 **1.4.11 Non-text Contrast**, 3:1, for control boundaries; **1.4.3** 4.5:1 for body text, unchanged |
---
## B. What v1.0.0 actually did
`tools/contrast_audit.py` measured control boundaries, printed that two of them
were below 3:1, and **exited 0**. Its own comment argued the position:
> in this design a control is identified by its *label*, which is measured above
> and passes, not by its edge. So a boundary below 3:1 is reported with its
> number and does not fail the run.
So the audit was a report, not a gate: no palette change could ever fail it on a
boundary. The two numbers it printed were **1.33:1** (`--border` on
`--bg-panel`) and **1.75:1** (`--border-bright`), against a floor of 3.0.
**WP-E overturns that argument.** 1.4.11 covers the visual information needed to
identify a component *and its boundary*; a reader who cannot see where a text box
ends cannot see that there is a text box to type into, label or no label. The
tokens were raised rather than the criterion re-argued.
---
## C. Token inventory
| Token | v1.0.0 | v1.1 | Why |
| --- | --- | --- | --- |
| `--border` | `#2b2b3d` | **`#676792`** | every control's resting edge |
| `--border-bright` | `#3d3d55` | **`#7a7aaa`** | hover edges, the composer's resting edge, the open panel tab |
| `--bg-panel`, `--bg-input`, `--bg`, `--text`, `--text-dim`, `--accent*`, `--danger`, `--warning`, `--player`, `--chart-*` | — | **unchanged** | WP-E is a boundary package; no text or accent colour moved |
**The floor is taken against `--bg-input`, not `--bg-panel`.** Inputs and buttons
are drawn on `--bg-input` (`styles/forms.css`), which is lighter than
`--bg-panel` and therefore the harder case. The audit had been checking only
`--bg-panel`, so a token could have passed the audit while the real control
failed. Measured on the new values:
| | vs `--bg-input` | vs `--bg-panel` | vs `--bg` |
| --- | --- | --- | --- |
| `--border` | **3.21:1** | 3.44:1 | 3.70:1 |
| `--border-bright` | **4.24:1** | 4.55:1 | 4.88:1 |
Two properties were preserved deliberately: the rest→hover step is the same size
as before (1.318 → 1.320), so hover still reads as a change rather than a jump;
and `--border-bright` stays *below* body text against the same panel (2.98:1
between them), so no edge outshines the words inside it.
---
## D. Component inventory
Where these tokens are actually drawn, from the stylesheets:
| Control | Rule | Rest | Hover / active | Focus |
| --- | --- | --- | --- | --- |
| Story composer | `.input-bar` (story.css) | `--border-bright` on `--bg-panel` | — | `--accent-dim` + `--accent-glow` ring |
| Story controls | `.story-controls button` | `--border` on `--bg-panel` | `--border-bright` | (M11 focus check) |
| Fields and buttons | `forms.css` | `--border` on `--bg-input` | `--accent-dim` | `--accent-dim` + ring |
| Panel tabs | `.panel-tabs button` | **`transparent`** | `--border` on hover, `--border-bright` when active | — |
| Top navigation | `.topnav` | `--border` bottom edge on **`--bg-panel-glass`** | — | — |
| Scrollbar thumb | `base.css` | `--border-bright` as a *fill* on `--bg` | `--accent-dim` | — |
Two of these cannot be answered by token arithmetic at all, and both are
measured in the browser instead (§G): the nav sits on a translucent panel, and
the panel tab's edge is `transparent` until the panel is open.
---
## E. The audit is now a gate
`tools/contrast_audit.py`:
1. **Boundary pairs fail.** `text` and `boundary` rows are both pass/fail; the
advisory branch is gone. The run returns 1 if either kind falls short.
2. **Eight boundary pairs replace two.** Each border is checked against every
background it is drawn on — `--bg-input`, `--bg-panel` and `--bg` — plus the
focused edge (`--accent-dim`) on both panel and field backgrounds.
3. **The verdict is taken on the number that is printed** (rounded to two
decimals), so a pair shown as `3.00:1` is not failed for arithmetic the
reader cannot see.
4. The comment block that argued the old position is replaced by one recording
what changed and why, including why `--bg-panel-glass` is not in the list.
## F. Gate tests
`backend/tests/test_v11_e_contrast.py` — **11 passed**. The threshold is
exercised from both sides, on real token files:
| Test | Result |
| --- | --- |
| A boundary at **2.99:1** against `--bg-input` fails the run (exit 1) | PASS |
| A boundary at **3.00:1** passes (exit 0) | PASS |
| The **v1.0.0 value** `#2b2b3d` fails, at 1.24:1 against `--bg-input` | PASS |
| A dimmed `--text-dim` still fails as a *text* pair | PASS |
| A renamed token is a failure, not a silent skip | PASS |
| The shipped palette passes both criteria | PASS |
| Every boundary pair is measured against the background it is drawn on | PASS |
| The M11 text baselines are unchanged: 14.57 / 13.57 / 5.48 / 5.88 | PASS |
| The hover edge stays brighter than the resting edge | PASS |
| No boundary becomes as loud as body text | PASS |
| The WCAG ratio formula is anchored on known values (21:1, 1:1, symmetry) | PASS |
The 2.99 and 3.00 values are worth noting: **both clear 3:1 against
`--bg-panel`** (3.21 and 3.22). They decide the gate only because the floor is
now taken against the background the control is really on — so these two tests
also prove §C's change is doing work.
---
## G. Browser measurement
`tools/m11_browser.py` gains a third suite, **WP-E**, counted separately from
M11's 38 and WP-C's 53. It measures the *rendered* edge — `borderColor` from
`getComputedStyle` — against what is actually behind it, with every translucent
layer composited bottom-up.
A boundary is measured against **both** adjacent colours (the control's own fill
inside it, the background outside it) and passes on the better of the two: an
edge that matches its fill but contrasts with the page is still a visible
outline. What 1.4.11 asks is that the component's extent be perceivable.
Two harness capabilities were added for this (`tools/m11_webdriver.py`):
- **`hover()`** moves a real pointer through the WebDriver Actions API.
Dispatching a `mouseover` event from JavaScript does *not* trigger CSS
`:hover`, so a synthetic event would have re-measured the resting edge and
reported it as the hover edge.
- **`screenshot()`** writes the viewport as a PNG, for the before/after evidence.
**A defect this found in my own first measurement.** The first run reported the
hover edge as `rgb(114, 114, 160)` and the focused edge as `rgb(144, 120, 81)` —
neither of which is any token. Both controls carry `transition: border-color
0.15s`, so the measurement was taken mid-animation, on a colour no state
actually has. `_settled()` now polls until the computed edge colour is the same
on two consecutive reads before measuring (polled, not slept, per this harness's
own rule). After the fix the same edges read exactly `rgb(122, 122, 170)`
(`--border-bright`) and `rgb(150, 119, 58)` (`--accent-dim`).
---
## H. Before and after, measured in the browser
Both passes were taken the same way — `--only boundaries --no-narrator`, the
production build — with only `tokens.css` differing. Evidence under
`$HOME/v11-evidence/wp-e/before/` and `.../after/`.
| Control (state) | Before | After | Floor |
| --- | --- | --- | --- |
| Story composer — resting edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
| Story control — resting edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
| Open panel tab — resting edge | **1.75:1** FAIL | **4.55:1** pass | 3.0 |
| Story control — hover edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
| Top navigation — translucent edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
| Story composer — focused edge | 4.70:1 pass | 4.70:1 pass | 3.0 |
**Suite result: before 5 passed / 5 failed; after 10 passed / 0 failed / 0
skipped.**
Two things this table says that a summary would blur:
- **Focus was never the defect.** The focused edge (`--accent-dim`) already
cleared 3:1 in v1.0.0 at 4.70:1, and WP-E did not change it. What failed was
rest and hover — the states a reader spends all their time in.
- **The translucent edge is real evidence.** The nav's background composited to
`rgb(17, 17, 29)` — `--bg-panel-glass` (rgba 19,19,32 @ 0.82) over
`rgb(10, 10, 15)` — not a fallback. That is the case token arithmetic cannot
reach, and it moved from 1.43:1 to 3.70:1.
## I. Screenshots
| File | |
| --- | --- |
| `before/control-boundaries.png` | 131,175 bytes, 1366×682 |
| `before/control-boundaries-nav.png` | 63,376 bytes, 1366×682 |
| `after/control-boundaries.png` | 132,430 bytes, 1366×682 |
| `after/control-boundaries-nav.png` | 63,541 bytes, 1366×682 |
All four are PNG, 1366×682, taken on the production build through the same
harness path, differing only in `tokens.css`. The play-page pair shows the
composer, the story controls and the open panel tab; the nav pair shows the
translucent top edge on the library route.
## J. A finding this package created and fixed
Raising `--border-bright` broke something that had nothing to do with control
boundaries. `.slice-7` in the context inspector's token breakdown was painted
with `var(--border-bright)`, so it followed the token to `#7a7aaa` — an OKLab ΔE
of **0.035** from `.slice-6` (`#7c86b8`), making two neighbouring chart slices
effectively the same colour. The other slices sit **0.100–0.119** from their
nearest neighbour.
`.slice-7` is now pinned to `#3d3d55`, the literal value it already rendered, so
its appearance is unchanged from v1.0.0 and its separation (ΔE **0.251**) is the
widest in the set. A chart fill and a control edge have different jobs and should
not share a token.
Reassigning it to a fresh hue was considered and rejected on evidence: inside the
palette's own chroma (0.045–0.120) and lightness (0.586–0.804) bands, the only
hues clearing the set's 0.100 separation floor are pinks near 14°, which is
`--danger`'s territory. Painting an ordinary prompt section in the colour this
application reserves for failure would trade an accessibility fix for a semantic
lie.
*(A first attempt, `#5d7f9e`, was rejected by the same measurement at ΔE 0.061 —
below every real slice. It is recorded here because it was written into the file
before it was measured.)*
## K. Text contrast regression
Unchanged, and asserted so in §F: **14.57:1** body text on the page, **13.57:1**
in a panel, **5.48:1** secondary text in a panel, **5.88:1** on the page. Every
text pair still clears 1.4.3, and no text token was touched.
## L. Browser release regression
The full harness, all three suites, against a real narrator on the production
build. Evidence: `$HOME/v11-evidence/wp-e/release/browser-report.json`.
| | |
| --- | --- |
| Kind | **`release regression`** — not `partial`, not `development (only …)` |
| Narrator | `qwen2.5:3b-instruct`, the reference model, over plain HTTP on the LAN GPU host |
| Browser | Firefox 155.0.1, geckodriver 0.37.1 |
| Served | FastAPI on loopback, the **built** SPA (`dist` 2026-09-15T21:08:52) |
| Duration | 71 s |
| Turns played | **8**, of which `turns_not_clean` **0** and `protocol_shapes_in_narration` **0** |
**Counted per suite, as the brief requires:**
| Suite | Passed | Failed | Skipped |
| --- | --- | --- | --- |
| **M11** (the v1 release regression) | **38** | **0** | **0** |
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
| **WP-E** (control boundaries) | **10** | **0** | **0** |
| **Total** | **101** | **0** | **0** |
M11's 38 and WP-C's 53 are unchanged in count and in name: WP-E added a suite
beside them rather than altering either. The `kind` field is quoted above
because a fast run invites the question — 71 s for 101 checks including 8
narrated turns is the GPU host being quick with a 3B model, and the 8 recorded
turns with no unclean accounting are what rule out narration having been
skipped.
The WP-E rows in this run are the same ten as the standalone capture in §H,
re-measured with a narrator present and a full story on the page.
## M. Full regression
| Suite | Result |
| --- | --- |
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,054.7 s) |
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
| **Production build** (`npm run build`) | succeeded |
| **Browser harness** (M11 + WP-C + WP-E) | **101 passed, 0 failed, 0 skipped** (§L) |
| **Contrast audit** (`python tools/contrast_audit.py`) | **exit 0** — every text pair and every boundary pair passes |
**The backend count reconciles exactly.** WP-D's tree was 1,712; WP-E adds the
11 in `test_v11_e_contrast.py`. 1,712 + 11 = **1,723**. The 17 skips are the same
environment-gated real-model tests recorded since B.1 — WP-E used no model and
added no skip.
**The plan's regression requirements for WP-E** were the frontend suite and
lint, and the harness's accessibility checks. All three pass: 175 and exit 0
above, and M11's `A11y` rows — accessible names, visible keyboard focus, no
positive tabindex, nothing revealed only on hover, and the four rendered text
contrasts — are inside the 38/38 in §L.
**Lint detail.** The 15 warnings are the same pre-existing
`only-export-components` and unused-import kind recorded at WP-C and WP-D, and
**none is in a file WP-E changed**. `tokens.css` and `context.css` are not
flagged.
## N. Residual risks
1. **The screenshots are unapproved.** The measurements say every boundary now
clears 3:1; whether the result *looks* right in this design is a judgment
the numbers cannot make. Recorded as PENDING in §O, not assumed.
2. **Five controls are measured in the browser; the rest inherit.** The audit
checks token pairs, and the harness measures the composer, a story control,
the open panel tab, the nav edge and the focused composer. Every other
bordered surface — modals, cards, the knowledge and context panels — draws
the same two tokens, so it moves with them, but none is individually
measured. A component that overrides a border with a literal colour would
not be caught by either check.
3. **`--bg-panel-glass` is outside the token audit by nature.** It is rgba over
a gradient, so no token pair can express it; its edge is covered only by the
browser measurement, which runs in the harness rather than in CI.
4. **Disabled controls are deliberately not measured.** `.story-controls
button:disabled` carries `opacity: 0.35`, so a disabled control's rendered
edge is dimmer than any value here. WCAG 1.4.11 exempts inactive components,
and the harness selects `:not(:disabled)` on purpose — stated so that the
exclusion is visible rather than looking like an oversight.
5. **The chart set was re-checked only where WP-E disturbed it.** `.slice-7`'s
separation and colour-blind distance were measured against the other seven
(§J); the set as a whole was not re-audited, which is outside this package.
6. **`a11y.test.jsx` was not extended**, though the plan listed it as likely
affected. It asserts structure and names, not colours, and adding colour
assertions in jsdom would test the stylesheet's text rather than a rendered
result. Boundary contrast is asserted instead where it can be measured: the
gate tests (§F) and the browser (§G).
**One risk that turned out not to exist.** §G's rule — a boundary passes on the
better of its two adjacent colours — was written to avoid failing an edge that
contrasts with the page but matches its own fill. In the event it never did any
work: every measured boundary clears 3:1 against **both** neighbours (composer
4.55/4.88, story control 3.44/3.70, panel tab 4.24/4.55, hover 4.55/4.88, focus
4.38/4.70, nav 3.51/3.70). The stricter reading would have produced the same
verdict on every row.
## O. Final decision
**Against the plan's acceptance criteria** (§ WP-E, *Control-boundary contrast*):
| # | Criterion | Verdict |
| --- | --- | --- |
| 1 | `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and 0 on the package tree | **PASS** — 2.99:1 exits 1 and 3.00:1 exits 0 (§F); the v1.0.0 value exits 1; the package tree exits 0 |
| 2 | Every text pair still clears 4.5:1, and the rendered text contrasts do not fall below v1's 14.57 / 5.48 / 13.57 / 5.88 | **PASS** — asserted as exact baselines in §F and re-measured on the rendered page in §L's M11 rows. No text token changed |
| 3 | The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1; the focus indicator is still visible | **PASS** — composer 4.88:1, story control 3.70:1 at rest and 4.88:1 on hover (§H), and M11's visible-focus check passes in §L |
| 4 | The owner approves the before-and-after screenshots, and the report records the approval | **PENDING** — not a criterion I can satisfy. The four PNGs are in §I |
**Where I exceeded the criterion, said plainly.** The plan asks for 3:1 "against
their panel" and names two controls. This package measures against
**`--bg-input`** as well — the lighter background inputs and buttons are really
drawn on, and the one that decides the gate — and adds the open panel tab, the
translucent navigation edge and the focused state. The stricter floor is the
reason the two threshold tests in §F are decided by `--bg-input` rather than
`--bg-panel`, where both would have passed.
**What I got wrong and corrected.** Three of my own claims failed checking and
were fixed rather than softened: the first hover and focus measurements were
taken mid-transition and reported colours no state has (§G); two contrast
figures were written into `context.css` before being measured, and were wrong
(§J); and my first gate check reported the current tokens as failing because my
harness crashed on a shallow path, not because of any contrast.
**Not done, and not claimed:** owner approval (criterion 4), individual
measurement of every bordered component (§N.2), and any palette work beyond the
two boundary tokens and the one chart fill that borrowed from them.
```text
WP-E CONTRAST GATE: PASS
WP-E BOUNDARY MEASUREMENT: PASS
WP-E OVERALL:
PASS, pending owner approval of the screenshots (criterion 4)
```
All WP-E changes are **staged and uncommitted**. No commit, no push, no tag.
v1.1 release validation has not begun.
```text
OWNER SCREENSHOT APPROVAL: APPROVED
```
The before/after screenshots in §I are the evidence for a change a reader judges
by looking at it. The measurements say every boundary now clears 3:1; whether the
result looks right in this design was the owner's call.
**Approved by the owner on 2026-09-16**, in the v1.1 release-validation brief,
after reviewing the before/after pair in `$HOME/v11-evidence/wp-e/`. This line
was `PENDING` in the signed commit `87a4032` because the report predated that
review; it is updated here as part of the release closeout, with its source and
date recorded rather than the approval being assumed. No visual code changed
during release validation, so the approval stands (release report §Q).