Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
db7b309e3d | ||
|
|
87a40326a2 | ||
|
|
59b5ebc2d8 | ||
|
|
0c1ba836ba | ||
|
|
beb17ada10 | ||
|
|
d63804f22e | ||
|
|
ac465ed867 |
+106
-5
@@ -408,9 +408,17 @@ AIDND_TEST_ENDPOINT=... AIDND_TEST_MODEL=... AIDND_TEST_EMBED_MODEL=... \
|
||||
# What that campaign is worth on a machine that has never seen it (I01-I07).
|
||||
.venv/bin/python -m tools.m11_recovery --bundle "$HOME/m11-evidence/m01/bundle.json" --out "$HOME/m11-evidence/m01"
|
||||
|
||||
# The browser release regression and the accessibility measurements. Needs
|
||||
# `frontend/dist` built and geckodriver on PATH.
|
||||
.venv/bin/python -m tools.m11_browser --out "$HOME/m11-evidence/browser"
|
||||
# The browser release regression and the accessibility measurements: M11's 38
|
||||
# checks plus v1.1 WP-C's reader workflows (Retry, Save Points, state correction,
|
||||
# narration length, failed generation, export download). Needs `frontend/dist`
|
||||
# built, geckodriver on PATH, and --out under $HOME (the downloads land inside
|
||||
# it). Release evidence needs the narrator over trusted-LAN HTTPS.
|
||||
AIDND_TEST_ENDPOINT=https://... AIDND_TEST_MODEL=qwen2.5:3b-instruct \
|
||||
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/browser"
|
||||
# Without a narrator (a partial smoke run, not evidence), or one scenario while
|
||||
# developing (--only takes: shell, history, markdown, hidden, context, csp, a11y,
|
||||
# retry, savepoint, state, length, failure, export).
|
||||
.venv/bin/python -m tools.m11_browser --out "$HOME/v11-evidence/wp-c/smoke" --no-narrator
|
||||
|
||||
# A container with no network at all: the offline run and the packaging path.
|
||||
.venv/bin/python -m tools.m11_offline --out "$HOME/m11-evidence/offline"
|
||||
@@ -473,8 +481,11 @@ nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,p
|
||||
# The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling
|
||||
nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/gpu-dmon-$(date +%F-%H%M).log"
|
||||
|
||||
# Kernel and Ollama messages, live
|
||||
journalctl -f -k -u ollama | tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log"
|
||||
# Kernel and Ollama messages, live. The `+` is an OR: `journalctl -k -u ollama`
|
||||
# asks for messages that are both kernel messages and the ollama unit's, which
|
||||
# is none, and writes an empty log.
|
||||
journalctl -f -o short-iso _TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service \
|
||||
| tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log"
|
||||
```
|
||||
|
||||
If the GPU faults, find the moment and then read what the card was doing just
|
||||
@@ -496,6 +507,33 @@ It exists so browser evidence needs no Selenium in the dependency surface, and
|
||||
it documents the one environment quirk that matters here: a snap Firefox will
|
||||
not open a file the driver names under `/tmp`, but will under `$HOME`.
|
||||
|
||||
**Downloads in the browser harness (v1.1 WP-C).** The export checks click the
|
||||
real Export controls and wait for the file on disk, so the browser has to save
|
||||
without asking. `m11_webdriver.firefox_download_prefs` gives the WebDriver
|
||||
session a profile that does that:
|
||||
- `browser.download.folderList` 2, `browser.download.dir` the run's
|
||||
`downloads/` folder, `browser.download.useDownloadDir` true;
|
||||
- no "always ask", and `application/json` saved to disk.
|
||||
|
||||
It works on the snap Firefox this machine has (155.0.1, geckodriver 0.37.1), and
|
||||
no separate Firefox is needed. The same sandbox rule applies as for opening
|
||||
files: the download folder must be under `$HOME`, and the harness refuses one
|
||||
that is not.
|
||||
|
||||
A download counts as finished only when all of these hold at once
|
||||
(`m11_webdriver.wait_for_download`):
|
||||
- a new name has appeared;
|
||||
- no `*.part` file is left;
|
||||
- the file is more than zero bytes;
|
||||
- its size is the same across consecutive polls.
|
||||
|
||||
The toast that says "Campaign exported." is not evidence.
|
||||
|
||||
**Waiting.** Nothing in the harness sleeps before an assertion. Every wait is on
|
||||
something the page, the browser or the filesystem shows. A condition that
|
||||
already holds before the action it waits for does not count as waiting for that
|
||||
action; the harness defects found in M8, M11 and WP-C were all of that shape.
|
||||
|
||||
## Backing up, and getting a campaign back
|
||||
|
||||
There are two recovery tools and they answer different questions. Using the
|
||||
@@ -510,6 +548,40 @@ together.
|
||||
| Taken from | Export, on a campaign | Settings → *Back up everything on this machine* |
|
||||
| Restored by | Import campaign, on the library screen | replacing the database file, below |
|
||||
|
||||
### How large an export can get
|
||||
|
||||
The importer accepts a request body up to **20 MB**
|
||||
(`backend/app/limits.py`, `MAX_IMPORT_BODY_BYTES`), and v1.1 does not change it.
|
||||
What that means for a campaign, measured rather than guessed:
|
||||
|
||||
- the M11 evidence campaign came to roughly **13 kB per action** in its bundle;
|
||||
- M9's conservative estimate from that figure is about **279 turns** before a
|
||||
bundle approaches the limit.
|
||||
|
||||
Both are measurements of particular campaigns, **not a turn limit**. What a
|
||||
campaign actually weighs depends on how long its turns are, how much imported
|
||||
knowledge travels with it, and how many attempts each turn kept. A campaign of
|
||||
400 short turns can be well inside the limit; one of 200 long ones with a large
|
||||
library may not be.
|
||||
|
||||
**v1.1 (WP-D) makes the individual case visible.** Every export reports its own
|
||||
serialised size and whether this version could import it back:
|
||||
|
||||
```text
|
||||
X-Export-Bytes the bundle's size, as the importer would weigh it
|
||||
X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
|
||||
X-Importable-By-This-Version true / false
|
||||
X-Export-Warning present only when it is false
|
||||
```
|
||||
|
||||
The export always succeeds and the file is always delivered — it is complete and
|
||||
undamaged; what it exceeds is this version's import ceiling. Both Export
|
||||
controls show the warning when there is one. The size compared is the compact
|
||||
serialisation the browser would POST back, which is smaller than the
|
||||
pretty-printed file on disk.
|
||||
|
||||
Raising the limit, or streaming an import past it, is deferred to v1.2.
|
||||
|
||||
### Exporting and importing a campaign
|
||||
|
||||
Export is on each campaign in the library, and in the campaign's own Settings
|
||||
@@ -620,6 +692,35 @@ campaign gets less history than the setting asks for, which is a visible,
|
||||
explicable loss rather than a silent one, and Settings' **Test connection**
|
||||
reports the window it found or says plainly that it could not check.
|
||||
|
||||
**It also keeps a margin, and checks the server's own count (v1.1).** The
|
||||
application counts tokens with `cl100k_base`, and your model counts them with
|
||||
its own tokenizer. The two disagree slightly, so the prompt is built to leave
|
||||
`max(256, 5% of the window)` tokens free on top of the reply: 256 at 4,096, and
|
||||
820 at 16,384. After each turn the server's reported prompt-token count is
|
||||
compared with what was sent. The context inspector shows the result for any
|
||||
past turn:
|
||||
|
||||
- **The server read the whole prompt:** the ordinary case.
|
||||
- **The server did not say how much it read:** the server reported no usage.
|
||||
Nothing is wrong, and nothing is confirmed either.
|
||||
- **The server may have cut the start of the prompt:** it read far fewer tokens
|
||||
than were sent. Ollama does this, silently, to a prompt larger than the window
|
||||
the model was loaded with. The turn is kept. Check the window with the
|
||||
commands above.
|
||||
- **The prompt was larger than the server allowed for:** its count and the reply
|
||||
together exceed the window. The reply may have been cut short. The turn is
|
||||
kept.
|
||||
|
||||
The last two also appear in the server log as a warning.
|
||||
|
||||
**A model that is not loaded yet is loaded first.** Before a turn, if the
|
||||
application cannot read the window because your model isn't in memory, it asks
|
||||
the same Ollama to load it once. That is a `POST /api/generate` naming only the
|
||||
model, which generates no text. It then reads the window again, so the first
|
||||
turn of a session is built to the window the model really has rather than to
|
||||
your setting. If loading fails, or the window still can't be read, the turn goes
|
||||
ahead exactly as before, unverified, and the check above still applies.
|
||||
|
||||
That does not make the window *bigger*, and the rest of this section is still
|
||||
how you do that.
|
||||
|
||||
|
||||
@@ -80,6 +80,14 @@ that isn't the live one starts a new branch.
|
||||
in settings lets you state the window so the prompt is still capped. A window the server
|
||||
itself reported always wins over that, and a declared one is never reported as verified.
|
||||
|
||||
Since v1.1 the prompt also stops short of that window on purpose. It leaves
|
||||
`max(256, 5% of the window)` tokens free, because your model counts tokens differently from
|
||||
the application, and the v1 evidence came within 23 tokens of the edge. After each turn,
|
||||
the server's own count of what it read is compared with what was sent. A turn the server
|
||||
appears to have truncated is kept, flagged and shown in the context inspector, not left to
|
||||
pass silently. A model that isn't loaded yet, and so cannot report its window, is loaded
|
||||
once before the turn is built, so the first turn of a session gets the real window too.
|
||||
|
||||
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
|
||||
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
|
||||
card used to arrive in front of it as a world fact with no class, no visibility, no source and
|
||||
@@ -179,9 +187,9 @@ that isn't the live one starts a new branch.
|
||||
|
||||
None yet. The inherited screenshots showed upstream's UI — a Scripts tab, Log in and Sign up,
|
||||
a guest banner, scripting demo scenarios — none of which this fork has since M2, so they were
|
||||
removed rather than left standing as a picture of a product that no longer exists. The M4
|
||||
closeout drove the real application in a real browser, so the screens exist and work; taking
|
||||
presentable screenshots of them is a job for the UI pass in M8.
|
||||
removed rather than left standing as a picture of a product that no longer exists. The
|
||||
screens exist and are driven in a real browser by the release harness
|
||||
(`backend/tools/m11_browser.py`). Presentable screenshots of them have not been taken.
|
||||
|
||||
## Quick start
|
||||
|
||||
@@ -285,7 +293,7 @@ player input
|
||||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||||
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
|
||||
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile
|
||||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
|
||||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (94 and counting)
|
||||
├─ endpoints.py the inference-endpoint address policy
|
||||
├─ contextwindow.py what the server will actually accept, and the cap
|
||||
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
|
||||
@@ -310,7 +318,8 @@ development, Vite proxies `/api` to FastAPI.
|
||||
|
||||
## Tests
|
||||
|
||||
1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||||
1,524 backend tests (1,507 run everywhere, 17 need a real local model and skip without one): unit
|
||||
tests plus full HTTP integration through the real turn engine, with
|
||||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||||
proves nothing. A further handful need a real local model and skip without one; they exist
|
||||
@@ -346,9 +355,20 @@ most interesting engineering in the repo.
|
||||
|
||||
## Repo notes
|
||||
|
||||
- **Status:** milestones M1-M11 are complete. The v1 release gate passed on
|
||||
2026-09-14 (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
|
||||
§T). Passing the gate is not a release: there is no `v1.0.0` tag yet.
|
||||
- **Status:** **v1.0.0 remains the released version.** Milestones M1-M11 are
|
||||
complete, and the v1 release gate passed (see [`planning/reports/M11-IMPLEMENTATION-REPORT.md`](planning/reports/M11-IMPLEMENTATION-REPORT.md),
|
||||
§T). The signed tag `v1.0.0` and `main` both point at the signed release
|
||||
commit `432f041`.
|
||||
|
||||
**v1.1 is implemented and validated, but not yet released.** All six work
|
||||
packages (WP-A1, WP-A2, WP-B, WP-C, WP-D, WP-E) are complete and accepted on
|
||||
the `v1.1-development` branch, and integrated release validation passed on
|
||||
candidate `87a4032` — see
|
||||
[`planning/reports/v1.1/V1.1-RELEASE-REPORT.md`](planning/reports/v1.1/V1.1-RELEASE-REPORT.md).
|
||||
WP-B ships with a documented reference-model memory limitation, recorded in
|
||||
that report. **No `v1.1.0` tag exists and `main` is unchanged**; the release
|
||||
commit, `main` and the tag are the owner's to make. The plan is
|
||||
[`planning/V1.1-PLAN.md`](planning/V1.1-PLAN.md).
|
||||
|
||||
- `planning/` is this fork's own package: the product specification, the architecture
|
||||
decisions, the milestone plan, the acceptance contract, and a review report for every
|
||||
|
||||
@@ -46,7 +46,14 @@ from .narrative import model as narrative_model
|
||||
# token accounting. Each attempt is its own API call, and a retry is the call
|
||||
# most likely to read the prompt back out of cache. Everything else in a snapshot
|
||||
# is the prompt, which is assembled once per turn.
|
||||
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage")
|
||||
#
|
||||
# v1.1 WP-A1: `accounting` is one attempt's too. It compares the server's count
|
||||
# for *that* call with the turn's estimate. Left out of this tuple, it was
|
||||
# treated as part of the shared prompt, so moving the live flag handed the
|
||||
# superseded attempt's accounting to the new live one and threw the new one's
|
||||
# away. Found by the A2 long run: two retries and one take selection left three
|
||||
# attempts reporting no accounting, or another attempt's.
|
||||
ATTEMPT_KEYS = ("world_state", "narrative_state", "raw_output", "usage", "accounting")
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ reading
|
||||
|
||||
+20
-9
@@ -39,8 +39,9 @@ turn is blocked.
|
||||
never leaves a half-written file wearing a backup's name. `os.replace` is
|
||||
atomic on the same filesystem, which is why the temporary sits in the
|
||||
destination's own directory rather than in `/tmp`.
|
||||
3. `PRAGMA quick_check` runs against the finished copy, opened as its own
|
||||
database, before it is renamed. A backup nobody verified is a belief.
|
||||
3. `PRAGMA integrity_check` runs against the finished copy, opened as its own
|
||||
database, before it is renamed. A backup nobody verified is a belief. v1.1
|
||||
WP-D made this the full check rather than `quick_check`; see `_verify`.
|
||||
4. An existing file is never overwritten. Each run writes a new name stamped
|
||||
with the time, so yesterday's backup survives today's mistake — which is most
|
||||
of what a backup is for.
|
||||
@@ -189,18 +190,28 @@ def _copy(source_path: Path, working: Path) -> int:
|
||||
|
||||
|
||||
def _verify(working: Path) -> str:
|
||||
"""Runs `PRAGMA quick_check` against the finished copy.
|
||||
"""Runs `PRAGMA integrity_check` against the finished copy.
|
||||
|
||||
Opened as its own connection, so what is checked is the file on disk rather
|
||||
than any page cache the copy left behind. `quick_check` rather than
|
||||
`integrity_check` because it does the structural work — every page reachable,
|
||||
every record readable — without the full index cross-check, which on a large
|
||||
database is minutes rather than moments. A backup nobody verified is a
|
||||
belief; a backup verified slowly enough that nobody takes one is worse.
|
||||
than any page cache the copy left behind.
|
||||
|
||||
**v1.1 WP-D: the full check, not `quick_check`.** M9 chose `quick_check` for
|
||||
its speed, on the argument that a backup verified slowly enough that nobody
|
||||
takes one is worse than a fast one. The measurements say the trade was not
|
||||
needed here: `quick_check` omits the cross-check between a table and its
|
||||
indexes, and that is a real class of damage it reports as `ok`. A copy whose
|
||||
index disagrees with its table restores into a database that answers queries
|
||||
with rows that are not there — the failure a backup exists to prevent.
|
||||
|
||||
The cost is small at the sizes this application produces: on the 100-turn
|
||||
evidence campaign both checks are a few milliseconds, and on a synthetic
|
||||
database two orders of magnitude larger the difference is still short of a
|
||||
second (WP-D report §E). A backup nobody verified is a belief; this is the
|
||||
check that makes it a fact.
|
||||
"""
|
||||
connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute("PRAGMA quick_check").fetchall()
|
||||
rows = connection.execute("PRAGMA integrity_check").fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
result = ", ".join(str(row[0]) for row in rows) if rows else "no result"
|
||||
|
||||
@@ -25,6 +25,7 @@ from sqlalchemy.orm import object_session
|
||||
|
||||
from .. import contextwindow, derived, models, narrative, summaries, worldstate
|
||||
from ..knowledge import inject as knowledge_inject
|
||||
from ..providers.openai_compatible import CHAT_CONTINUE_HINT
|
||||
from ..knowledge import records as knowledge_records
|
||||
from . import encoding, history
|
||||
|
||||
@@ -106,11 +107,22 @@ BAND_FLOOR_SHARE = 0.5
|
||||
# Built from the table vendored in `encoding.py`, not fetched: the upstream
|
||||
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
|
||||
# is called on every turn.
|
||||
# M6: added to the configured reply budget when reserving output space. It
|
||||
# absorbs the section separators added after budgeting and the drift between
|
||||
# this tokenizer and the serving model's. Fixed rather than proportional: what
|
||||
# it covers does not grow with the size of the budget.
|
||||
OUTPUT_SAFETY_MARGIN = 64
|
||||
#
|
||||
# v1.1 WP-A1: `OUTPUT_SAFETY_MARGIN = 64` was here. M6 added it to the reply
|
||||
# budget to absorb two unrelated things, and v1.1 separates them:
|
||||
#
|
||||
# * **Text the application adds after pricing.** The separators between
|
||||
# sections, and `CHAT_CONTINUE_HINT`, which the provider appends to every chat
|
||||
# request and nothing counted. That is not drift, it is our own text, so it is
|
||||
# now priced exactly (`transport` below).
|
||||
# * **The drift between this tokenizer and the narrator's.** That is what the
|
||||
# 64 tokens were really for, and the v1 evidence showed it was too small. It
|
||||
# is now `contextwindow.safety_reserve`, sized to the window.
|
||||
#
|
||||
#: Story sections that can be joined by `SEPARATOR` after pricing: history,
|
||||
#: author's note, recent history, summary, lore, memories, state, front memory,
|
||||
#: length hint, refusals, reminder. Knowledge and history rows price their own.
|
||||
STORY_SECTION_SLOTS = 11
|
||||
|
||||
|
||||
class ContextOverflow(RuntimeError):
|
||||
@@ -180,19 +192,19 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
words = min(words, band_ceiling)
|
||||
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
|
||||
tail = (
|
||||
" Finish the narration and append the state block well inside the limit."
|
||||
" " + narrative.extract.LENGTH_HINT_TAIL
|
||||
)
|
||||
if floor < MIN_LENGTH_FLOOR_WORDS:
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
|
||||
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
||||
)
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
|
||||
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
||||
f"the scene genuinely needs more.{tail}]"
|
||||
)
|
||||
tail = " Finish the narration and append the state block well inside the limit."
|
||||
tail = " " + narrative.extract.LENGTH_HINT_TAIL
|
||||
|
||||
# State the number as a ceiling, never as a budget. In measurements, the
|
||||
# wording "keep this turn under about N words" read to the model as a target
|
||||
@@ -204,7 +216,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
|
||||
if floor < MIN_LENGTH_FLOOR_WORDS:
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words. Write only as "
|
||||
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
||||
)
|
||||
# Both numbers are bounds, and the wording is deliberately asymmetric. The
|
||||
@@ -216,7 +228,7 @@ def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
|
||||
# so a terse model reading the same clause stops at the floor rather than at
|
||||
# forty words.
|
||||
return (
|
||||
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
||||
f"{narrative.extract.LENGTH_HINT_OPENING} this turn must not exceed {words} words, and it should not "
|
||||
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
||||
f"the scene genuinely needs more.{tail}]"
|
||||
)
|
||||
@@ -625,12 +637,19 @@ def build_context(
|
||||
# truncated turn on a model whose window is the budget
|
||||
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
|
||||
#
|
||||
# The margin covers what is added after this arithmetic — the separators
|
||||
# between sections, and the difference between our tokenizer's count and the
|
||||
# serving model's. It is small and fixed rather than proportional, because
|
||||
# what it absorbs does not scale with the budget.
|
||||
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
|
||||
protected = reserved + output_reserve
|
||||
# v1.1 WP-A1: the reply allocation is exactly the reply cap. The text this
|
||||
# application adds after pricing — separators, and the chat hint the
|
||||
# provider appends — is counted as `transport`. What neither can know, the
|
||||
# narrator's tokenizer disagreeing with `cl100k_base`, is the safety reserve,
|
||||
# which is sized to the window and taken before any history is chosen.
|
||||
output_reserve = max(0, settings.max_output_tokens)
|
||||
separator_tokens = count_tokens(SEPARATOR)
|
||||
transport = (
|
||||
separator_tokens * (len(system_sections) + STORY_SECTION_SLOTS)
|
||||
+ count_tokens(CHAT_CONTINUE_HINT)
|
||||
)
|
||||
safety = contextwindow.safety_reserve(budget)
|
||||
protected = reserved + transport + output_reserve + safety
|
||||
if protected >= budget:
|
||||
# Failing here is the point. The alternative — carrying on with a token
|
||||
# or two of history — builds a prompt that is known to overflow, and
|
||||
@@ -638,8 +657,9 @@ def build_context(
|
||||
# gracefully if protected context alone is too large."
|
||||
raise ContextOverflow(
|
||||
f"The protected context needs {protected} tokens "
|
||||
f"({reserved} of prompt plus {output_reserve} reserved for the "
|
||||
f"reply) but the context budget is {budget}. "
|
||||
f"({reserved} of prompt, {transport} of formatting, {output_reserve} "
|
||||
f"reserved for the reply and a {safety}-token safety margin) but the "
|
||||
f"context budget is {budget}. "
|
||||
+ (
|
||||
"That budget is what this server was found to accept, so raising "
|
||||
"the setting alone will not help — load the model with a larger "
|
||||
@@ -827,6 +847,7 @@ def build_context(
|
||||
story_text = SEPARATOR.join(s.text for s in story_sections)
|
||||
|
||||
all_sections = [s for s in system_sections if s.text] + story_sections
|
||||
total_tokens = count_tokens(system_text) + count_tokens(story_text)
|
||||
report = {
|
||||
"sections": [
|
||||
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
|
||||
@@ -837,13 +858,23 @@ def build_context(
|
||||
# what the history was actually allowed to spend after everything
|
||||
# protected was subtracted.
|
||||
"tokens": {
|
||||
"total": count_tokens(system_text) + count_tokens(story_text),
|
||||
"total": total_tokens,
|
||||
"budget": budget,
|
||||
"configured_budget": settings.context_token_budget,
|
||||
"output_reserve": output_reserve,
|
||||
"protected": reserved,
|
||||
"available_for_history": available,
|
||||
"history_spent": spent,
|
||||
# v1.1 WP-A1. `transport` is the formatting priced in above;
|
||||
# `estimate` is what this application believes it actually sent,
|
||||
# the assembled text plus what the provider adds to it, and is what
|
||||
# the server's own count is compared against after the reply.
|
||||
"transport": transport,
|
||||
"safety_reserve": safety,
|
||||
"estimate": total_tokens + (
|
||||
count_tokens(CHAT_CONTINUE_HINT) if settings.api_mode != "completion"
|
||||
else separator_tokens
|
||||
),
|
||||
},
|
||||
# M11: what the server was found to accept, and how. `verified` false
|
||||
# means nobody could check — the prompt was built to the configured
|
||||
|
||||
@@ -86,6 +86,7 @@ one.
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import math
|
||||
import re
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
@@ -132,6 +133,11 @@ class Window:
|
||||
model_max: int | None = None
|
||||
#: Why the window is unknown, or how it was found. Shown to the user.
|
||||
detail: str = ""
|
||||
#: v1.1: the server answered a discovery request at all, whatever it said.
|
||||
#: A server that answered but could not report a window may simply not have
|
||||
#: the model loaded yet, which `ensure_window` can fix; one that did not
|
||||
#: answer cannot be helped by asking it to load anything.
|
||||
reachable: bool = False
|
||||
|
||||
@property
|
||||
def verified(self) -> bool:
|
||||
@@ -180,6 +186,116 @@ def effective_budget(configured: int, window: Window | int | None) -> int:
|
||||
return min(configured, tokens)
|
||||
|
||||
|
||||
#: v1.1 WP-A1: the tokens kept free below the effective window, beyond the reply.
|
||||
#:
|
||||
#: The builder counts with `cl100k_base`; the narrator counts with its own
|
||||
#: tokenizer. The v1 evidence put the largest prompts 23-42 real tokens from the
|
||||
#: edge of a 16,384 window, and Ollama does not refuse a prompt past the edge —
|
||||
#: measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt came back 200
|
||||
#: with `prompt_tokens` 2,050. So the reserve is deliberate and sized to the
|
||||
#: window: the larger of a floor and a share, **rounded up to a whole token**.
|
||||
#:
|
||||
#: 4,096 -> 256 8,192 -> 410 16,384 -> 820
|
||||
#:
|
||||
#: A fixed, documented tolerance, owner-chosen for v1.1. It is not a setting and
|
||||
#: it is not calibrated per model.
|
||||
SAFETY_RESERVE_FLOOR = 256
|
||||
SAFETY_RESERVE_PERCENT = 5
|
||||
|
||||
|
||||
def safety_reserve(effective_window: int) -> int:
|
||||
"""`max(256, ceil(5% of the effective window))`, in tokens.
|
||||
|
||||
The effective window is the budget the prompt is actually built to — the
|
||||
verified or declared window when there is one, the configured budget
|
||||
otherwise — so a 16,384 setting against a 4,096 server reserves 256, not 820.
|
||||
Integer arithmetic, so the rounding is exact rather than a float's.
|
||||
"""
|
||||
share = math.ceil(max(0, effective_window) * SAFETY_RESERVE_PERCENT / 100)
|
||||
return max(SAFETY_RESERVE_FLOOR, share)
|
||||
|
||||
|
||||
#: v1.1 WP-A1: what the server's own count says about a turn that was sent.
|
||||
FITS = "fits"
|
||||
EXCEEDED = "exceeded"
|
||||
TRUNCATION_SUSPECTED = "truncation_suspected"
|
||||
#: `UNKNOWN` above: the server reported no usable count.
|
||||
|
||||
|
||||
def classify_usage(usage: dict | None, *, estimate: int, budget: int,
|
||||
max_output_tokens: int, window_verified: bool) -> dict:
|
||||
"""Sets the server's reported prompt count against what the application sent.
|
||||
|
||||
The order of the checks is the order of what they prove:
|
||||
|
||||
``unknown``
|
||||
No positive integer `prompt_tokens`. Nothing can be said, and nothing
|
||||
is claimed: an absent count is never read as a prompt that fitted.
|
||||
``truncation_suspected``
|
||||
The server read fewer tokens than were sent by more than the safety
|
||||
reserve. A tokenizer thriftier than `cl100k_base` may honestly count a
|
||||
little less; a shortfall larger than the tolerance the application keeps
|
||||
for drift is the signature of a server that cut the prompt — the real
|
||||
shape was 6,316 sent and 2,050 read.
|
||||
``exceeded``
|
||||
The server's count plus the reply allocation is more than the window
|
||||
the prompt was built for. The drift was larger than the whole reserve,
|
||||
so the reply may have been cut short.
|
||||
``fits``
|
||||
Otherwise.
|
||||
|
||||
`observed_margin` is what was left beside the reply by the server's count:
|
||||
`budget - max_output_tokens - server_prompt_tokens`. The safety reserve is
|
||||
the tolerance, so a margin between 0 and the reserve is still `fits`.
|
||||
|
||||
A discrepancy is recorded, never acted on: the reply has already streamed
|
||||
to the reader and is accepted story.
|
||||
"""
|
||||
prompt = usage.get("prompt_tokens") if isinstance(usage, dict) else None
|
||||
reserve = safety_reserve(budget)
|
||||
verified_note = "" if window_verified else (
|
||||
" The window itself was not verified for this turn.")
|
||||
record = {
|
||||
"status": UNKNOWN,
|
||||
"server_prompt_tokens": None,
|
||||
"estimate": estimate,
|
||||
"difference": None,
|
||||
"budget": budget,
|
||||
"max_output_tokens": max_output_tokens,
|
||||
"safety_reserve": reserve,
|
||||
"observed_margin": None,
|
||||
"window_verified": bool(window_verified),
|
||||
"detail": "",
|
||||
}
|
||||
if type(prompt) is not int or prompt <= 0:
|
||||
record["detail"] = ("The server reported no prompt token count, so nothing "
|
||||
"confirms the whole prompt was read." + verified_note)
|
||||
return record
|
||||
|
||||
record["server_prompt_tokens"] = prompt
|
||||
record["difference"] = prompt - estimate
|
||||
record["observed_margin"] = budget - max_output_tokens - prompt
|
||||
if prompt + reserve < estimate:
|
||||
record["status"] = TRUNCATION_SUSPECTED
|
||||
record["detail"] = (
|
||||
f"The server read {prompt:,} prompt tokens of the {estimate:,} sent, a "
|
||||
f"shortfall larger than the {reserve:,}-token safety reserve. A server "
|
||||
"that cuts an over-window prompt reports exactly this, and what it cuts "
|
||||
"is the start: the narrator's rules and the canon." + verified_note)
|
||||
elif prompt + max_output_tokens > budget:
|
||||
record["status"] = EXCEEDED
|
||||
record["detail"] = (
|
||||
f"The server counted {prompt:,} prompt tokens; with {max_output_tokens:,} "
|
||||
f"for the reply that is more than the {budget:,}-token window the prompt "
|
||||
"was built for, so the reply may have been cut short." + verified_note)
|
||||
else:
|
||||
record["status"] = FITS
|
||||
record["detail"] = (
|
||||
f"The server read {prompt:,} prompt tokens, leaving "
|
||||
f"{record['observed_margin']:,} beside the reply." + verified_note)
|
||||
return record
|
||||
|
||||
|
||||
def cache_clear() -> None:
|
||||
"""Forgets what was learned. Called when the endpoint or model changes."""
|
||||
_cache.clear()
|
||||
@@ -223,6 +339,7 @@ def _declared_or(declared: int | None, discovered: Window) -> Window:
|
||||
declared, DECLARED, discovered.model_max,
|
||||
f"{declared:,} tokens, declared in settings — the server was not able "
|
||||
f"to say ({discovered.detail})",
|
||||
reachable=discovered.reachable,
|
||||
)
|
||||
|
||||
|
||||
@@ -266,6 +383,7 @@ async def _ask(endpoint_url: str, model: str) -> Window:
|
||||
return Window(
|
||||
tokens, LOADED, ceiling,
|
||||
f"{tokens:,} tokens, reported by the running model",
|
||||
reachable=True,
|
||||
)
|
||||
return await _declared_window(client, base, model)
|
||||
except (httpx.HTTPError, ValueError, TypeError, KeyError) as exc:
|
||||
@@ -293,6 +411,7 @@ async def _declared_window(client, base: str, model: str) -> Window:
|
||||
return Window(
|
||||
None, UNKNOWN,
|
||||
detail=f"the server did not describe the model (HTTP {resp.status_code})",
|
||||
reachable=True,
|
||||
)
|
||||
body = resp.json() or {}
|
||||
ceiling = _architecture_ceiling(body.get("model_info") or {})
|
||||
@@ -304,14 +423,103 @@ async def _declared_window(client, base: str, model: str) -> Window:
|
||||
"the model sets no num_ctx, so the server will load it at its own "
|
||||
"default — which is 4,096 where there is no VRAM"
|
||||
),
|
||||
reachable=True,
|
||||
)
|
||||
tokens = min(declared, ceiling) if ceiling else declared
|
||||
return Window(
|
||||
tokens, PARAMETERS, ceiling,
|
||||
f"{tokens:,} tokens, from the model's own num_ctx",
|
||||
reachable=True,
|
||||
)
|
||||
|
||||
|
||||
#: v1.1 WP-A1 corrective: loading the configured model so its window can be read.
|
||||
#:
|
||||
#: The first real turn of the A1 evidence found a cold model: `/api/ps` knew
|
||||
#: nothing, `/api/show` found no `num_ctx`, so the window was unverified and the
|
||||
#: prompt was built to the configured 16,384. Ollama loaded the model at its own
|
||||
#: 4,096 default, kept 2,050 of 13,875 tokens and answered 200. That case is
|
||||
#: preventable, because the window becomes readable the moment the model is
|
||||
#: resident. Ollama's native `POST /api/generate` with a model and **no prompt**
|
||||
#: loads the model and generates nothing — measured on Ollama 0.33: HTTP 200,
|
||||
#: `"response": ""`, `"done_reason": "load"`, and `/api/ps` then reported the
|
||||
#: window. The OpenAI-compatible request that followed did not reload it.
|
||||
WARM_PATH = "/api/generate"
|
||||
|
||||
|
||||
async def warm(endpoint_url: str, model: str, *, timeout: float) -> tuple[bool, str]:
|
||||
"""Asks the configured server to load `model`. One request, no story text.
|
||||
|
||||
Held to the same endpoint policy and TLS trust as inference and the probe, and
|
||||
sent to the same host the probe asks. The body names the model and nothing
|
||||
else: no prompt, so nothing is generated, and no `options` or `keep_alive`, so
|
||||
the model loads the way the server would load it for the turn itself.
|
||||
|
||||
Returns `(loaded, detail)`. Every failure is `(False, why)` and never raises:
|
||||
a server that will not load the model on request will fail the turn's own
|
||||
call the ordinary way, which is where that failure belongs.
|
||||
"""
|
||||
reason = endpoints.rejection_reason(endpoint_url)
|
||||
if reason is not None:
|
||||
return False, f"endpoint not allowed — {reason}"
|
||||
base = native_base(endpoint_url)
|
||||
try:
|
||||
async with httpx.AsyncClient(
|
||||
timeout=httpx.Timeout(timeout, connect=CONNECT_TIMEOUT),
|
||||
verify=tlstrust.ssl_context(),
|
||||
) as client:
|
||||
resp = await client.post(f"{base}{WARM_PATH}", json={"model": model})
|
||||
except httpx.HTTPError as exc:
|
||||
log.debug("model warm-up failed for %s: %s", base, exc)
|
||||
return False, f"could not ask the server to load the model ({type(exc).__name__})"
|
||||
if resp.status_code != 200:
|
||||
return False, f"the server did not load the model (HTTP {resp.status_code})"
|
||||
try:
|
||||
body = resp.json() or {}
|
||||
except ValueError:
|
||||
return False, "the server answered the load request with something that was not JSON"
|
||||
return True, f"the server loaded the model ({body.get('done_reason') or 'done'})"
|
||||
|
||||
|
||||
async def ensure_window(endpoint_url: str, model: str, *, declared: int | None = None,
|
||||
warm_timeout: float = 300.0) -> tuple[Window, dict]:
|
||||
"""The window for a turn about to be generated, loading the model once if that is what it takes.
|
||||
|
||||
1. Probe as before.
|
||||
2. If the window is not verified, the server answered, and there is a model to
|
||||
load: one bounded `warm` request.
|
||||
3. If the model loaded, probe again, bypassing the cache that still holds the
|
||||
unverified answer.
|
||||
|
||||
Whatever the second probe says is the answer. There is no retry loop, no
|
||||
guessed window, and no hard-coded 4,096: a window still unverified leaves the
|
||||
configured budget standing, exactly as before, and the turn's accounting
|
||||
still catches a server that cut the prompt.
|
||||
|
||||
Returns the window and a `preflight` record for the turn's provenance.
|
||||
Not used by the context dry run: loading a model is a side effect, and
|
||||
opening a panel should not cause one.
|
||||
"""
|
||||
window = await probe(endpoint_url, model, declared=declared)
|
||||
preflight = {"attempted": False, "loaded": None, "verified_before": window.verified,
|
||||
"verified_after": window.verified, "detail": ""}
|
||||
if window.verified:
|
||||
preflight["detail"] = "the window was already verified"
|
||||
return window, preflight
|
||||
if not (endpoint_url and model):
|
||||
preflight["detail"] = "no endpoint or model configured"
|
||||
return window, preflight
|
||||
if not window.reachable:
|
||||
preflight["detail"] = "the server did not answer, so no model was loaded"
|
||||
return window, preflight
|
||||
loaded, detail = await warm(endpoint_url, model, timeout=warm_timeout)
|
||||
preflight.update(attempted=True, loaded=loaded, detail=detail)
|
||||
if loaded:
|
||||
window = await probe(endpoint_url, model, declared=declared, use_cache=False)
|
||||
preflight["verified_after"] = window.verified
|
||||
return window, preflight
|
||||
|
||||
|
||||
def _num_ctx(parameters) -> int | None:
|
||||
"""Reads `num_ctx` out of the plain-text parameter block Ollama returns."""
|
||||
if not isinstance(parameters, str):
|
||||
|
||||
@@ -129,6 +129,34 @@ MAX_BODY_BYTES = 2 * 1024 * 1024
|
||||
MAX_IMPORT_BODY_BYTES = 20 * 1024 * 1024
|
||||
|
||||
|
||||
def import_limit_label(limit: int | None = None) -> str:
|
||||
"""The import ceiling as a reader would say it, e.g. "20 MB".
|
||||
|
||||
Derived from the constant rather than written beside it, so the refusal, the
|
||||
export warning and the documentation cannot drift apart from each other or
|
||||
from what the middleware actually enforces (v1.1 WP-D).
|
||||
"""
|
||||
size = MAX_IMPORT_BODY_BYTES if limit is None else limit
|
||||
megabytes = size / (1024 * 1024)
|
||||
return f"{megabytes:.0f} MB" if abs(megabytes - round(megabytes)) < 0.05 else f"{megabytes:.1f} MB"
|
||||
|
||||
|
||||
def oversized_export_warning(export_bytes: int, limit: int | None = None) -> str:
|
||||
"""What to tell a reader whose export is larger than import will accept.
|
||||
|
||||
v1.1 WP-D. The file is written and is not damaged: what it exceeds is this
|
||||
version's import ceiling, so it cannot be brought back in *here*. Saying that
|
||||
plainly is the whole point — the alternative is a reader who finds out when
|
||||
they try to restore it.
|
||||
"""
|
||||
size = MAX_IMPORT_BODY_BYTES if limit is None else limit
|
||||
return (
|
||||
f"This export is larger than this version's {import_limit_label(size)} import "
|
||||
f"limit ({export_bytes:,} bytes). The file was exported successfully, but this "
|
||||
f"version cannot import it."
|
||||
)
|
||||
|
||||
|
||||
class BodySizeLimitMiddleware:
|
||||
"""Rejects oversized request bodies by their declared `Content-Length`.
|
||||
|
||||
|
||||
+450
-76
@@ -17,9 +17,10 @@ database session. It does three things:
|
||||
then evicts the bank down to its capacity. Evicted memories are marked as
|
||||
forgotten and kept so that the UI can still show them.
|
||||
|
||||
When the app generates a turn, `retrieve_memories` embeds the recent story text
|
||||
and ranks the bank by cosine similarity. The highest-ranked memories become the
|
||||
Memories section of the context.
|
||||
When the app generates a turn, `retrieve_memories` embeds the player's input
|
||||
and the current scene, and ranks the bank by a fixed mix of cosine similarity
|
||||
and rarity-weighted word overlap with the input (v1.1 WP-B.2). The
|
||||
highest-ranked memories become the Memories section of the context.
|
||||
|
||||
Every AI call in this module is best-effort. A failure is logged to the debug
|
||||
page and retried on a later turn, because the cursors advance only after a call
|
||||
@@ -28,6 +29,7 @@ succeeds.
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import math
|
||||
from array import array
|
||||
from collections import OrderedDict
|
||||
|
||||
@@ -36,6 +38,7 @@ from sqlalchemy.orm import Session, defer, object_session
|
||||
|
||||
from . import derived, models, summaries, tree, vectors
|
||||
from .context import (
|
||||
count_tokens,
|
||||
cursors,
|
||||
history,
|
||||
lineage,
|
||||
@@ -43,8 +46,11 @@ from .context import (
|
||||
story_actions,
|
||||
truncate_to_last_tokens,
|
||||
)
|
||||
from .context.builder import _encoding as _token_encoding
|
||||
from .database import SessionLocal
|
||||
from .knowledge import embeddings as knowledge_embeddings
|
||||
from .knowledge import fts
|
||||
from .narrative import model as narrative_model
|
||||
from .providers import OpenAICompatibleProvider, ProviderError
|
||||
from .vectors import cosine # re-exported: the ranking lives here, the maths there
|
||||
|
||||
@@ -55,10 +61,14 @@ MEMORY_START = 12 # first memory once the adventure reaches this many actions
|
||||
SUMMARY_INTERVAL = 15 # actions between Story Summary updates
|
||||
MAX_MEMORIES_PER_RUN = 5 # cap catch-up work (e.g. imported adventures) per turn
|
||||
MAX_EMBED_BATCH = 32
|
||||
RETRIEVAL_WINDOW_TOKENS = 600 # recent story text used as the similarity query
|
||||
RETRIEVAL_WINDOW_ACTIONS = 4 # ...taken from this many of the newest actions
|
||||
SUMMARY_MAX_WORDS = 250
|
||||
MEMORY_EXCERPT_TOKENS = 2000 # of the block, when a block is longer than this
|
||||
MEMORY_EXCERPT_TOKENS = 2000 # the most of a block the summariser is shown
|
||||
|
||||
# v1.1 WP-B.2: what stands between the two parts of a block too long to send
|
||||
# whole. It says a part is missing, so the summariser does not read the end as
|
||||
# following straight on from the opening, and `summarize_block` removes it from
|
||||
# anything the model repeats back.
|
||||
EXCERPT_OMISSION_MARKER = "[… the middle of this stretch of story is left out here …]"
|
||||
|
||||
# How much story has to sit past a block before that block is summarized.
|
||||
#
|
||||
@@ -278,9 +288,10 @@ def set_vector(memory: models.Memory, vector: list[float] | None) -> None:
|
||||
"""
|
||||
memory.embedding_blob = None if vector is None else vectors.pack(vector)
|
||||
memory.embedded = vector is not None
|
||||
cached = _vector_cache.get(memory.adventure_id)
|
||||
if cached is not None:
|
||||
cached.pop(memory.id, None)
|
||||
for cache in (_vector_cache, _terms_cache):
|
||||
cached = cache.get(memory.adventure_id)
|
||||
if cached is not None:
|
||||
cached.pop(memory.id, None)
|
||||
|
||||
|
||||
# ---------- The vector cache ----------
|
||||
@@ -309,6 +320,37 @@ _vector_cache: OrderedDict[int, dict[int, array]] = OrderedDict()
|
||||
VECTOR_CACHE_ADVENTURES = 8 # ~600 KB each at a 100-memory bank
|
||||
|
||||
|
||||
# v1.1 WP-B.2: each memory's lexical terms, held the same way and by the same
|
||||
# two rules as its vector. `set_vector` is also where a memory's text changes
|
||||
# (an edit clears the vector to re-embed it), so dropping the entry there covers
|
||||
# a rewritten text as well as a rewritten vector. Text is read only for memories
|
||||
# not already held, and only on a turn whose input has words to match.
|
||||
_terms_cache: OrderedDict[int, dict[int, frozenset[str]]] = OrderedDict()
|
||||
|
||||
|
||||
def _terms_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, frozenset[str]]:
|
||||
"""The lexical terms for `ids`, reading text only for the ones not already held."""
|
||||
cached = _terms_cache.get(adventure_id)
|
||||
if cached is None:
|
||||
cached = _terms_cache[adventure_id] = {}
|
||||
_terms_cache.move_to_end(adventure_id)
|
||||
while len(_terms_cache) > VECTOR_CACHE_ADVENTURES:
|
||||
_terms_cache.popitem(last=False)
|
||||
|
||||
wanted = set(ids)
|
||||
for gone in set(cached) - wanted:
|
||||
del cached[gone]
|
||||
missing = [memory_id for memory_id in ids if memory_id not in cached]
|
||||
if missing:
|
||||
rows = db.execute(
|
||||
select(models.Memory.id, models.Memory.text)
|
||||
.where(models.Memory.id.in_(missing))
|
||||
).all()
|
||||
for memory_id, text in rows:
|
||||
cached[memory_id] = lexical_terms(text or "")
|
||||
return cached
|
||||
|
||||
|
||||
def forget_cached_vectors(adventure_id: int) -> None:
|
||||
"""Drops an adventure's cached vectors.
|
||||
|
||||
@@ -316,6 +358,7 @@ def forget_cached_vectors(adventure_id: int) -> None:
|
||||
corrects itself, as described in the comment above.
|
||||
"""
|
||||
_vector_cache.pop(adventure_id, None)
|
||||
_terms_cache.pop(adventure_id, None)
|
||||
|
||||
|
||||
def _vectors_for(db: Session, adventure_id: int, ids: list[int]) -> dict[int, array]:
|
||||
@@ -535,6 +578,199 @@ def cast_brief(adventure: models.Adventure, text: str) -> str:
|
||||
|
||||
# ---------- Retrieval (runs inside the turn, before build_context) ----------
|
||||
|
||||
# v1.1 WP-B.2: what the retrieval query is made of, and how a memory is scored
|
||||
# against it (CONTEXT-AND-MEMORY §18, §20).
|
||||
#
|
||||
# WP-B.1 measured the v1.0.0 query, the newest four actions cut to 600 tokens,
|
||||
# against a planted early fact. The player's one-line question arrived after
|
||||
# three turns of narration, so the embedding mostly described the narration: a
|
||||
# direct question about the fact fell from cosine 0.708 on its own to 0.241 in
|
||||
# that query, and a real 100-turn campaign ranked the only memory of the fact
|
||||
# 10th of 19 against a `memory_top_k` of 4.
|
||||
#
|
||||
# The query is now two short texts, embedded in one call:
|
||||
#
|
||||
# input the player's own action this turn, when there is one
|
||||
# context the current scene from the authoritative state (summary, location,
|
||||
# who is present), then the end of the newest narration
|
||||
#
|
||||
# The context is still there because a question often cannot be read without
|
||||
# it ("I ask her where she hid it"), and §18 says retrieval must not rely on raw
|
||||
# input alone. It is bounded so it can resolve a reference but cannot outweigh
|
||||
# the question by sheer length.
|
||||
#
|
||||
# A memory's score is
|
||||
#
|
||||
# semantic_score = INPUT_WEIGHT * cos(input, memory)
|
||||
# + (1 - INPUT_WEIGHT) * cos(context, memory)
|
||||
# lexical_score = rarity-weighted share of the input's words the memory holds
|
||||
# final_score = semantic_score + LEXICAL_WEIGHT * lexical_score
|
||||
#
|
||||
# With no player input (a continue, or a dry run from Insights) the semantic
|
||||
# score is the context cosine alone and the lexical score is 0. Pins are
|
||||
# unchanged: a pinned memory is always used and counts toward `memory_top_k`.
|
||||
INPUT_TYPES = ("do", "say", "story") # player actions that carry words to search for
|
||||
QUERY_INPUT_TOKENS = 200 # of the player's action; a long `story` entry is cut
|
||||
QUERY_SCENE_TOKENS = 60 # of the state's scene line
|
||||
QUERY_NARRATION_TOKENS = 120 # from the end of the newest narration
|
||||
INPUT_WEIGHT = 0.6
|
||||
# Chosen by sweep (0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5) over the deterministic
|
||||
# ranking fixtures, recorded in the WP-B.2 report (§C, §D). The two-part query
|
||||
# alone already ranks the planting-era memory first; 0.15 is the smallest weight
|
||||
# at which the lexical term by itself also lifts it into `memory_top_k` against
|
||||
# the v1.0.0 narration-filled query, and no rare-word negative control put an
|
||||
# unrelated memory above it. At 0.5 an incidental shared word was enough to
|
||||
# select it for an unrelated question, which is the failure a larger weight buys.
|
||||
LEXICAL_WEIGHT = 0.15
|
||||
# `fts.terms` drops these already; the plural fold below is the only stemming.
|
||||
_MIN_FOLD_LENGTH = 5
|
||||
|
||||
|
||||
def _fold(word: str) -> str:
|
||||
"""One term, reduced so "shelves'" and "shelf" do not meet, but "teapots"
|
||||
and "teapot" do. Possessives lose their `'s`, and a trailing `s` goes from a
|
||||
word long enough to be a plural and not ending in `ss`. Deliberately no more
|
||||
than that: a stemmer is a dependency, and a wrong fold merges two words."""
|
||||
word = word.split("'", 1)[0]
|
||||
if len(word) >= _MIN_FOLD_LENGTH and word.endswith("s") and not word.endswith("ss"):
|
||||
word = word[:-1]
|
||||
return word
|
||||
|
||||
|
||||
def lexical_terms(text: str) -> frozenset[str]:
|
||||
"""The words of `text` that lexical matching compares, folded.
|
||||
|
||||
The tokenizer and stop list are imported knowledge's (`knowledge.fts`), so
|
||||
the two retrieval paths agree on what a word is.
|
||||
"""
|
||||
return frozenset(t for t in (_fold(w) for w in fts.terms(text)) if len(t) >= fts.MIN_TERM_LENGTH)
|
||||
|
||||
|
||||
def lexical_scores(input_terms: frozenset[str], terms_of: dict[int, frozenset[str]]) -> dict[int, float]:
|
||||
"""Each candidate's share of the input's rarity, in [0, 1].
|
||||
|
||||
A term's weight is `ln((N + 1) / (df + 1))`: N candidates, df of them holding
|
||||
it. A word every candidate holds weighs exactly 0, so a protagonist's name or
|
||||
a word the whole bank shares moves nothing, and a word no candidate holds
|
||||
weighs the most. The share is taken over **all** the input's terms, so a
|
||||
memory that happens to hold one rare word of a longer question gets that
|
||||
word's part of the question, not the whole of it. The weights live only for
|
||||
this call, over this candidate set: no index, no stored field.
|
||||
"""
|
||||
if not input_terms or not terms_of:
|
||||
return {memory_id: 0.0 for memory_id in terms_of}
|
||||
n = len(terms_of)
|
||||
weight = {
|
||||
term: math.log((n + 1) / (sum(1 for terms in terms_of.values() if term in terms) + 1))
|
||||
for term in input_terms
|
||||
}
|
||||
total = sum(weight.values())
|
||||
if total <= 0:
|
||||
return {memory_id: 0.0 for memory_id in terms_of}
|
||||
return {
|
||||
memory_id: min(1.0, sum(w for term, w in weight.items() if term in terms) / total)
|
||||
for memory_id, terms in terms_of.items()
|
||||
}
|
||||
|
||||
|
||||
def _scene_text(state) -> str:
|
||||
"""The scene as the authoritative state has it: summary, location, who is present.
|
||||
|
||||
Names only, read straight off the document. The full entity list is left
|
||||
out on purpose: a campaign with a large cast would turn every query into a
|
||||
search for everyone.
|
||||
"""
|
||||
if not isinstance(state, dict):
|
||||
return ""
|
||||
scene = state.get("scene")
|
||||
if not isinstance(scene, dict):
|
||||
return ""
|
||||
pieces: list[str] = []
|
||||
summary = scene.get("summary")
|
||||
if isinstance(summary, str) and summary.strip():
|
||||
pieces.append(summary.strip())
|
||||
location = scene.get("location")
|
||||
if isinstance(location, str) and location.strip():
|
||||
pieces.append(narrative_model.entity_name(state, location.strip()))
|
||||
present = scene.get("present")
|
||||
if isinstance(present, list):
|
||||
names = [narrative_model.entity_name(state, key) for key in present[:8]
|
||||
if isinstance(key, str) and key.strip()]
|
||||
if names:
|
||||
pieces.append(", ".join(names))
|
||||
return truncate_to_last_tokens(". ".join(pieces), QUERY_SCENE_TOKENS)
|
||||
|
||||
|
||||
def retrieval_query(adventure: models.Adventure, exclude_action_id: int | None = None) -> dict:
|
||||
"""The two texts a turn's memory retrieval embeds, and the words it matches.
|
||||
|
||||
Returns `{"input", "context", "input_terms"}`. `input` is empty when the
|
||||
newest action is not a player action with text, which is a continue turn or a
|
||||
dry run. `context` is empty only for a story with no scene and no narration.
|
||||
"""
|
||||
recent = history.tail(adventure, 2, exclude_action_id)
|
||||
newest = recent[-1] if recent else None
|
||||
player_input = ""
|
||||
if newest is not None and newest.type in INPUT_TYPES:
|
||||
player_input = truncate_to_last_tokens(newest.text.strip(), QUERY_INPUT_TOKENS)
|
||||
narration = recent[0].text if len(recent) > 1 else ""
|
||||
else:
|
||||
narration = newest.text if newest is not None else ""
|
||||
context = "\n".join(part for part in (
|
||||
_scene_text(adventure.narrative_state),
|
||||
truncate_to_last_tokens(narration.strip(), QUERY_NARRATION_TOKENS),
|
||||
) if part.strip())
|
||||
return {
|
||||
"input": player_input,
|
||||
"context": context,
|
||||
"input_terms": sorted(lexical_terms(player_input)),
|
||||
}
|
||||
|
||||
|
||||
def score_candidates(
|
||||
ids: list[int],
|
||||
held: dict,
|
||||
terms_of: dict[int, frozenset[str]],
|
||||
input_vec,
|
||||
context_vec,
|
||||
input_terms,
|
||||
) -> list[tuple[float, int, float, float]]:
|
||||
"""`(final_score, memory_id, semantic_score, lexical_score)`, best first.
|
||||
|
||||
Ties on the final score are broken by id, so the order never depends on the
|
||||
order the database returned rows in.
|
||||
"""
|
||||
lexical = lexical_scores(frozenset(input_terms), {i: terms_of.get(i, frozenset()) for i in ids})
|
||||
rows = []
|
||||
for memory_id in ids:
|
||||
vector = held[memory_id]
|
||||
if input_vec is not None and context_vec is not None:
|
||||
semantic = (INPUT_WEIGHT * cosine(input_vec, vector)
|
||||
+ (1.0 - INPUT_WEIGHT) * cosine(context_vec, vector))
|
||||
else:
|
||||
semantic = cosine(input_vec if input_vec is not None else context_vec, vector)
|
||||
lex = lexical.get(memory_id, 0.0)
|
||||
rows.append((semantic + LEXICAL_WEIGHT * lex, memory_id, semantic, lex))
|
||||
rows.sort(key=lambda row: (-row[0], row[1]))
|
||||
return rows
|
||||
|
||||
|
||||
def select_memories(scored, pinned_of, held, authority_of, top_k):
|
||||
"""Pins first, then the best-scoring rest, skipping repeats (§22).
|
||||
|
||||
Returns `(used, suppressed)`. `used` is `(final_score, memory_id, pinned)`
|
||||
rows, best first.
|
||||
"""
|
||||
rows = [(final, memory_id, pinned_of[memory_id]) for final, memory_id, _, _ in scored]
|
||||
used = [row for row in rows if row[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
candidates = [row for row in rows if not row[2]]
|
||||
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
|
||||
used += kept
|
||||
used.sort(key=lambda row: (-row[0], row[1]))
|
||||
return used, suppressed
|
||||
|
||||
|
||||
async def retrieve_memories(
|
||||
adventure: models.Adventure,
|
||||
settings: models.Settings,
|
||||
@@ -544,16 +780,17 @@ async def retrieve_memories(
|
||||
"""Returns the memories to inject, or None when the bank is off.
|
||||
|
||||
The result is a dict of the form
|
||||
`{"used": [{id, text, similarity, pinned}], "error": str | None}`. It is
|
||||
None when the memory bank is disabled for this adventure.
|
||||
`{"used": [{id, text, similarity, semantic_score, lexical_score,
|
||||
final_score, pinned, authority, source}], "query": {...}, "error": str | None}`.
|
||||
`similarity` is the semantic score, under the name the inspector has always
|
||||
shown. It is None when the memory bank is disabled for this adventure.
|
||||
|
||||
This only reads. A turn counts the memories it used with `record_use`, just
|
||||
before the commit that saves the turn; see that function for why the count
|
||||
cannot be written here.
|
||||
|
||||
`exclude_action_id` removes the action being retried from the similarity
|
||||
query, so that a discarded attempt cannot influence which memories are
|
||||
returned.
|
||||
`exclude_action_id` removes the action being retried from the query, so that
|
||||
a discarded attempt cannot influence which memories are returned.
|
||||
"""
|
||||
if not adventure.memory_bank_enabled:
|
||||
return None
|
||||
@@ -585,39 +822,33 @@ async def retrieve_memories(
|
||||
if not catalogue:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
recent = history.tail(adventure, RETRIEVAL_WINDOW_ACTIONS, exclude_action_id)
|
||||
query = truncate_to_last_tokens(
|
||||
"\n\n".join(a.text for a in recent), RETRIEVAL_WINDOW_TOKENS
|
||||
)
|
||||
if not query.strip():
|
||||
query = retrieval_query(adventure, exclude_action_id)
|
||||
texts = [t for t in (query["input"], query["context"]) if t.strip()]
|
||||
if not texts:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
try:
|
||||
[query_vec] = await embedding_provider(settings).embed([query])
|
||||
embedded = await embedding_provider(settings).embed(texts)
|
||||
except ProviderError as exc:
|
||||
return {"used": [], "error": str(exc)}
|
||||
vectors_by_text = dict(zip(texts, embedded))
|
||||
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
|
||||
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
|
||||
|
||||
held = _vectors_for(db, adventure.id, [memory_id for memory_id, _, _ in catalogue])
|
||||
ids = [memory_id for memory_id, _, _ in catalogue]
|
||||
held = _vectors_for(db, adventure.id, ids)
|
||||
# Memory text is read only when there are input words to match against, and
|
||||
# then only for memories not already held (see `_terms_for`).
|
||||
terms_of = _terms_for(db, adventure.id, ids) if query["input_terms"] else {}
|
||||
authority_of = {memory_id: authority for memory_id, _, authority in catalogue}
|
||||
scored = sorted(
|
||||
(
|
||||
(cosine(query_vec, held[memory_id]), memory_id, pinned)
|
||||
for memory_id, pinned, _ in catalogue
|
||||
if memory_id in held
|
||||
),
|
||||
key=lambda row: row[0],
|
||||
reverse=True,
|
||||
pinned_of = {memory_id: pinned for memory_id, pinned, _ in catalogue}
|
||||
scored = score_candidates(
|
||||
[memory_id for memory_id in ids if memory_id in held],
|
||||
held, terms_of, input_vec, context_vec, query["input_terms"],
|
||||
)
|
||||
# Pinned memories are always used, and they count toward `top_k`, so the
|
||||
# injected set stays within the budget unless the pinned memories alone
|
||||
# exceed it.
|
||||
top_k = max(1, settings.memory_top_k)
|
||||
used = [row for row in scored if row[2]]
|
||||
remaining = max(0, top_k - len(used))
|
||||
candidates = [row for row in scored if not row[2]]
|
||||
kept, suppressed = _drop_redundant(candidates, held, authority_of, remaining)
|
||||
used += kept
|
||||
used.sort(key=lambda row: row[0], reverse=True)
|
||||
components = {memory_id: (semantic, lex) for _, memory_id, semantic, lex in scored}
|
||||
used, suppressed = select_memories(
|
||||
scored, pinned_of, held, authority_of, max(1, settings.memory_top_k))
|
||||
if not used:
|
||||
return {"used": [], "error": None}
|
||||
|
||||
@@ -638,14 +869,19 @@ async def retrieve_memories(
|
||||
).where(models.Memory.id.in_(used_ids))
|
||||
).all()
|
||||
}
|
||||
texts = {memory_id: row.text for memory_id, row in detail.items()}
|
||||
texts_of = {memory_id: row.text for memory_id, row in detail.items()}
|
||||
|
||||
return {
|
||||
"used": [
|
||||
{
|
||||
"id": memory_id,
|
||||
"text": texts.get(memory_id, ""),
|
||||
"similarity": round(score, 4),
|
||||
"text": texts_of.get(memory_id, ""),
|
||||
"similarity": round(components[memory_id][0], 4),
|
||||
# v1.1 WP-B.2: the parts of the score, so an inspector can see
|
||||
# why this memory beat the ones below it.
|
||||
"semantic_score": round(components[memory_id][0], 4),
|
||||
"lexical_score": round(components[memory_id][1], 4),
|
||||
"final_score": round(final, 4),
|
||||
"pinned": pinned,
|
||||
# M6: what weight this carries, and where it came from.
|
||||
"authority": getattr(detail.get(memory_id), "authority", ACCEPTED_STORY),
|
||||
@@ -656,7 +892,7 @@ async def retrieve_memories(
|
||||
"source_end": getattr(detail.get(memory_id), "source_end", None),
|
||||
},
|
||||
}
|
||||
for score, memory_id, pinned in used
|
||||
for final, memory_id, pinned in used
|
||||
],
|
||||
"considered": len(catalogue),
|
||||
# M6: how many candidates were set aside as repeating one already
|
||||
@@ -665,6 +901,15 @@ async def retrieve_memories(
|
||||
{"id": memory_id, "duplicate_of": kept_id}
|
||||
for memory_id, kept_id in suppressed
|
||||
],
|
||||
# v1.1 WP-B.2: what was searched for. Recorded per turn, like the rest.
|
||||
"query": {
|
||||
"input": query["input"],
|
||||
"context": query["context"],
|
||||
"input_terms": query["input_terms"],
|
||||
"input_weight": INPUT_WEIGHT if input_vec is not None and context_vec is not None
|
||||
else (1.0 if input_vec is not None else 0.0),
|
||||
"lexical_weight": LEXICAL_WEIGHT,
|
||||
},
|
||||
"error": None,
|
||||
}
|
||||
|
||||
@@ -822,6 +1067,61 @@ async def _guarded(db: Session, adventure_id: int, kind: str, coro) -> None:
|
||||
db.commit()
|
||||
|
||||
|
||||
def _excerpt_encoding():
|
||||
return _token_encoding()
|
||||
|
||||
|
||||
def excerpt_split(budget: int = MEMORY_EXCERPT_TOKENS) -> tuple[int, int]:
|
||||
"""`(head_tokens, tail_tokens)` for a block longer than `budget`.
|
||||
|
||||
The marker and the blank lines around it are paid for first; what is left is
|
||||
halved, and an odd token goes to the tail, the most recent part. So the two
|
||||
parts plus the marker come to exactly `budget`.
|
||||
"""
|
||||
room = max(0, budget - count_tokens(f"\n\n{EXCERPT_OMISSION_MARKER}\n\n"))
|
||||
head = room // 2
|
||||
return head, room - head
|
||||
|
||||
|
||||
def memory_excerpt(raw: str, budget: int = MEMORY_EXCERPT_TOKENS) -> str:
|
||||
"""What the summariser is shown of one block.
|
||||
|
||||
v1.1 WP-B.2. A block that fits in `budget` tokens is sent whole, exactly as
|
||||
before. A longer block used to be cut to its last `budget` tokens, and B.1
|
||||
showed that a fact near its start then never reached the summariser at all.
|
||||
It is now sent as its opening and its end, in order, with
|
||||
`EXCERPT_OMISSION_MARKER` between them, still inside `budget`.
|
||||
|
||||
Rejoining two token runs can tokenise a little differently at the seams, so
|
||||
the result is measured, and the head gives up tokens until it fits. A fact in
|
||||
the middle of a very long block is still left out: this bounds the input, it
|
||||
does not summarise everything.
|
||||
"""
|
||||
enc = _excerpt_encoding()
|
||||
tokens = enc.encode(raw)
|
||||
if len(tokens) <= budget:
|
||||
return raw
|
||||
head_n, tail_n = excerpt_split(budget)
|
||||
while True:
|
||||
excerpt = (f"{enc.decode(tokens[:head_n]).rstrip()}\n\n{EXCERPT_OMISSION_MARKER}\n\n"
|
||||
f"{enc.decode(tokens[-tail_n:]).lstrip()}" if tail_n else
|
||||
enc.decode(tokens[:head_n]))
|
||||
over = count_tokens(excerpt) - budget
|
||||
if over <= 0 or head_n == 0:
|
||||
return excerpt
|
||||
head_n = max(0, head_n - over)
|
||||
|
||||
|
||||
def memory_user_prompt(brief: str, excerpt: str) -> str:
|
||||
"""The user message of a memory call: the cast brief, then the excerpt.
|
||||
|
||||
Kept apart from `summarize_block` so an evaluation can send a model exactly
|
||||
what the application sends (v1.1 WP-B.2, `tools/memory_fidelity.py`).
|
||||
"""
|
||||
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
|
||||
return f"{brief}\n\n{prompt}" if brief else prompt
|
||||
|
||||
|
||||
async def summarize_block(
|
||||
adventure: models.Adventure,
|
||||
provider: OpenAICompatibleProvider,
|
||||
@@ -840,15 +1140,16 @@ async def summarize_block(
|
||||
old text in place and moves on.
|
||||
"""
|
||||
raw = "\n\n".join(a.text for a in block)
|
||||
excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)
|
||||
excerpt = memory_excerpt(raw)
|
||||
# Match the cast against the untruncated block. The excerpt is what the
|
||||
# model reads, but a character named in the part that was trimmed is still
|
||||
# model reads, but a character named in the part that was left out is still
|
||||
# one the memory may have to name.
|
||||
brief = cast_brief(adventure, raw)
|
||||
prompt = f"Story excerpt:\n\n{excerpt}\n\nMemory:"
|
||||
return await provider.complete(
|
||||
MEMORY_SYSTEM_PROMPT, f"{brief}\n\n{prompt}" if brief else prompt
|
||||
)
|
||||
text = await provider.complete(MEMORY_SYSTEM_PROMPT, memory_user_prompt(brief, excerpt))
|
||||
# The marker is an instruction to the summariser, never a fact of the story.
|
||||
if text and EXCERPT_OMISSION_MARKER in text:
|
||||
text = " ".join(text.replace(EXCERPT_OMISSION_MARKER, " ").split())
|
||||
return text
|
||||
|
||||
|
||||
async def _create_due_memories(
|
||||
@@ -1028,11 +1329,93 @@ async def _embed_pending(
|
||||
return len(pending)
|
||||
|
||||
|
||||
def eviction_order(rows, limit: int) -> list[int]:
|
||||
"""The ids eviction would take, first to last, at most `limit` of them.
|
||||
|
||||
v1.1 WP-B.2. `rows` are the active memories of one adventure, each with
|
||||
`id`, `pinned`, `source_start`, `source_end`, `last_used_at`, `created_at`
|
||||
and `use_count`. Nothing here reads a vector or the database, so the same
|
||||
function is what the eviction pass runs and what a diagnostic reports.
|
||||
|
||||
WP-B.1 showed what pure least-recently-used order does to a long campaign.
|
||||
Retrieval is steered by the present scene, so a memory of an early stretch
|
||||
nothing recent resembles stops being used. It then becomes the least
|
||||
recently used row, and it goes first, while the bank keeps several memories
|
||||
of the last few scenes that the history window still holds in full. The
|
||||
rule below keeps the bank spread over the whole story instead.
|
||||
|
||||
**Coverage.** Memories with a source range say which stretch of the story
|
||||
they describe. A memory is judged by the hole its removal would leave: the
|
||||
number of depths between the end of the nearest memory before it and the
|
||||
start of the nearest memory after it. The smallest hole goes first, so the
|
||||
bank thins where it is densest. A memory whose start another memory shares
|
||||
(a retried or re-played stretch, or a sibling line) leaves no hole, and is
|
||||
the first kind to go. Pinned memories count as coverage, since they stay.
|
||||
|
||||
**Boundaries.** The earliest and the latest memory by position leave a hole
|
||||
with no memory on one side: removing the first loses the only record of the
|
||||
opening, and removing the last loses the only record of the most recent
|
||||
stretch, which is also what keeps a memory written this turn from being
|
||||
evicted by the pass that wrote it (the frozen bank, below). Boundaries are
|
||||
not coverage candidates.
|
||||
|
||||
**Recency.** Among memories whose removal leaves the same hole, the least
|
||||
recently used goes first (`coalesce(last_used_at, created_at)`), then the
|
||||
less used, then the lower id. Ties are therefore never left to the order the
|
||||
database returned rows in.
|
||||
|
||||
**Fallback.** When no memory is a coverage candidate — memories typed by the
|
||||
player or migrated from before coordinates have no range, and a bank can be
|
||||
all boundaries — the rest are taken least recently used first, exactly as
|
||||
v1.0.0 did. The bank stays bounded either way. Pinned memories are never
|
||||
taken; if every active memory is pinned, capacity yields to the pins.
|
||||
|
||||
Recomputed after each pick, because removing one memory widens the holes
|
||||
of its neighbours.
|
||||
"""
|
||||
remaining = {row.id: row for row in rows if not row.pinned}
|
||||
coverers = {row.id: row for row in rows
|
||||
if row.source_start is not None and row.source_end is not None}
|
||||
|
||||
def recency(row):
|
||||
return (row.last_used_at or row.created_at, row.use_count or 0, row.id)
|
||||
|
||||
order: list[int] = []
|
||||
while remaining and len(order) < limit:
|
||||
spans = sorted(coverers.values(), key=lambda r: (r.source_start, r.source_end, r.id))
|
||||
starts: dict[int, int] = {}
|
||||
for row in spans:
|
||||
starts[row.source_start] = starts.get(row.source_start, 0) + 1
|
||||
best = None
|
||||
furthest_end = None # the largest source_end before index i
|
||||
for i, row in enumerate(spans):
|
||||
if row.id in remaining:
|
||||
if starts[row.source_start] > 1:
|
||||
cost = 0
|
||||
elif i == 0 or i == len(spans) - 1:
|
||||
cost = None # a boundary
|
||||
else:
|
||||
cost = max(0, spans[i + 1].source_start - furthest_end - 1)
|
||||
if cost is not None:
|
||||
key = (cost, *recency(row))
|
||||
if best is None or key < best[0]:
|
||||
best = (key, row.id)
|
||||
furthest_end = row.source_end if furthest_end is None else max(furthest_end, row.source_end)
|
||||
if best is None:
|
||||
victim = min(remaining.values(), key=recency).id
|
||||
else:
|
||||
victim = best[1]
|
||||
order.append(victim)
|
||||
del remaining[victim]
|
||||
coverers.pop(victim, None)
|
||||
return order
|
||||
|
||||
|
||||
def _evict_over_capacity(
|
||||
adventure: models.Adventure, settings: models.Settings, db: Session
|
||||
) -> None:
|
||||
# The database performs both the count and the ranking, and returns neither
|
||||
# the rows nor the vectors. Counting by walking `adventure.memories` fetched
|
||||
# The database performs the count, and the rows read for ordering carry
|
||||
# neither text nor vectors. Counting by walking `adventure.memories` fetched
|
||||
# every vector in the bank on every turn, whether or not the bank was over
|
||||
# capacity.
|
||||
in_this_bank = (models.Memory.adventure_id == adventure.id,
|
||||
@@ -1043,32 +1426,23 @@ def _evict_over_capacity(
|
||||
overflow = active - max(1, settings.memory_bank_capacity)
|
||||
if overflow <= 0:
|
||||
return
|
||||
# Evict the least recently used memory first, and use the use count only to
|
||||
# break ties.
|
||||
# v1.1 WP-B.2: the order is `eviction_order`, coverage first and recency
|
||||
# second. It replaces least recently used alone; see that function.
|
||||
#
|
||||
# Ordering by use count first froze the bank. A memory written on this turn
|
||||
# has never been used, so once every other memory had been retrieved at
|
||||
# least once, the new memory held the lowest count in the bank. The same
|
||||
# post-turn run that wrote it then evicted it, one pass after embedding it.
|
||||
# Use counts only increase, so the bank never recovered. An adventure kept
|
||||
# whatever memories it held when the bank first filled, and every later
|
||||
# memory was summarized, marked as forgotten, and never ranked.
|
||||
#
|
||||
# Ordering by recency avoids that. A new memory carries the newest
|
||||
# timestamp, so it is the last row to be evicted rather than the first, and
|
||||
# it remains until other memories are used. Demoting the use count costs
|
||||
# little, because retrieving a useful memory also makes it recent. The two
|
||||
# orderings differ only for memories that were used once and have not been
|
||||
# retrieved since, which are the rows a full bank should evict.
|
||||
doomed = db.execute(
|
||||
select(models.Memory.id)
|
||||
.where(*in_this_bank, models.Memory.pinned.is_(False))
|
||||
.order_by(
|
||||
func.coalesce(models.Memory.last_used_at, models.Memory.created_at),
|
||||
models.Memory.use_count,
|
||||
)
|
||||
.limit(overflow)
|
||||
).scalars().all()
|
||||
# What the old ordering fixed still holds. Ordering by use count first froze
|
||||
# the bank: a memory written on this turn has never been used, so once every
|
||||
# other memory had been retrieved at least once, the new memory held the
|
||||
# lowest count in the bank, and the same post-turn run that wrote it evicted
|
||||
# it. Counts only increase, so the bank never recovered. Under the coverage
|
||||
# rule the newest memory is the latest boundary, so it is not a coverage
|
||||
# candidate, and in the fallback it carries the newest timestamp.
|
||||
rows = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.last_used_at,
|
||||
models.Memory.created_at, models.Memory.use_count)
|
||||
.where(*in_this_bank)
|
||||
).all()
|
||||
doomed = eviction_order(rows, overflow)
|
||||
if not doomed:
|
||||
return # Every active memory is pinned, so the pins override capacity.
|
||||
db.execute(
|
||||
|
||||
@@ -35,6 +35,8 @@ creating a second, empty Mara.
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
|
||||
# Field types the schema layer enforces. Kept deliberately small: a narrative
|
||||
# state event carries names, labels and plain values, and nothing here needs a
|
||||
# nested structure a model could hide something inside.
|
||||
@@ -184,11 +186,30 @@ def vocabulary_for_prompt() -> str:
|
||||
Generated from `SPECS` rather than written out beside it, so the model can
|
||||
never be told about an event the application does not implement — the drift
|
||||
that would produce proposals rejected for reasons nobody could see.
|
||||
|
||||
v1.1 WP-A2: each event is shown as the object the model must put in the
|
||||
`events` list, with its required fields, not as `name(field, …)`. The call
|
||||
notation was never the wire format, and a 3B narrator copied it into its
|
||||
prose as `> set_possession(silver-key, "alice")`. An object copied into prose
|
||||
is a proposal the extractor already recognises and removes; a call is not.
|
||||
"""
|
||||
lines = []
|
||||
for name, definition in SPECS.items():
|
||||
fields = list(definition["required"]) + [
|
||||
f"{field}?" for field in definition["optional"]
|
||||
]
|
||||
lines.append(f' {name}({", ".join(fields)}) — {definition["summary"]}')
|
||||
shape = {"type": name}
|
||||
for field, kind in definition["required"].items():
|
||||
shape[field] = _PLACEHOLDER[kind]
|
||||
body = json.dumps(shape, ensure_ascii=False, separators=(",", ":"))
|
||||
line = f" {body} — {definition['summary']}"
|
||||
if definition["optional"]:
|
||||
line += f" (optional: {', '.join(definition['optional'])})"
|
||||
lines.append(line)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
#: What a field of each kind looks like in the prompt's vocabulary. Placeholders,
|
||||
#: never example identifiers, so the vocabulary names nothing a story could copy.
|
||||
#: A list field is shown as a list, so the model is told its shape; every other
|
||||
#: field is an ellipsis. Measured: `"<key>"`-style placeholders with spaced
|
||||
#: separators cost 456 tokens against v1.0.0's 258; this form costs about 380,
|
||||
#: and every line is still the object the model must send.
|
||||
_PLACEHOLDER = {KEY: "…", TEXT: "…", VALUE: "…", LABELS: ["…"]}
|
||||
|
||||
@@ -37,17 +37,17 @@ EMIT_RULE = (
|
||||
"appeared, record it.\n"
|
||||
"\n"
|
||||
"Every value is ABSOLUTE — the new state of things, never a change or a "
|
||||
"difference. Use only these events:\n"
|
||||
"difference. Use only these events, in exactly this shape:\n"
|
||||
f"{events.vocabulary_for_prompt()}\n"
|
||||
"\n"
|
||||
"Identifiers are short lower-case slugs (mara, silver-key, old-abbey) and must "
|
||||
"match the ones already in the state you were shown. Introduce a person, place "
|
||||
"or thing with create_entity before referring to it. If the turn established "
|
||||
"nothing, send an empty events list.\n"
|
||||
"Identifiers are short lower-case slugs and must match the ones already in the "
|
||||
"state you were shown; the example's identifiers are placeholders. Introduce a "
|
||||
"person, place or thing with create_entity before referring to it. If the turn "
|
||||
"established nothing, send an empty events list.\n"
|
||||
"Example:\n"
|
||||
'```state\n'
|
||||
'{"events": [{"type": "set_possession", "item": "silver-key", "owner": "aldric"},'
|
||||
' {"type": "set_current_location", "entity": "aldric", "location": "old-abbey"}]}\n'
|
||||
'{"events": [{"type": "set_possession", "item": "item-1", "owner": "character-1"},'
|
||||
' {"type": "set_current_location", "entity": "character-1", "location": "location-1"}]}\n'
|
||||
'```'
|
||||
)
|
||||
|
||||
@@ -58,6 +58,47 @@ EMIT_REMINDER = (
|
||||
"nothing changed.]"
|
||||
)
|
||||
|
||||
# v1.1 WP-A2: the length hint's own words, named once. `builder.length_hint`
|
||||
# builds the hint from these, and the extractor recognises an echo of it by
|
||||
# them, so the two cannot drift apart.
|
||||
LENGTH_HINT_OPENING = "[Hard limit:"
|
||||
LENGTH_HINT_TAIL = "Finish the narration and append the state block well inside the limit."
|
||||
#: The application's wording inside a hint. A 3B narrator reworded the front
|
||||
#: ("your next turn") and the end ("This story ends here."), and kept one or the
|
||||
#: other of these every time.
|
||||
_LENGTH_HINT_PHRASE_RE = re.compile(
|
||||
r"append the state block|turn must not exceed \d+ words", re.IGNORECASE
|
||||
)
|
||||
|
||||
#: v1.1 WP-A2: the rules that remove protocol a narrator copied, named so the
|
||||
#: replay tool and the report can say which removed what.
|
||||
RULE_EVENT_CALL = "event_call_line"
|
||||
RULE_LENGTH_HINT = "echoed_length_hint"
|
||||
RULE_SCENE_LINE = "rendered_scene_line"
|
||||
RULE_EMPTY_FENCE = "empty_dangling_fence"
|
||||
RULE_INSTRUCTION_TAIL = "echoed_instruction_tail"
|
||||
|
||||
#: v1.1 WP-A2 corrective (R5). The sentence `CHAT_CONTINUE_HINT` in
|
||||
#: `providers/openai_compatible.py` carries, which a narrator echoed with the rest
|
||||
#: of the hint reworded around it. Kept as a copy rather than an import, so the
|
||||
#: narrative package does not depend on the provider; a test pins that the
|
||||
#: hint still contains it.
|
||||
CONTINUE_HINT_PHRASE = "Output only story text"
|
||||
|
||||
# R1. A whole line opening with a call to an event this protocol has. The names
|
||||
# come from the vocabulary, so a call-shaped line naming anything else — a
|
||||
# character's `open_door(north)` — is not matched.
|
||||
_EVENT_CALL_LINE_RE = re.compile(
|
||||
r"^[ \t]*(?:>[ \t]*)?(?:"
|
||||
+ "|".join(re.escape(name) for name in events.SPECS)
|
||||
+ r")[ \t]*\(",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# R3. The renderer's scene line carries its location this way.
|
||||
_RENDERED_SCENE_LOCATION_RE = re.compile(r"\(at [^()\n]+\)\s*$")
|
||||
# R4. An opener with nothing after it.
|
||||
_EMPTY_FENCE_LINE_RE = re.compile(r"```(?:json)?[ \t]*", re.IGNORECASE)
|
||||
|
||||
# Three patterns, and the difference between them is the whole of this module's
|
||||
# safety. A story is allowed to contain code, and taking a code block out of
|
||||
# someone's prose is a worse failure than leaving a stray proposal in it.
|
||||
@@ -128,11 +169,42 @@ def _is_echoed_instruction(inner: str) -> bool:
|
||||
# opening words only, because the echo is often cut off before it ends.
|
||||
if low.lstrip().startswith("continue the story directly"):
|
||||
return True
|
||||
# v1.1 WP-A2 corrective (R5): the same hint, reworded at the front. The M11
|
||||
# closeout-era identity re-run stored "[You don't need to continue; … Continue
|
||||
# the story here, directly. Output only story text.]" as the last line of a
|
||||
# reply, and because nothing recognised it, nothing above it was trailing.
|
||||
if CONTINUE_HINT_PHRASE.lower() in low:
|
||||
return True
|
||||
# v1.1 WP-A2 (R2): the length hint, which names "state block" but not
|
||||
# "events list", so it passed every check above.
|
||||
if _is_length_hint(inner):
|
||||
return True
|
||||
# The reminder names both; prose about the protocol rarely names either the
|
||||
# way the instruction does, and effectively never both.
|
||||
return "state block" in low and "events list" in low
|
||||
|
||||
|
||||
def _opens_like_length_hint(inner: str) -> bool:
|
||||
"""R5. The bracket opens with the length hint's own `Hard limit:`, whatever follows.
|
||||
|
||||
Never enough on its own: an in-world "[Hard limit: forty days]" opens the same
|
||||
way. `_clean` takes it only directly above an echoed instruction it has already
|
||||
removed from the end of the same reply.
|
||||
"""
|
||||
return inner.lstrip().lower().startswith(LENGTH_HINT_OPENING[1:].lower())
|
||||
|
||||
|
||||
def _is_length_hint(inner: str) -> bool:
|
||||
"""Whether a bracket's contents are `builder.length_hint`, however reworded.
|
||||
|
||||
It must open the way the hint opens *and* carry the hint's own wording. An
|
||||
in-world "Hard limit: forty days" has the opening and none of the wording.
|
||||
"""
|
||||
opening = LENGTH_HINT_OPENING[1:].lower()
|
||||
return (inner.lstrip().lower().startswith(opening)
|
||||
and bool(_LENGTH_HINT_PHRASE_RE.search(inner)))
|
||||
|
||||
|
||||
# A heading the model writes above a block it did not fence: `State`, sometimes
|
||||
# as `State:`, `**State**` or `### State`. It is removed only in two places:
|
||||
# directly above a proposal that is removed, and as the last line of the reply.
|
||||
@@ -145,7 +217,7 @@ _LINE_OBJECT_RE = re.compile(r"^[ \t]*(?:>[ \t]*)?\{", re.MULTILINE)
|
||||
_QUOTE_PREFIX_RE = re.compile(r"^[ \t]*>[ \t]?")
|
||||
|
||||
|
||||
def _clean(prose: str) -> str:
|
||||
def _clean(prose: str, *, after_block: bool = False) -> str:
|
||||
"""Removes protocol the block extraction could not, and nothing else.
|
||||
|
||||
Found by the M5 realistic-context run (§12), which is the failure class
|
||||
@@ -160,9 +232,23 @@ def _clean(prose: str) -> str:
|
||||
story after it. Stored text is replayed as history, so every leak also
|
||||
showed the next prompt a second, older account of the state, which is what
|
||||
M5 review Finding 4 removed from replayed history.
|
||||
|
||||
v1.1 WP-A2 added four shapes, from the M11 closeout's identity run and the
|
||||
v1 corpus, each anchored to something the application owns rather than to
|
||||
what prose looks like: a line opening with a vocabulary call (R1), the
|
||||
length hint echoed at the end (R2), the renderer's scene line left last
|
||||
(R3), and an empty fence opener left last (R4). `after_block` says a
|
||||
proposal block was already taken out of this reply, which is what lets R3
|
||||
remove a bare scene line that sat above it.
|
||||
"""
|
||||
cleaned, _found = _inline_proposals(prose)
|
||||
cleaned, calls_removed = _strip_event_call_lines(prose)
|
||||
cleaned, _found = _inline_proposals(cleaned)
|
||||
cleaned = _strip_echoed_state(cleaned)
|
||||
protocol_cut = after_block or calls_removed
|
||||
# R5: set once an echoed instruction bracket has come off the end. Only then
|
||||
# may a bracket that merely opens the way the length hint opens be taken as
|
||||
# part of the same echoed tail.
|
||||
instruction_cut = False
|
||||
# The end of the reply is cut until nothing more comes off, because one kind
|
||||
# of leftover can hide another. In a real reply, a `State` heading sat above
|
||||
# a block the model never finished, and a parroted reminder sat above an
|
||||
@@ -171,7 +257,12 @@ def _clean(prose: str) -> str:
|
||||
before = cleaned
|
||||
for pattern in (_TRAILING_BRACKET_RE, _UNCLOSED_BRACKET_RE):
|
||||
bracket = pattern.search(cleaned)
|
||||
if bracket is not None and _is_echoed_instruction(bracket.group(1)):
|
||||
if bracket is None:
|
||||
continue
|
||||
if _is_echoed_instruction(bracket.group(1)):
|
||||
cleaned = cleaned[: bracket.start()]
|
||||
instruction_cut = True
|
||||
elif instruction_cut and _opens_like_length_hint(bracket.group(1)):
|
||||
cleaned = cleaned[: bracket.start()]
|
||||
cleaned = _DANGLING_STATE_RE.sub("", cleaned)
|
||||
dangling = _DANGLING_JSON_RE.search(cleaned)
|
||||
@@ -182,10 +273,93 @@ def _clean(prose: str) -> str:
|
||||
cleaned = _strip_trailing_state_heading(cleaned).rstrip()
|
||||
# A bare quote marker, the start of a quoted block that never came.
|
||||
cleaned = re.sub(r"\n[ \t]*>[ \t]*\Z", "", cleaned)
|
||||
cleaned = _strip_empty_dangling_fence(cleaned)
|
||||
if cleaned.rstrip() != before.rstrip():
|
||||
protocol_cut = True
|
||||
cleaned = _strip_trailing_scene_line(cleaned, protocol_cut)
|
||||
if cleaned == before:
|
||||
return cleaned.strip()
|
||||
|
||||
|
||||
def _strip_event_call_lines(text: str) -> tuple[str, bool]:
|
||||
"""R1. Removes whole lines that open with a call to a vocabulary event.
|
||||
|
||||
A line inside a fenced code block is the story's own code and is never
|
||||
examined. Returns the text and whether anything was removed.
|
||||
"""
|
||||
kept: list[str] = []
|
||||
in_fence = False
|
||||
removed = False
|
||||
for line in text.split("\n"):
|
||||
if line.lstrip().startswith("```"):
|
||||
in_fence = not in_fence
|
||||
kept.append(line)
|
||||
continue
|
||||
if not in_fence and _EVENT_CALL_LINE_RE.match(line):
|
||||
removed = True
|
||||
continue
|
||||
kept.append(line)
|
||||
if not removed:
|
||||
return text, False
|
||||
return re.sub(r"\n{3,}", "\n\n", "\n".join(kept)), True
|
||||
|
||||
|
||||
def _strip_empty_dangling_fence(text: str) -> str:
|
||||
"""R4. A ```` ```json ```` or ```` ``` ```` opener as the last line, with nothing after it.
|
||||
|
||||
Only an *opener*: the fence lines are counted, and an even count means the
|
||||
last one closes a story's own code block, which stays.
|
||||
"""
|
||||
lines = text.rstrip().split("\n")
|
||||
if len(lines) < 2 or not _EMPTY_FENCE_LINE_RE.fullmatch(lines[-1].strip()):
|
||||
return text
|
||||
fences = sum(1 for line in lines if line.lstrip().startswith("```"))
|
||||
if fences % 2 == 0:
|
||||
return text
|
||||
return "\n".join(lines[:-1]).rstrip()
|
||||
|
||||
|
||||
def _strip_trailing_scene_line(text: str, protocol_cut: bool) -> str:
|
||||
"""R3. The renderer's scene line, left as the last line of the reply.
|
||||
|
||||
Taken when it carries the renderer's own `(at <location>)`, or when protocol
|
||||
was already cut from this reply, which makes a bare scene line part of the
|
||||
same pasted tail. A final screenplay-style "Scene: …" line in a reply with
|
||||
no protocol in it stays, and so does any scene line with story after it.
|
||||
"""
|
||||
lines = text.rstrip().split("\n")
|
||||
if len(lines) < 2:
|
||||
return text
|
||||
last = lines[-1].strip()
|
||||
if not last.startswith(render.HEADING_SCENE + " "):
|
||||
return text
|
||||
if not (_RENDERED_SCENE_LOCATION_RE.search(last) or protocol_cut):
|
||||
return text
|
||||
return "\n".join(lines[:-1]).rstrip()
|
||||
|
||||
|
||||
def explain_removed_line(line: str) -> str | None:
|
||||
"""Which v1.1 rule removes a line of this shape, for the replay report.
|
||||
|
||||
None means no v1.1 rule explains it, which the replay treats as a failure.
|
||||
"""
|
||||
stripped = line.strip()
|
||||
if _EVENT_CALL_LINE_RE.match(line):
|
||||
return RULE_EVENT_CALL
|
||||
if stripped.startswith("["):
|
||||
inner = stripped[1:]
|
||||
inner = inner[:-1] if inner.endswith("]") else inner
|
||||
if _is_length_hint(inner):
|
||||
return RULE_LENGTH_HINT
|
||||
if _is_echoed_instruction(inner) or _opens_like_length_hint(inner):
|
||||
return RULE_INSTRUCTION_TAIL
|
||||
if stripped.startswith(render.HEADING_SCENE + " "):
|
||||
return RULE_SCENE_LINE
|
||||
if _EMPTY_FENCE_LINE_RE.fullmatch(stripped):
|
||||
return RULE_EMPTY_FENCE
|
||||
return None
|
||||
|
||||
|
||||
def _is_state_heading(line: str) -> bool:
|
||||
return bool(_STATE_HEADING_RE.match(line))
|
||||
|
||||
@@ -451,7 +625,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
if matches:
|
||||
match = matches[-1]
|
||||
raw = match.group(1).strip()
|
||||
prose = _clean(text[: match.start()] + text[match.end():])
|
||||
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
||||
return prose, _tolerant_load(raw), raw
|
||||
|
||||
# A `json` or unlabelled fence is ours only when its contents are this
|
||||
@@ -465,7 +639,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
raw = match.group(1).strip()
|
||||
parsed = _tolerant_load(raw)
|
||||
if _looks_like_proposal(parsed) or _reads_as_protocol(raw):
|
||||
prose = _clean(text[: match.start()] + text[match.end():])
|
||||
prose = _clean(text[: match.start()] + text[match.end():], after_block=True)
|
||||
return prose, parsed, raw
|
||||
|
||||
match = _TRAILING_RE.search(text)
|
||||
@@ -473,7 +647,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
raw = match.group(1)
|
||||
parsed = _tolerant_load(raw)
|
||||
if _looks_like_proposal(parsed):
|
||||
return _clean(text[: match.start()]), parsed, raw
|
||||
return _clean(text[: match.start()], after_block=True), parsed, raw
|
||||
|
||||
# An unfenced proposal on its own lines but not at the end: quoted, or
|
||||
# followed by more story. The last one is the turn's proposal, as with
|
||||
@@ -481,7 +655,7 @@ def split(text: str) -> tuple[str, dict | None, str]:
|
||||
without, found = _inline_proposals(text)
|
||||
if found:
|
||||
parsed, raw = found[-1]
|
||||
return _clean(without), parsed, raw
|
||||
return _clean(without, after_block=True), parsed, raw
|
||||
|
||||
# No block at all — but the reply may still carry protocol the model wrote
|
||||
# as prose, or a fence it never closed.
|
||||
|
||||
@@ -26,6 +26,10 @@ EMBED_READ_TIMEOUT = 60.0
|
||||
|
||||
|
||||
|
||||
#: v1.1 WP-A1: ask a stream to report its token usage. Without it Ollama sends
|
||||
#: none, and a prompt the server cut cannot be told from one it read whole.
|
||||
STREAM_OPTIONS = {"include_usage": True}
|
||||
|
||||
# Completion endpoints have no roles, so a chat has to be flattened into one
|
||||
# labeled transcript that ends on "Assistant:" for the model to continue.
|
||||
_ROLE_LABELS = {"system": "System", "user": "User", "assistant": "Assistant"}
|
||||
@@ -84,10 +88,14 @@ class OpenAICompatibleProvider(Provider):
|
||||
def _record_usage(self, payload: dict) -> None:
|
||||
"""Records the endpoint's own token accounting, if it reported any.
|
||||
|
||||
OpenRouter now always reports usage, and `usage: {include: true}` and
|
||||
`stream_options` are deprecated and do nothing. In a stream the usage
|
||||
arrives on a final chunk that carries no choices, which is why this is
|
||||
read separately from the text extraction.
|
||||
In a stream the usage arrives on a final chunk that carries no choices,
|
||||
which is why this is read separately from the text extraction.
|
||||
|
||||
v1.1 WP-A1: Ollama sends that chunk only when asked. Measured on Ollama
|
||||
0.33: a stream with no `stream_options` carried no usage at all, and not
|
||||
one of the 514 AI turns in the v1 evidence had a count stored. Every
|
||||
streaming body therefore sets `stream_options.include_usage`
|
||||
(`STREAM_OPTIONS`), and the turn compares the count with what it sent.
|
||||
"""
|
||||
usage = payload.get("usage")
|
||||
if isinstance(usage, dict) and usage:
|
||||
@@ -102,6 +110,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
else:
|
||||
url = f"{self.base_url}/chat/completions"
|
||||
@@ -114,6 +123,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
return url, body
|
||||
|
||||
@@ -183,6 +193,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
else:
|
||||
url = f"{self.base_url}/chat/completions"
|
||||
@@ -192,6 +203,7 @@ class OpenAICompatibleProvider(Provider):
|
||||
"temperature": temperature,
|
||||
"max_tokens": max_tokens,
|
||||
"stream": True,
|
||||
"stream_options": dict(STREAM_OPTIONS),
|
||||
}
|
||||
async for event in self._stream(url, body):
|
||||
yield event
|
||||
|
||||
@@ -28,7 +28,9 @@ repair. Refusing a whole campaign because a search index would not build would
|
||||
trade the valuable thing for the cheap one.
|
||||
"""
|
||||
|
||||
from fastapi import Body, Depends, Request
|
||||
import json
|
||||
|
||||
from fastapi import Body, Depends, Request, Response
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from ... import bundle, head, limits, models, schemas
|
||||
@@ -46,8 +48,38 @@ def export_adventure(
|
||||
|
||||
`app/bundle.py` owns the format, in all three of its versions. A backup
|
||||
outlives the schema, so no call site decides anything about its shape.
|
||||
|
||||
**v1.1 WP-D: the export also says whether this version could import it back.**
|
||||
A campaign large enough to pass `limits.MAX_IMPORT_BODY_BYTES` still exports —
|
||||
the file is complete and not damaged, and refusing to write it would destroy
|
||||
the only copy the reader was trying to make. What it cannot do is come back
|
||||
in here, and the reader is told that at the moment they take it rather than
|
||||
at the moment they need it.
|
||||
|
||||
It travels in headers, not in the body. The body is the bundle, the browser
|
||||
saves exactly those bytes as the file, and a warning inside it would become
|
||||
part of a portable story file and of every checksum taken over one.
|
||||
|
||||
The size measured is the compact serialisation, because that is both what
|
||||
this response sends and what the browser POSTs back on import, which is what
|
||||
`BodySizeLimitMiddleware` weighs. The pretty-printed file the reader
|
||||
downloads is larger, and is not what import reads.
|
||||
"""
|
||||
return bundle.export(db, adv)
|
||||
payload = bundle.export(db, adv)
|
||||
# Serialised exactly as Starlette's JSONResponse would, so the bytes counted
|
||||
# are the bytes sent.
|
||||
body = json.dumps(payload, ensure_ascii=False, allow_nan=False,
|
||||
separators=(",", ":")).encode("utf-8")
|
||||
limit = limits.MAX_IMPORT_BODY_BYTES
|
||||
importable = len(body) <= limit
|
||||
headers = {
|
||||
"X-Export-Bytes": str(len(body)),
|
||||
"X-Import-Limit-Bytes": str(limit),
|
||||
"X-Importable-By-This-Version": "true" if importable else "false",
|
||||
}
|
||||
if not importable:
|
||||
headers["X-Export-Warning"] = limits.oversized_export_warning(len(body), limit)
|
||||
return Response(content=body, media_type="application/json", headers=headers)
|
||||
|
||||
|
||||
@router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201)
|
||||
|
||||
@@ -6,6 +6,7 @@ lock guards one set only while one module owns it. And a test that replaces
|
||||
`OpenAICompatibleProvider` or `generate_turn` patches this module, which every
|
||||
caller reads through.
|
||||
"""
|
||||
import logging
|
||||
import threading
|
||||
|
||||
from fastapi import Depends, HTTPException, Request
|
||||
@@ -28,6 +29,8 @@ from .deps import CurrentUser, current_adventure, router
|
||||
from .nodes import _move_to_after, next_depth
|
||||
from .paging import annotate_takes
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def world_delta_of(snapshot: dict | None) -> dict | None:
|
||||
"""Returns the bulk-read slice of a context snapshot, for `Action.world_delta`.
|
||||
@@ -212,8 +215,18 @@ async def _generate_turn(
|
||||
# network calls — and cached per endpoint and model, so it costs one short
|
||||
# request per session rather than one per turn. An unverified window does
|
||||
# not block the turn; it is recorded as unverified in the snapshot below.
|
||||
window = await contextwindow.probe(settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override)
|
||||
#
|
||||
# v1.1 WP-A1 corrective: a model that is not resident cannot report its window,
|
||||
# and a turn built to the configured budget against it was silently cut in the
|
||||
# A1 evidence (13,875 tokens sent, 2,050 read). So an unverified window gets
|
||||
# one bounded attempt to load the model, and one more probe, before the
|
||||
# prompt is assembled. No story text is generated by it and nothing is
|
||||
# written. A window still unverified afterwards changes nothing below.
|
||||
window, preflight = await contextwindow.ensure_window(
|
||||
settings.endpoint_url, settings.model,
|
||||
declared=settings.context_window_override,
|
||||
warm_timeout=float(settings.model_timeout_seconds or 300),
|
||||
)
|
||||
try:
|
||||
system_text, story_text, snapshot = build_context(
|
||||
adventure,
|
||||
@@ -232,6 +245,9 @@ async def _generate_turn(
|
||||
yield turn_error(str(exc))
|
||||
return
|
||||
|
||||
if isinstance(snapshot.get("window"), dict):
|
||||
snapshot["window"]["preflight"] = preflight
|
||||
|
||||
parts = PromptParts(system=system_text, story=story_text)
|
||||
|
||||
provider = OpenAICompatibleProvider(
|
||||
@@ -318,6 +334,24 @@ async def _generate_turn(
|
||||
# prompt came from cache rather than being billed in full. This is recorded
|
||||
# per attempt, next to the prompt it priced.
|
||||
snapshot["usage"] = provider.last_usage
|
||||
# v1.1 WP-A1: what the server says it read, against what was sent. Recorded
|
||||
# and shown, never acted on: the narration has already streamed to the
|
||||
# reader, and discarding an accepted turn over an accounting discrepancy
|
||||
# would lose story to hide a problem. A server that cut the prompt answers
|
||||
# 200 either way, so this record is the only place the cut is visible.
|
||||
tokens = snapshot.get("tokens") or {}
|
||||
accounting = contextwindow.classify_usage(
|
||||
provider.last_usage,
|
||||
estimate=tokens.get("estimate") or tokens.get("total") or 0,
|
||||
budget=tokens.get("budget") or settings.context_token_budget,
|
||||
max_output_tokens=settings.max_output_tokens,
|
||||
window_verified=bool((snapshot.get("window") or {}).get("verified")),
|
||||
)
|
||||
snapshot["accounting"] = accounting
|
||||
if accounting["status"] in (contextwindow.EXCEEDED,
|
||||
contextwindow.TRUNCATION_SUSPECTED):
|
||||
log.warning("turn accounting for adventure %s: %s — %s",
|
||||
adventure.id, accounting["status"], accounting["detail"])
|
||||
|
||||
reasoning = "".join(reasoning_chunks).strip() or None
|
||||
ai_action = models.Action(
|
||||
@@ -381,7 +415,8 @@ async def _generate_turn(
|
||||
db.commit()
|
||||
db.refresh(ai_action)
|
||||
yield _SAVED
|
||||
yield sse({"type": "done", "action": action_json(ai_action, db)})
|
||||
yield sse({"type": "done", "action": action_json(ai_action, db),
|
||||
"accounting": accounting})
|
||||
# Phase 6: schedule summarization and embedding without waiting for them.
|
||||
# The task opens its own database session.
|
||||
memorybank.schedule_post_turn(adventure)
|
||||
|
||||
@@ -246,11 +246,28 @@ def test_the_story_prompt_keeps_its_prefix_across_a_new_turn(saturated):
|
||||
assert report_before["history"]["floor_depth"] is not None, (
|
||||
"this fixture is meant to be over budget; trimming never engaged")
|
||||
|
||||
_play_one_more(db, adventure, 60)
|
||||
_, after, report_after = _builder.build_context(adventure, settings)
|
||||
# v1.1 WP-A1: the fixture used to be positioned so that the very next turn
|
||||
# held the floor. The safety reserve takes 256 tokens of this 2,048 budget,
|
||||
# the block is now the minimum of two, and the next turn is a step. So walk
|
||||
# forward until a turn holds, requiring every move on the way to be exactly
|
||||
# one block: a window that slides by one action every turn fails either way.
|
||||
held = None
|
||||
depth = 60
|
||||
for _ in range(4):
|
||||
_play_one_more(db, adventure, depth)
|
||||
depth += 1
|
||||
_, after, report_after = _builder.build_context(adventure, settings)
|
||||
floor_before = report_before["history"]["floor_depth"]
|
||||
floor_after = report_after["history"]["floor_depth"]
|
||||
block = report_after["history"]["trim_block"]
|
||||
assert floor_after - floor_before in (0, block), (floor_before, floor_after, block)
|
||||
if floor_after == floor_before:
|
||||
held = (before, after)
|
||||
break
|
||||
before, report_before = after, report_after
|
||||
|
||||
assert report_after["history"]["floor_depth"] == report_before["history"]["floor_depth"]
|
||||
assert _shared_prefix(before, after) > 0.85
|
||||
assert held is not None, "the floor never held across a turn"
|
||||
assert _shared_prefix(*held) > 0.85
|
||||
|
||||
|
||||
def test_without_a_stable_floor_the_prefix_collapses(saturated):
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
"""v1.1 WP-B.1: the long run's `recovered_through_memory_independent` verdict.
|
||||
|
||||
The new verdict must never be reported when anything other than memory could
|
||||
have carried the fact. Each precondition is named when it fails. The existing M04
|
||||
verdicts keep their meaning exactly.
|
||||
|
||||
python -m pytest tests/test_v11_b1_long_run_verdict.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
|
||||
from tools import m11_long_run as lr
|
||||
|
||||
GOOD = {
|
||||
"independent_planted_depth": 3,
|
||||
"planted_turn_outside_history": True,
|
||||
"absent_from_state": True,
|
||||
"absent_from_summary": True,
|
||||
"absent_from_knowledge": True,
|
||||
"absent_from_later_narration": True,
|
||||
"memory_covering_planting_carries_fact": True,
|
||||
"memory_forgotten": False,
|
||||
"memory_injected": True,
|
||||
}
|
||||
|
||||
|
||||
def test_every_precondition_and_an_injected_memory_is_the_new_verdict():
|
||||
assert lr._independent_memory_verdict(GOOD) == "recovered_through_memory_independent"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
|
||||
def test_a_failed_precondition_is_named_and_never_a_recovery(name):
|
||||
assert lr._independent_memory_verdict({**GOOD, name: False}) == f"precondition_failed:{name}"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", lr.INDEPENDENT_PRECONDITIONS)
|
||||
def test_an_unmeasured_precondition_is_unknown_not_a_pass(name):
|
||||
assert lr._independent_memory_verdict({**GOOD, name: None}) == f"precondition_unknown:{name}"
|
||||
|
||||
|
||||
def test_no_planted_depth_is_unknown():
|
||||
assert lr._independent_memory_verdict({**GOOD, "independent_planted_depth": None}) == \
|
||||
"precondition_unknown:planted_depth"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("change, verdict", [
|
||||
({"memory_covering_planting_carries_fact": False}, "not_recovered:not_created"),
|
||||
({"memory_forgotten": True}, "not_recovered:evicted"),
|
||||
({"memory_injected": False}, "not_recovered:not_injected"),
|
||||
])
|
||||
def test_the_failing_memory_stage_is_named(change, verdict):
|
||||
assert lr._independent_memory_verdict({**GOOD, **change}) == verdict
|
||||
|
||||
|
||||
def test_preconditions_are_judged_before_memory():
|
||||
"""A carried fact disqualifies the run even when memory also failed."""
|
||||
both = {**GOOD, "absent_from_state": False, "memory_covering_planting_carries_fact": False}
|
||||
assert lr._independent_memory_verdict(both) == "precondition_failed:absent_from_state"
|
||||
|
||||
|
||||
def test_the_fact_is_matched_as_whole_words():
|
||||
assert lr._mentions_fact("She hid the amber Sundial.")
|
||||
assert lr._mentions_fact("a cracked TEAPOT on the shelf")
|
||||
assert not lr._mentions_fact("teapots") # a different word, not the fact's
|
||||
assert not lr._mentions_fact("the sun dialled down")
|
||||
|
||||
|
||||
def test_the_m04_verdicts_are_unchanged():
|
||||
base = {"planted_turn_in_history_window": False, "in_memories_section": False,
|
||||
"in_summary_section": False, "in_state_section": False}
|
||||
assert lr._m04_verdict(base) == "not_recovered"
|
||||
assert lr._m04_verdict({**base, "in_state_section": True}) == "recovered_through_state_only"
|
||||
assert lr._m04_verdict({**base, "in_memories_section": True}) == \
|
||||
"recovered_through_memory_or_summary"
|
||||
assert lr._m04_verdict({**base, "planted_turn_in_history_window": True}) == \
|
||||
"precondition_not_met"
|
||||
|
||||
|
||||
def test_the_independent_fact_is_not_in_any_imported_knowledge_file():
|
||||
for text in (lr.CANON_MD, lr.REFERENCE_MD, lr.INSPIRATION_MD, *lr.BEATS):
|
||||
assert not lr._mentions_fact(text)
|
||||
|
||||
|
||||
def test_the_planting_text_and_recall_carry_the_fact():
|
||||
assert lr._mentions_fact(lr.INDEPENDENT_FACT_TEXT)
|
||||
assert lr._mentions_fact(lr.INDEPENDENT_RECALL_TEXT)
|
||||
@@ -0,0 +1,434 @@
|
||||
"""v1.1 WP-B.1: the memory-retention diagnostic, deterministically.
|
||||
|
||||
B.1 changes no memory behaviour. These tests prove two things about the
|
||||
diagnostic in `tools/memory_diagnostic.py`:
|
||||
|
||||
1. **It measures what it claims.**
|
||||
- The fixture keeps the planted fact out of every layer except memory.
|
||||
- Each stage (created, retained, ranked, injected) is reported from the rows
|
||||
and the recall turn's own stored context.
|
||||
- Its ranking agrees with the selection production stored.
|
||||
2. **What it finds on this tree.** The scenarios run with a best-case summariser,
|
||||
one that keeps a fact if and only if the fact reached it. Any failure is
|
||||
therefore the application's mechanism, not a model's writing.
|
||||
- WP-B.1 marked the criteria v1.0.0 did not meet `xfail(strict=True)`. WP-B.2
|
||||
fixed ranking, eviction and the creation excerpt, and those tests are now
|
||||
ordinary passes; the v1.0.0 results are recorded in the WP-B.2 report.
|
||||
- The same file is run unchanged against v1.0.0 for the baseline.
|
||||
|
||||
python -m pytest tests/test_v11_b1_memory_diagnostic.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
from tools import memory_diagnostic as md
|
||||
|
||||
_results: dict = {}
|
||||
|
||||
|
||||
def scenario(name: str) -> dict:
|
||||
"""Runs a named scenario once per session and keeps the result."""
|
||||
if name not in _results:
|
||||
_results[name] = md.run_scenario(md.SCENARIOS[name])
|
||||
return _results[name]
|
||||
|
||||
|
||||
# ------------------------------------------------------- fixture preconditions
|
||||
|
||||
def test_the_fact_is_planted_early_and_recalled_past_depth_one_hundred():
|
||||
result = scenario("independent_default")
|
||||
assert result["plant_depth"] is not None and result["plant_depth"] <= 3
|
||||
assert result["recall_depth"] >= 100
|
||||
|
||||
|
||||
@pytest.mark.parametrize("check", ["state_document", "state_snapshots", "later_narration",
|
||||
"summary", "knowledge", "recent_history", "state_section"])
|
||||
def test_no_layer_but_memory_carries_the_fact(check):
|
||||
"""A test where another layer carries F is not evidence about memory."""
|
||||
isolation = scenario("independent_default")["isolation"]
|
||||
assert isolation["checks"][check]["ok"], isolation["checks"][check]
|
||||
assert isolation["ok"]
|
||||
|
||||
|
||||
def test_the_isolation_check_fails_when_another_layer_carries_the_fact():
|
||||
"""The negative control for the precondition itself: a state fact naming F."""
|
||||
fact = md.FACT_F
|
||||
with SessionLocal() as db:
|
||||
Base.metadata.create_all(bind=engine)
|
||||
try:
|
||||
user = models.User(is_guest=False, email="b1-iso@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
adventure = models.Adventure(user_id=user.id, title="iso")
|
||||
adventure.narrative_state = {"facts": [{"id": "x", "predicate": "hidden",
|
||||
"value": "the amber sundial is in the teapot"}]}
|
||||
db.add(adventure)
|
||||
db.commit()
|
||||
result = md.isolation(db, adventure, fact, 1)
|
||||
assert result["ok"] is False
|
||||
assert result["checks"]["state_document"]["ok"] is False
|
||||
finally:
|
||||
db.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- stages
|
||||
|
||||
def test_creation_is_reported_with_the_covering_memory_and_what_the_summariser_saw():
|
||||
created = scenario("independent_default")["diagnosis"]["created"]
|
||||
assert created["yes"] is True
|
||||
assert created["source_start"] <= scenario("independent_default")["plant_depth"] <= created["source_end"]
|
||||
assert md.FACT_F.carried_by(created["memory_text"])
|
||||
covering = [c for c in created["covering_memories"] if c["memory_id"] == created["memory_id"]]
|
||||
assert covering and covering[0]["fact_in_block"] and covering[0]["fact_in_summariser_excerpt"]
|
||||
|
||||
|
||||
def test_retention_is_reported_with_the_bank_and_its_eviction_order():
|
||||
retained = scenario("independent_default")["diagnosis"]["retained"]
|
||||
assert retained["yes"] is True and retained["forgotten"] is False
|
||||
assert retained["on_active_lineage"] is True
|
||||
assert retained["active_memories"] <= retained["memory_bank_capacity"]
|
||||
assert retained["eviction_position"] is not None
|
||||
|
||||
|
||||
def test_ranking_is_production_ranking_and_agrees_with_the_stored_selection():
|
||||
ranked = scenario("independent_default")["diagnosis"]["ranked"]
|
||||
assert ranked["replica_matches_stored_selection"] is True
|
||||
assert ranked["top_k_cutoff"] == 5
|
||||
assert ranked["yes"] is True and ranked["selected"] is True
|
||||
assert 1 <= ranked["rank"] <= ranked["top_k_cutoff"]
|
||||
# v1.1 WP-B.2: every part of the score is reported, and they add up.
|
||||
assert 0.0 <= ranked["lexical_score"] <= 1.0
|
||||
assert ranked["final_score"] == pytest.approx(
|
||||
ranked["semantic_score"] + memorybank.LEXICAL_WEIGHT * ranked["lexical_score"], abs=2e-4)
|
||||
assert ranked["query"]["input"].endswith(md.SCENARIOS["independent_default"].recall_text)
|
||||
|
||||
|
||||
def test_injection_is_read_from_the_recall_turns_own_context():
|
||||
diagnosis = scenario("independent_default")["diagnosis"]
|
||||
assert diagnosis["injected"]["yes"] is True
|
||||
assert diagnosis["injected"]["context_component"] == md.MEMORIES_LABEL
|
||||
assert diagnosis["injected"]["token_count"] > 0
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_ranking_variants_direct_paraphrase_and_unrelated():
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["direct"]["rank"] == 1 and variants["direct"]["selected"]
|
||||
assert variants["paraphrase"]["rank"] == 1 and variants["paraphrase"]["selected"]
|
||||
assert (variants["direct"]["final_score"] > variants["paraphrase"]["final_score"]
|
||||
> variants["unrelated"]["final_score"])
|
||||
|
||||
|
||||
def test_retrieval_still_fills_top_k_whatever_the_similarity():
|
||||
"""There is still no relevance floor: an unrelated question selects a full
|
||||
`memory_top_k`. B.2 changed which memories those are, not how many — the
|
||||
early fact is no longer carried along by an unrelated question."""
|
||||
variants = scenario("independent_default")["ranking_variants"]
|
||||
assert variants["unrelated"]["selected_count"] == 5
|
||||
assert variants["unrelated"]["selected"] is False
|
||||
assert variants["unrelated"]["rank"] > 5
|
||||
|
||||
|
||||
# ------------------------------------------------- WP-B.2 ranking acceptance
|
||||
|
||||
@pytest.mark.parametrize("name", ["ranking_crowded", "ranking_context_dependent"])
|
||||
def test_acceptance_the_early_memory_is_ranked_and_injected_below_capacity(name):
|
||||
"""The B.1 ranking failure, made deterministic. On v1.0.0 both fixtures are
|
||||
`retained_but_not_ranked` (ranks 7 and 6 of 17 against a top-k of 4)."""
|
||||
result = scenario(name)
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
diagnosis = result["diagnosis"]
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert diagnosis["retained"]["active_memories"] <= diagnosis["retained"]["memory_bank_capacity"]
|
||||
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["rank"] <= 4
|
||||
assert diagnosis["ranked"]["replica_matches_stored_selection"] is True
|
||||
assert diagnosis["injected"]["yes"] is True
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_the_crowded_fixture_is_won_by_the_players_question():
|
||||
ranked = scenario("ranking_crowded")["diagnosis"]["ranked"]
|
||||
assert ranked["rank"] == 1
|
||||
assert ranked["lexical_score"] > 0 # "sundial" and "amber" are in the question
|
||||
|
||||
|
||||
def test_a_paraphrase_is_found_by_meaning_not_by_shared_words():
|
||||
"""Lexical matching must not replace semantic retrieval. The paraphrase
|
||||
shares none of F's distinctive words, yet ranks first."""
|
||||
for name in ("ranking_crowded", "independent_default"):
|
||||
paraphrase = scenario(name)["ranking_variants"]["paraphrase"]
|
||||
assert paraphrase["rank"] == 1 and paraphrase["selected"]
|
||||
# Only "Mara" is shared, which is far less than the direct question holds.
|
||||
direct = scenario(name)["ranking_variants"]["direct"]
|
||||
assert paraphrase["lexical_score"] < direct["lexical_score"] / 2
|
||||
|
||||
|
||||
def test_a_context_dependent_question_needs_the_scene():
|
||||
""""I ask her what she keeps up there" names nothing F's memory holds. The
|
||||
scene the last narration set up (Mara, the top shelf, a kettle) is what
|
||||
finds it; without that context it ranks last."""
|
||||
result = scenario("ranking_context_dependent")
|
||||
ranked = result["diagnosis"]["ranked"]
|
||||
assert ranked["lexical_score"] == 0.0
|
||||
assert ranked["rank"] <= 4
|
||||
assert "top shelf" in ranked["query"]["context"]
|
||||
assert result["ranking_variants"]["input_only"]["rank"] > 4
|
||||
|
||||
|
||||
def test_an_unrelated_rare_word_does_not_outrank_the_relevant_memory():
|
||||
"""Negative control: the paraphrase plus a place only one other memory
|
||||
holds. The decoy gains lexical score, and still ranks below F."""
|
||||
for name in ("ranking_crowded", "independent_default"):
|
||||
control = scenario(name)["ranking_variants"]["rare_word_with_paraphrase"]
|
||||
assert control["decoy_lexical_score"] > control["lexical_score"]
|
||||
assert control["rank"] == 1
|
||||
assert control["decoy_rank"] > control["rank"]
|
||||
|
||||
|
||||
def test_common_words_contribute_nothing():
|
||||
common = scenario("independent_default")["ranking_variants"]["common_words"]
|
||||
assert common["lexical_score"] == 0.0
|
||||
|
||||
|
||||
# ---------------------------------------------------------- capacity/eviction
|
||||
|
||||
@pytest.mark.parametrize("name", ["past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"])
|
||||
def test_past_capacity_the_early_memory_is_retained(name):
|
||||
"""v1.1 WP-B.2. On v1.0.0 all three are `created_but_evicted`: F was the
|
||||
least recently used row once recent narration stopped retrieving it, and
|
||||
went first (turns 21, 21 and 36). Coverage-first eviction keeps the only
|
||||
memory of the opening, so it stays active and is recalled at depth 106."""
|
||||
result = scenario(name)
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
diagnosis = result["diagnosis"]
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert result["eviction"]["f_evicted_at_turn"] is None
|
||||
# The bank really was past capacity, and stayed bounded.
|
||||
assert result["eviction"]["first_eviction_turn"] is not None
|
||||
assert all(t["active"] <= result["scenario"]["capacity"] + (1 if result["scenario"]["pin_first_memory"] else 0)
|
||||
for t in result["trace"])
|
||||
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
|
||||
|
||||
|
||||
def test_past_capacity_the_bank_still_describes_the_whole_story():
|
||||
"""What the rule buys in general, not only for F: the active bank reaches
|
||||
from the opening to the newest block, and no stretch between them goes
|
||||
undescribed for more than twice the average spacing a bank of this capacity
|
||||
can afford (story span / capacity). On v1.0.0 these banks began at depths 36
|
||||
and 18: the opening was simply gone."""
|
||||
for name in ("past_capacity", "past_capacity_low_top_k"):
|
||||
result = scenario(name)
|
||||
cover = result["trace"][-1]["coverage"]
|
||||
assert cover["first_start"] == 0
|
||||
assert cover["last_end"] >= result["recall_depth"] - 2 * memorybank.MEMORY_INTERVAL
|
||||
assert cover["largest_gap"] <= 2 * (cover["last_end"] + 1) / result["scenario"]["capacity"]
|
||||
|
||||
|
||||
def test_no_memory_is_evicted_by_the_same_pass_that_created_it():
|
||||
"""The frozen-bank regression the v1.0.0 rule fixed, still holding."""
|
||||
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
|
||||
assert scenario(name)["eviction"]["created_and_evicted_same_turn"] == []
|
||||
|
||||
|
||||
def test_a_pinned_memory_survives_capacity():
|
||||
eviction = scenario("past_capacity_pinned")["eviction"]
|
||||
assert eviction["pinned_memory_id"] is not None
|
||||
assert eviction["pinned_memory_forgotten"] is False
|
||||
|
||||
|
||||
def test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity():
|
||||
"""Was `xfail(strict=True)` in WP-B.1; B.2 fixed the eviction rule."""
|
||||
for name in ("past_capacity", "past_capacity_pinned", "past_capacity_low_top_k"):
|
||||
assert scenario(name)["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ---------------------------------------------------------- creation window
|
||||
|
||||
def test_a_fact_early_in_a_long_block_now_reaches_the_summariser():
|
||||
"""v1.1 WP-B.2. On v1.0.0 this block (2,079 tokens) was cut to its last
|
||||
2,000, the fact at its start was never seen, and the stage was
|
||||
`not_created`. The excerpt is now the block's opening and end."""
|
||||
result = scenario("long_block_fact_early")
|
||||
created = result["diagnosis"]["created"]
|
||||
covering = created["covering_memories"]
|
||||
assert covering, "the long block must have been summarised"
|
||||
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert covering[0]["fact_in_block"] is True
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is True
|
||||
assert created["yes"] is True
|
||||
assert created["source_start"] <= result["plant_depth"] <= created["source_end"]
|
||||
assert memorybank.EXCERPT_OMISSION_MARKER not in created["memory_text"]
|
||||
|
||||
|
||||
def test_the_same_fact_late_in_the_same_sized_block_does():
|
||||
result = scenario("long_block_fact_late")
|
||||
covering = result["diagnosis"]["created"]["covering_memories"]
|
||||
assert covering[0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert covering[0]["fact_in_summariser_excerpt"] is True
|
||||
assert result["diagnosis"]["created"]["yes"] is True
|
||||
|
||||
|
||||
def test_acceptance_a_fact_early_in_a_long_block_is_remembered():
|
||||
"""Was `xfail(strict=True)` in WP-B.1; B.2 changed the excerpt."""
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["created"]["yes"] is True
|
||||
assert scenario("long_block_fact_early")["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ------------------------------------------ WP-B.2 full deterministic acceptance
|
||||
|
||||
def test_acceptance_full_isolation_holds_on_every_turn():
|
||||
"""`independent_full`: long blocks, a crowded query, a bank past capacity.
|
||||
F must be carried by memory alone for the whole run, not only at recall."""
|
||||
result = scenario("independent_full")
|
||||
assert result["plant_depth"] <= 3 and result["recall_depth"] >= 100
|
||||
assert not any(t["f_in_state"] for t in result["trace"])
|
||||
assert not any(t["f_in_summary"] for t in result["trace"])
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
for check in ("state_document", "state_snapshots", "later_narration", "summary",
|
||||
"knowledge", "recent_history", "state_section"):
|
||||
assert result["isolation"]["checks"][check]["ok"], check
|
||||
|
||||
|
||||
def test_acceptance_full_every_stage_passes_past_capacity_with_long_blocks():
|
||||
"""On v1.0.0 this fixture fails at creation: every block is over 2,000
|
||||
tokens, and the fact at the start of the first one is never summarised."""
|
||||
result = scenario("independent_full")
|
||||
diagnosis = result["diagnosis"]
|
||||
assert result["trace"][-1]["total"] > result["scenario"]["capacity"]
|
||||
assert diagnosis["created"]["covering_memories"][0]["block_tokens"] > memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert diagnosis["created"]["yes"] and diagnosis["retained"]["yes"]
|
||||
assert diagnosis["ranked"]["yes"] and diagnosis["ranked"]["replica_matches_stored_selection"]
|
||||
assert diagnosis["injected"]["yes"]
|
||||
assert diagnosis["verdict"] == "injected"
|
||||
|
||||
|
||||
def test_acceptance_full_provenance_resolves_to_the_planting_turn():
|
||||
provenance = scenario("independent_full")["provenance"]
|
||||
assert provenance["recorded"] is not None
|
||||
assert provenance["range_covers_plant"] and provenance["matches_row"]
|
||||
assert provenance["source_block_holds_planting"] is True
|
||||
assert provenance["recorded"]["authority"] == memorybank.ACCEPTED_STORY
|
||||
|
||||
|
||||
def test_acceptance_full_is_the_long_run_independent_memory_verdict():
|
||||
"""The same measurements, judged by the long-run tool's own verdict."""
|
||||
from tools import m11_long_run as lr
|
||||
|
||||
result = scenario("independent_full")
|
||||
checks = result["isolation"]["checks"]
|
||||
diagnosis = result["diagnosis"]
|
||||
verdict = lr._independent_memory_verdict({
|
||||
"independent_planted_depth": result["plant_depth"],
|
||||
"planted_turn_outside_history": checks["recent_history"]["ok"],
|
||||
"absent_from_state": checks["state_document"]["ok"] and checks["state_snapshots"]["ok"]
|
||||
and not any(t["f_in_state"] for t in result["trace"]),
|
||||
"absent_from_summary": checks["summary"]["ok"]
|
||||
and not any(t["f_in_summary"] for t in result["trace"]),
|
||||
"absent_from_knowledge": checks["knowledge"]["ok"],
|
||||
"absent_from_later_narration": checks["later_narration"]["ok"],
|
||||
"memory_covering_planting_carries_fact": diagnosis["created"]["yes"],
|
||||
"memory_forgotten": not diagnosis["retained"]["yes"],
|
||||
"memory_injected": diagnosis["injected"]["yes"],
|
||||
})
|
||||
assert verdict == "recovered_through_memory_independent"
|
||||
|
||||
|
||||
# ------------------------------------------------------- lineage control (G)
|
||||
|
||||
def test_an_abandoned_lines_memory_is_stored_but_never_eligible_or_injected():
|
||||
g = scenario("lineage_control")["lineage_control"]
|
||||
assert g["memory_ids"], "G's memory must exist on line A before it is abandoned"
|
||||
assert sorted(g["stored"]) == sorted(g["memory_ids"])
|
||||
assert g["eligible_on_active_line"] == []
|
||||
assert g["g_text_ever_in_used_memories"] is False
|
||||
# Any turn that did name G's memory was on line A, before the divergence.
|
||||
assert g["eligible_after_returning_to_line_a"] == g["memory_ids"]
|
||||
|
||||
|
||||
def test_the_lineage_scenario_still_diagnoses_f_on_the_active_line():
|
||||
result = scenario("lineage_control")
|
||||
assert result["isolation"]["ok"], result["isolation"]
|
||||
assert result["diagnosis"]["verdict"] == "injected"
|
||||
|
||||
|
||||
# ----------------------------------------------------- authority control
|
||||
|
||||
@pytest.fixture()
|
||||
def authority_client(monkeypatch):
|
||||
embedder = md.ConceptEmbedder()
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email="b1-auth@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(user_id=user.id, model="script",
|
||||
endpoint_url="http://127.0.0.1:9/v1",
|
||||
embedding_model="concept-embed", memory_top_k=5))
|
||||
adventure = models.Adventure(user_id=user.id, title="auth", memory_bank_enabled=True,
|
||||
auto_summarize=True)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start", text="The tavern at dusk."))
|
||||
db.commit()
|
||||
adv, user_id = adventure.id, user.id
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", md.ScriptNarrator)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: embedder)
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: md.BestCaseSummariser())
|
||||
monkeypatch.setattr(memorybank, "schedule_post_turn", lambda a: None)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id))
|
||||
client = TestClient(app)
|
||||
client.adv = adv
|
||||
try:
|
||||
yield client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def test_a_memory_that_contradicts_state_loses_and_changes_nothing(authority_client):
|
||||
client, adv = authority_client, authority_client.adv
|
||||
corrected = client.post(f"/api/adventures/{adv}/state/corrections", json={"events": [
|
||||
{"type": "add_fact", "predicate": "the tavern lamp is lit", "fact_id": "lamp-lit"}]})
|
||||
assert corrected.status_code in (200, 201), corrected.text[:300]
|
||||
made = client.post(f"/api/adventures/{adv}/memories",
|
||||
json={"text": "The tavern lamp was never lit that night."})
|
||||
assert made.status_code == 201, made.text[:300]
|
||||
client.patch(f"/api/adventures/{adv}/memories/{made.json()['id']}", json={"pinned": True})
|
||||
asyncio.run(memorybank.run_post_turn(adv)) # embed it
|
||||
|
||||
before = client.get(f"/api/adventures/{adv}/state").json()["document"]
|
||||
md.ScriptNarrator.next_reply = 'The fire crackles.\n```state\n{"events": []}\n```'
|
||||
played = client.post(f"/api/adventures/{adv}/actions",
|
||||
json={"type": "do", "text": "I look at the lamp."})
|
||||
assert played.status_code == 200 and '"type": "error"' not in played.text
|
||||
after = client.get(f"/api/adventures/{adv}/state").json()["document"]
|
||||
assert after == before # retrieval mutated no state
|
||||
|
||||
with SessionLocal() as db:
|
||||
action = (db.query(models.Action).filter_by(adventure_id=adv, type="ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
snapshot = action.context_snapshot
|
||||
state_text = md._section(snapshot, md.STATE_LABEL)
|
||||
memory_text = md._section(snapshot, md.MEMORIES_LABEL)
|
||||
assert "the tavern lamp is lit" in state_text
|
||||
assert "never lit" in memory_text
|
||||
assert memory_text.startswith("Memories from earlier in the story")
|
||||
labels = [s["label"] for s in snapshot["sections"]]
|
||||
# State is read last of the live sections: it settles the conflict.
|
||||
assert labels.index(md.STATE_LABEL) > labels.index(md.MEMORIES_LABEL)
|
||||
@@ -0,0 +1,277 @@
|
||||
"""v1.1 WP-B.2 (B2.2): which memory a full bank lets go of.
|
||||
|
||||
v1.0.0 evicted the least recently used memory. B.1 showed that this discards
|
||||
the only memory of an early stretch first, because retrieval follows the present
|
||||
scene and nothing recent resembles it. `memorybank.eviction_order` now thins the
|
||||
bank where it is densest and keeps the opening and the newest stretch, with
|
||||
recency as the tie-break and least-recently-used as the fallback.
|
||||
|
||||
The scenario-level tests (the planted fact kept past capacity) are in
|
||||
`test_v11_b1_memory_diagnostic.py`; the v1.0.0 eviction tests in
|
||||
`test_memory_retrieval.py` still pass unchanged, because their memories carry
|
||||
no source range and take the fallback.
|
||||
|
||||
python -m pytest tests/test_v11_b2_memory_eviction.py -v
|
||||
"""
|
||||
|
||||
import random
|
||||
from collections import namedtuple
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import select
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.context import lineage
|
||||
from app.database import Base, SessionLocal, engine
|
||||
|
||||
T0 = datetime(2026, 1, 1, 12, 0, 0)
|
||||
Row = namedtuple("Row", "id pinned source_start source_end last_used_at created_at use_count")
|
||||
|
||||
|
||||
def row(id, start, end=None, *, pinned=False, used=None, created=None, uses=0):
|
||||
"""A memory as eviction sees it. Times are minutes after T0."""
|
||||
return Row(id, pinned, start, (start + 5) if end is None and start is not None else end,
|
||||
None if used is None else T0 + timedelta(minutes=used),
|
||||
T0 + timedelta(minutes=id if created is None else created), uses)
|
||||
|
||||
|
||||
def blocks(n, *, first_id=1):
|
||||
return [row(first_id + i, 6 * i) for i in range(n)]
|
||||
|
||||
|
||||
def largest_gap_from_opening(rows):
|
||||
"""The longest uncovered run of depths from depth 0 to the last memory."""
|
||||
ordered = sorted((r.source_start, r.source_end) for r in rows)
|
||||
gaps = [ordered[0][0]]
|
||||
reach = ordered[0][1]
|
||||
for start, end in ordered[1:]:
|
||||
gaps.append(max(0, start - reach - 1))
|
||||
reach = max(reach, end)
|
||||
return max(gaps)
|
||||
|
||||
|
||||
# ------------------------------------------------------------ the pure order
|
||||
|
||||
|
||||
def test_the_opening_and_the_newest_memory_are_kept():
|
||||
bank = blocks(7)
|
||||
doomed = memorybank.eviction_order(bank, 5)
|
||||
assert bank[0].id not in doomed and bank[-1].id not in doomed
|
||||
assert len(doomed) == 5
|
||||
|
||||
|
||||
def test_the_densest_stretch_is_thinned_first():
|
||||
# Memories every 6 depths to 30, then a sparse stretch. Removing one of the
|
||||
# dense ones leaves a 6-depth hole; removing a sparse one leaves far more.
|
||||
bank = [row(1, 0), row(2, 6), row(3, 12), row(4, 18), row(5, 60), row(6, 120), row(7, 180)]
|
||||
assert memorybank.eviction_order(bank, 1)[0] in {2, 3, 4}
|
||||
assert set(memorybank.eviction_order(bank, 2)) <= {2, 3, 4}
|
||||
|
||||
|
||||
def test_a_stretch_two_memories_describe_loses_one_of_them_first():
|
||||
"""A shared start (a re-played stretch, or a sibling line) leaves no hole.
|
||||
Of the two, the less recently used goes, even though a unique memory
|
||||
elsewhere is older and less used than both."""
|
||||
bank = [row(1, 0), row(2, 6, used=5), row(3, 12, used=50), row(4, 12, used=40), row(5, 18)]
|
||||
assert memorybank.eviction_order(bank, 1) == [4]
|
||||
|
||||
|
||||
def test_equal_holes_fall_to_the_least_recently_used():
|
||||
bank = [row(1, 0), row(2, 6, used=30), row(3, 12, used=10), row(4, 18, used=20), row(5, 24)]
|
||||
assert memorybank.eviction_order(bank, 1) == [3]
|
||||
|
||||
|
||||
def test_a_newborn_can_be_the_legitimate_first_to_go():
|
||||
"""The frozen bank is about a newborn losing to a count it cannot have yet.
|
||||
A newborn that only repeats a stretch another memory describes, one used
|
||||
after it was written, is legitimately the first to go."""
|
||||
bank = [row(1, 0), row(2, 6, used=100), row(3, 12), row(4, 6, created=90)]
|
||||
assert memorybank.eviction_order(bank, 1) == [4]
|
||||
|
||||
|
||||
def test_the_newest_memory_is_not_evicted_by_the_bank_it_joins():
|
||||
"""The frozen-bank regression under the new rule: every older memory has
|
||||
been used, the newborn never has, and it still stays."""
|
||||
bank = [row(i, 6 * (i - 1), used=200 + i, uses=3) for i in range(1, 6)]
|
||||
newborn = row(6, 30, created=300)
|
||||
assert newborn.id not in memorybank.eviction_order(bank + [newborn], 1)
|
||||
|
||||
|
||||
def test_pins_are_never_taken_but_still_count_as_coverage():
|
||||
bank = [row(1, 0), row(2, 6, pinned=True), row(3, 12), row(4, 18, pinned=True), row(5, 24)]
|
||||
doomed = memorybank.eviction_order(bank, 10)
|
||||
assert not {2, 4} & set(doomed)
|
||||
# With the pins covering 6 and 18, memory 3's hole is only its own block.
|
||||
assert doomed[0] == 3
|
||||
|
||||
|
||||
def test_memories_without_a_range_take_the_least_recently_used_fallback():
|
||||
hand_written = [row(1, None, None, used=30), row(2, None, None, used=10),
|
||||
row(3, None, None, used=20)]
|
||||
assert memorybank.eviction_order(hand_written, 3) == [2, 3, 1]
|
||||
|
||||
|
||||
def test_the_fallback_is_used_only_once_no_interior_memory_remains():
|
||||
bank = [row(1, 0, used=1), row(2, 6, used=90), row(3, 12, used=2),
|
||||
row(10, None, None, used=0)]
|
||||
order = memorybank.eviction_order(bank, 4)
|
||||
assert order[0] == 2 # the interior memory, although recently used
|
||||
assert order[1:] == [10, 1, 3] # then least recently used
|
||||
|
||||
|
||||
def test_the_order_does_not_depend_on_row_order():
|
||||
bank = [row(i, 6 * (i - 1), used=(i * 37) % 11, uses=i % 3) for i in range(1, 30)]
|
||||
bank += [row(40, 12), row(41, 12)] # a shared start with identical timestamps
|
||||
expected = memorybank.eviction_order(bank, 20)
|
||||
for seed in range(5):
|
||||
shuffled = bank[:]
|
||||
random.Random(seed).shuffle(shuffled)
|
||||
assert memorybank.eviction_order(shuffled, 20) == expected
|
||||
|
||||
|
||||
def test_a_tie_on_every_signal_is_broken_by_id():
|
||||
bank = [row(1, 0), row(9, 6, created=0), row(4, 12, created=0), row(20, 18)]
|
||||
assert memorybank.eviction_order(bank, 1) == [4]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("seed", range(8))
|
||||
def test_capacity_holds_and_pins_survive_for_any_bank(seed):
|
||||
rng = random.Random(seed)
|
||||
bank = []
|
||||
for i in range(1, rng.randint(2, 60)):
|
||||
start = None if rng.random() < 0.15 else rng.randrange(0, 400)
|
||||
bank.append(row(i, start, None if start is None else start + rng.choice([3, 5, 8]),
|
||||
pinned=rng.random() < 0.1, used=rng.choice([None, rng.randrange(500)]),
|
||||
uses=rng.randrange(4)))
|
||||
capacity = rng.randint(1, 30)
|
||||
overflow = len(bank) - capacity
|
||||
doomed = memorybank.eviction_order(bank, max(0, overflow))
|
||||
pinned = {r.id for r in bank if r.pinned}
|
||||
assert not pinned & set(doomed)
|
||||
assert len(set(doomed)) == len(doomed)
|
||||
remaining = len(bank) - len(doomed)
|
||||
assert remaining == max(capacity, len(pinned)) if overflow > 0 else remaining == len(bank)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("n, capacity, irregular", [(60, 10, False), (500, 80, False), (500, 80, True)])
|
||||
def test_a_long_bank_keeps_describing_the_whole_story(n, capacity, irregular):
|
||||
"""The general property, with nothing ever retrieved: memories arrive one
|
||||
block at a time and the bank is kept at capacity. The opening stays, and no
|
||||
stretch goes undescribed for more than twice the average spacing. Least
|
||||
recently used order, on the same arrivals, keeps only the newest stretch."""
|
||||
rng = random.Random(n)
|
||||
kept, lru = [], []
|
||||
depth = 0
|
||||
for i in range(1, n + 1):
|
||||
size = rng.choice([4, 6, 6, 9]) if irregular else 6
|
||||
memory = row(i, depth, depth + size - 1)
|
||||
depth += size
|
||||
kept.append(memory)
|
||||
lru.append(memory)
|
||||
if len(kept) > capacity:
|
||||
doomed = set(memorybank.eviction_order(kept, len(kept) - capacity))
|
||||
kept = [m for m in kept if m.id not in doomed]
|
||||
lru = sorted(lru, key=lambda m: (m.created_at, m.id))[len(lru) - capacity:]
|
||||
assert len(kept) == capacity
|
||||
assert min(m.source_start for m in kept) == 0
|
||||
assert max(m.id for m in kept) == n
|
||||
assert largest_gap_from_opening(kept) <= 2 * depth / capacity
|
||||
assert largest_gap_from_opening(lru) > depth / 2 # v1.0.0 order: the opening is gone
|
||||
|
||||
|
||||
# --------------------------------------------------------- on real rows
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
memorybank._vector_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def adventure(db):
|
||||
user = models.User(is_guest=False, email="b2-evict@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
settings = models.Settings(user_id=user.id, model="m", embedding_model="e",
|
||||
memory_bank_capacity=3)
|
||||
adv = models.Adventure(user_id=user.id, title="Evict", script_state={}, memory_bank_enabled=True)
|
||||
db.add_all([settings, adv])
|
||||
db.commit()
|
||||
adv.settings_row = settings
|
||||
return adv
|
||||
|
||||
|
||||
def test_the_pass_changes_nothing_but_forgotten(db, adventure):
|
||||
for i in range(6):
|
||||
memory = models.Memory(adventure_id=adventure.id, text=f"block {i}",
|
||||
source_start=6 * i, source_end=6 * i + 5, branch_id=None, depth=6 * i + 5)
|
||||
db.add(memory)
|
||||
db.commit()
|
||||
columns = (models.Memory.id, models.Memory.text, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.branch_id, models.Memory.depth,
|
||||
models.Memory.pinned, models.Memory.use_count, models.Memory.last_used_at)
|
||||
before = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
|
||||
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
|
||||
after = {r.id: tuple(r) for r in db.execute(select(*columns)).all()}
|
||||
assert before == after
|
||||
active = db.execute(select(models.Memory.id).where(models.Memory.forgotten.is_(False))).scalars().all()
|
||||
assert len(active) == 3
|
||||
assert min(active) == min(before) and max(active) == max(before) # the boundaries
|
||||
|
||||
|
||||
def test_eviction_does_not_make_an_abandoned_lines_memory_eligible(db, adventure):
|
||||
"""Eviction and lineage are separate: the pass decides only `forgotten`, so
|
||||
a memory on a line the story left is exactly as ineligible afterwards."""
|
||||
trunk = []
|
||||
for i in range(4):
|
||||
action = models.Action(adventure_id=adventure.id, type="ai", text=f"trunk {i}")
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
trunk.append(action)
|
||||
abandoned_node = models.Action(adventure_id=adventure.id, type="ai", text="the abandoned line")
|
||||
tree.place_action(db, adventure, abandoned_node)
|
||||
db.add(abandoned_node)
|
||||
db.flush()
|
||||
abandoned = models.Memory(adventure_id=adventure.id, text="on the abandoned line",
|
||||
source_start=4, source_end=4)
|
||||
tree.attach_memory(abandoned, abandoned_node)
|
||||
db.add(abandoned)
|
||||
db.commit()
|
||||
# Move the head back and diverge, so the abandoned node is off the path.
|
||||
adventure.head_depth = trunk[-1].depth
|
||||
db.commit()
|
||||
from app import head
|
||||
head.fork_if_behind_head(db, adventure)
|
||||
divergent = models.Action(adventure_id=adventure.id, type="ai", text="the new line")
|
||||
tree.place_action(db, adventure, divergent)
|
||||
db.add(divergent)
|
||||
db.flush()
|
||||
for i, node in enumerate(trunk + [divergent]):
|
||||
memory = models.Memory(adventure_id=adventure.id, text=f"active {i}",
|
||||
source_start=node.depth, source_end=node.depth)
|
||||
tree.attach_memory(memory, node)
|
||||
db.add(memory)
|
||||
db.commit()
|
||||
|
||||
def eligible():
|
||||
return set(db.execute(select(models.Memory.id).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory),
|
||||
models.Memory.forgotten.is_(False))).scalars().all())
|
||||
|
||||
assert abandoned.id not in eligible()
|
||||
memorybank._evict_over_capacity(adventure, adventure.settings_row, db)
|
||||
db.expire_all()
|
||||
assert abandoned.id not in eligible()
|
||||
assert len(db.execute(select(models.Memory.id).where(
|
||||
models.Memory.forgotten.is_(False))).scalars().all()) == 3
|
||||
@@ -0,0 +1,189 @@
|
||||
"""v1.1 WP-B.2 (B2.3): what the memory summariser is shown of a long block.
|
||||
|
||||
v1.0.0 sent the last 2,000 tokens of a block, so a fact early in a longer block
|
||||
never reached the summariser (B.1 §E). A block that fits is still sent whole. A
|
||||
longer one is now sent as its opening and its end, with a marker between them,
|
||||
inside the same 2,000-token budget.
|
||||
|
||||
The scenario-level test (the planted fact early in a long block, remembered) is
|
||||
in `test_v11_b1_memory_diagnostic.py`.
|
||||
|
||||
python -m pytest tests/test_v11_b2_memory_excerpt.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import random
|
||||
|
||||
import pytest
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine
|
||||
|
||||
BUDGET = memorybank.MEMORY_EXCERPT_TOKENS
|
||||
MARKER = memorybank.EXCERPT_OMISSION_MARKER
|
||||
FILLER = "The travellers walked the long grey road north past the salt market and the reed beds. "
|
||||
|
||||
|
||||
def words_to_tokens(tokens: int) -> str:
|
||||
"""Filler at least `tokens` long."""
|
||||
text = FILLER
|
||||
while builder.count_tokens(text) < tokens:
|
||||
text += FILLER
|
||||
return text
|
||||
|
||||
|
||||
# --------------------------------------------------------------- the excerpt
|
||||
|
||||
|
||||
def test_a_block_that_fits_is_sent_whole_and_unchanged():
|
||||
raw = words_to_tokens(BUDGET - 200)
|
||||
assert builder.count_tokens(raw) <= BUDGET
|
||||
assert memorybank.memory_excerpt(raw) == raw
|
||||
|
||||
|
||||
def test_a_block_of_exactly_the_budget_is_unchanged():
|
||||
raw = words_to_tokens(BUDGET)
|
||||
tokens = memorybank._excerpt_encoding().encode(raw)[:BUDGET]
|
||||
exact = memorybank._excerpt_encoding().decode(tokens)
|
||||
if builder.count_tokens(exact) == BUDGET:
|
||||
assert memorybank.memory_excerpt(exact) == exact
|
||||
|
||||
|
||||
def test_a_long_block_keeps_its_opening_and_its_end_in_order():
|
||||
opening = "Mara slipped the amber sundial inside the cracked teapot. "
|
||||
ending = "Aldric finally reached the north gate at dawn."
|
||||
raw = opening + words_to_tokens(3 * BUDGET) + ending
|
||||
excerpt = memorybank.memory_excerpt(raw)
|
||||
assert excerpt.startswith(opening)
|
||||
assert excerpt.endswith(ending)
|
||||
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
|
||||
assert tail, "the marker must sit between the two parts"
|
||||
assert excerpt.index(opening) < excerpt.index(MARKER) < excerpt.index(ending)
|
||||
|
||||
|
||||
def test_the_split_is_even_and_documented():
|
||||
raw = words_to_tokens(4 * BUDGET)
|
||||
head_budget, tail_budget = memorybank.excerpt_split(BUDGET)
|
||||
marker_tokens = builder.count_tokens(f"\n\n{MARKER}\n\n")
|
||||
assert head_budget + tail_budget + marker_tokens == BUDGET
|
||||
assert abs(head_budget - tail_budget) <= 1
|
||||
excerpt = memorybank.memory_excerpt(raw)
|
||||
head, _, tail = excerpt.partition(f"\n\n{MARKER}\n\n")
|
||||
# Each part is cut as a run of `head_budget` / `tail_budget` tokens. Measured
|
||||
# on its own, a cut run can come to one token more, because the text either
|
||||
# side of the cut tokenises differently once it is separated; the hard limit
|
||||
# is the whole excerpt, tested below.
|
||||
assert builder.count_tokens(head) <= head_budget + 1
|
||||
assert builder.count_tokens(tail) <= tail_budget + 1
|
||||
assert builder.count_tokens(excerpt) <= BUDGET
|
||||
|
||||
|
||||
@pytest.mark.parametrize("extra", [1, 7, 500, BUDGET, 9 * BUDGET])
|
||||
def test_the_excerpt_never_exceeds_the_budget(extra):
|
||||
raw = words_to_tokens(BUDGET + extra)
|
||||
excerpt = memorybank.memory_excerpt(raw)
|
||||
assert builder.count_tokens(excerpt) <= BUDGET
|
||||
assert memorybank.memory_excerpt(raw) == excerpt # deterministic
|
||||
|
||||
|
||||
@pytest.mark.parametrize("seed", range(4))
|
||||
def test_the_budget_holds_for_awkward_text(seed):
|
||||
"""Token boundaries can merge differently once the parts are rejoined, and
|
||||
text that is not plain English tokenises unevenly. The budget still holds."""
|
||||
rng = random.Random(seed)
|
||||
alphabet = "abcdefghij ÄÖÜ ßé漢字かな 🙂🐉 \n\t.,;:—'\""
|
||||
raw = "".join(rng.choice(alphabet) for _ in range(12000))
|
||||
assert builder.count_tokens(memorybank.memory_excerpt(raw)) <= BUDGET
|
||||
|
||||
|
||||
def test_a_fact_in_the_middle_of_a_very_long_block_is_still_omitted():
|
||||
"""The documented limit of a bounded excerpt: head and tail, not everything."""
|
||||
half = words_to_tokens(3 * BUDGET)
|
||||
raw = half + "Mara slipped the amber sundial inside the cracked teapot. " + half
|
||||
assert "sundial" not in memorybank.memory_excerpt(raw)
|
||||
|
||||
|
||||
# ------------------------------------------------------------ memory creation
|
||||
|
||||
|
||||
class EchoSummariser:
|
||||
"""Returns the whole excerpt it was given as the memory: the worst case for
|
||||
a marker leaking into stored text."""
|
||||
|
||||
def __init__(self):
|
||||
self.users: list[str] = []
|
||||
|
||||
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
|
||||
self.users.append(user)
|
||||
return user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def campaign(db, texts):
|
||||
user = models.User(is_guest=False, email="b2-excerpt@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
|
||||
adventure = models.Adventure(user_id=user.id, title="Excerpt", script_state={},
|
||||
auto_summarize=True, memory_bank_enabled=True)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
nodes = []
|
||||
for i, text in enumerate(texts):
|
||||
action = models.Action(adventure_id=adventure.id, type="ai" if i % 2 else "do", text=text)
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
nodes.append(action)
|
||||
db.commit()
|
||||
return adventure, nodes
|
||||
|
||||
|
||||
def write_memory(db, adventure, monkeypatch, summariser):
|
||||
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
|
||||
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: summariser)
|
||||
settings = db.query(models.Settings).first()
|
||||
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
|
||||
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
|
||||
|
||||
|
||||
def test_the_marker_is_never_stored_as_part_of_a_memory(db, monkeypatch):
|
||||
long = words_to_tokens(800)
|
||||
adventure, _ = campaign(db, [long] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK))
|
||||
summariser = EchoSummariser()
|
||||
[memory] = write_memory(db, adventure, monkeypatch, summariser)
|
||||
assert MARKER in summariser.users[0] # the summariser was told
|
||||
assert MARKER not in memory.text # and the memory does not repeat it
|
||||
assert "[…" not in memory.text and "omitted" not in memory.text
|
||||
|
||||
|
||||
def test_a_short_block_is_prompted_exactly_as_before(db, monkeypatch):
|
||||
texts = [f"Short action {i}." for i in range(memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK)]
|
||||
adventure, _ = campaign(db, texts)
|
||||
summariser = EchoSummariser()
|
||||
write_memory(db, adventure, monkeypatch, summariser)
|
||||
block = "\n\n".join(texts[:memorybank.MEMORY_INTERVAL])
|
||||
assert summariser.users[0] == f"Story excerpt:\n\n{block}\n\nMemory:"
|
||||
|
||||
|
||||
def test_a_long_blocks_memory_keeps_its_source_provenance(db, monkeypatch):
|
||||
early = "Mara slipped the amber sundial inside the cracked teapot. " + words_to_tokens(900)
|
||||
texts = [early] + [words_to_tokens(900)] * (memorybank.MEMORY_INTERVAL + memorybank.SETTLE_SLACK - 1)
|
||||
adventure, nodes = campaign(db, texts)
|
||||
[memory] = write_memory(db, adventure, monkeypatch, EchoSummariser())
|
||||
block = nodes[:memorybank.MEMORY_INTERVAL]
|
||||
assert "sundial" in memory.text # the early fact reached the summariser
|
||||
assert (memory.source_start, memory.source_end) == (block[0].depth, block[-1].depth)
|
||||
assert (memory.branch_id, memory.depth) == (block[-1].branch_id, block[-1].depth)
|
||||
@@ -0,0 +1,326 @@
|
||||
"""v1.1 WP-B.2 (B2.1): what memory retrieval searches for, and how it scores.
|
||||
|
||||
B.1 found the planting-era memory created and retained but ranked out of
|
||||
`memory_top_k`, because the query was three turns of narration with the player's
|
||||
question at the end. The query is now the player's input plus a short scene
|
||||
context, and the score adds one transparent lexical term over the input.
|
||||
|
||||
These tests pin the pieces. The end-to-end fixture tests (crowded bank,
|
||||
context-dependent question, negative controls) are in
|
||||
`test_v11_b1_memory_diagnostic.py`, beside the diagnostic they use.
|
||||
|
||||
python -m pytest tests/test_v11_b2_memory_ranking.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import math
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import event
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine
|
||||
|
||||
# ------------------------------------------------------------------ lexical
|
||||
|
||||
|
||||
def test_terms_are_folded_but_not_stemmed():
|
||||
terms = memorybank.lexical_terms("The tavern's teapots, the glass and the SUNDIAL")
|
||||
assert {"tavern", "teapot", "glass", "sundial"} <= terms
|
||||
assert "the" not in terms # the knowledge path's stop list
|
||||
assert "glas" not in terms # a word ending in "ss" is not a plural
|
||||
|
||||
|
||||
def test_a_word_every_candidate_holds_weighs_nothing():
|
||||
scores = memorybank.lexical_scores(
|
||||
frozenset({"travellers"}), {1: frozenset({"travellers", "road"}), 2: frozenset({"travellers"})})
|
||||
assert scores == {1: 0.0, 2: 0.0}
|
||||
|
||||
|
||||
def test_a_rarer_word_weighs_more_than_a_common_one():
|
||||
scores = memorybank.lexical_scores(
|
||||
frozenset({"sundial", "road"}),
|
||||
{1: frozenset({"sundial"}), 2: frozenset({"road"}), 3: frozenset({"road"}),
|
||||
4: frozenset({"gate"})})
|
||||
assert scores[1] > scores[2] == scores[3] > scores[4] == 0.0
|
||||
|
||||
|
||||
def test_a_single_rare_word_of_a_longer_question_is_only_its_share():
|
||||
"""The share is over the whole question, so one incidental word match
|
||||
cannot score like a memory that answers it."""
|
||||
question = frozenset({"where", "amber", "sundial", "fish"})
|
||||
scores = memorybank.lexical_scores(
|
||||
question, {1: frozenset({"fish"}), 2: frozenset({"amber", "sundial"}), 3: frozenset({"road"})})
|
||||
assert 0.0 < scores[1] < scores[2] <= 1.0
|
||||
assert scores[1] < 0.5
|
||||
|
||||
|
||||
def test_scores_are_bounded_and_empty_inputs_score_zero():
|
||||
candidates = {1: frozenset({"a1", "b2"}), 2: frozenset({"a1"})}
|
||||
assert all(0.0 <= v <= 1.0 for v in memorybank.lexical_scores(frozenset({"a1", "b2"}), candidates).values())
|
||||
assert memorybank.lexical_scores(frozenset(), candidates) == {1: 0.0, 2: 0.0}
|
||||
assert memorybank.lexical_scores(frozenset({"a1"}), {}) == {}
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ scoring
|
||||
|
||||
|
||||
def _unit(angle):
|
||||
return [math.cos(angle), math.sin(angle)]
|
||||
|
||||
|
||||
def test_ties_are_broken_by_id_not_by_row_order():
|
||||
held = {7: [1.0, 0.0], 3: [1.0, 0.0], 5: [1.0, 0.0]}
|
||||
rows = memorybank.score_candidates([7, 3, 5], held, {}, [1.0, 0.0], None, [])
|
||||
assert [row[1] for row in rows] == [3, 5, 7]
|
||||
|
||||
|
||||
def test_the_semantic_score_mixes_input_and_context_by_the_fixed_weight():
|
||||
held = {1: [1.0, 0.0]}
|
||||
[(final, _, semantic, lexical)] = memorybank.score_candidates(
|
||||
[1], held, {}, [1.0, 0.0], [0.0, 1.0], [])
|
||||
assert semantic == pytest.approx(memorybank.INPUT_WEIGHT)
|
||||
assert final == semantic and lexical == 0.0
|
||||
# Either part alone is used as it is.
|
||||
[(_, _, only_context, _)] = memorybank.score_candidates([1], held, {}, None, [0.0, 1.0], [])
|
||||
assert only_context == pytest.approx(0.0)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("margin, relevant_first", [(0.01, True), (-0.01, False)])
|
||||
def test_a_lexical_match_moves_a_memory_at_most_the_lexical_weight(margin, relevant_first):
|
||||
"""The bound that keeps rarity from overruling meaning: a memory more than
|
||||
`LEXICAL_WEIGHT` behind semantically cannot pass one ahead of it, however
|
||||
rare the word it shares."""
|
||||
decoy_cos = 1.0 - memorybank.LEXICAL_WEIGHT - margin
|
||||
held = {1: [1.0, 0.0], 2: _unit(math.acos(decoy_cos))}
|
||||
terms = {1: frozenset(), 2: frozenset({"zeppelin"})}
|
||||
rows = memorybank.score_candidates([1, 2], held, terms, [1.0, 0.0], None, ["zeppelin"])
|
||||
order = [row[1] for row in rows]
|
||||
assert (order[0] == 1) is relevant_first
|
||||
|
||||
|
||||
# ---------------------------------------------------- the query, on real rows
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
memorybank._terms_cache.clear()
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
memorybank._vector_cache.clear()
|
||||
memorybank._terms_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def restore_embedding_provider():
|
||||
real = memorybank.embedding_provider
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
memorybank.embedding_provider = real
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def settings(db):
|
||||
user = models.User(is_guest=False, email="b2-rank@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
row = models.Settings(user_id=user.id, model="m", embedding_model="stub-embed",
|
||||
memory_top_k=2, memory_bank_capacity=80)
|
||||
db.add(row)
|
||||
db.commit()
|
||||
return row
|
||||
|
||||
|
||||
SCENE_STATE = {
|
||||
"entities": {"mara": {"type": "character", "name": "Mara"},
|
||||
"tavern": {"type": "location", "name": "The Crooked Lantern"}},
|
||||
"scene": {"summary": "Closing time", "location": "tavern", "present": ["mara"]},
|
||||
}
|
||||
|
||||
|
||||
def make_adventure(db, settings, texts, state=None):
|
||||
"""`texts` is `[(type, text)]`, oldest first, each placed on the tree."""
|
||||
adventure = models.Adventure(user_id=settings.user_id, title="Rank", script_state={},
|
||||
memory_bank_enabled=True, narrative_state=state or {})
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
placed = []
|
||||
for kind, text in texts:
|
||||
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
placed.append(action)
|
||||
db.commit()
|
||||
return adventure, placed
|
||||
|
||||
|
||||
def test_the_query_is_the_players_input_and_the_scene(db, settings):
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("start", "Rain over the harbour."),
|
||||
("ai", "Mara wipes down the counter and glances up at the shelf."),
|
||||
("do", "> You ask Mara about the brass dial."),
|
||||
], state=SCENE_STATE)
|
||||
query = memorybank.retrieval_query(adventure)
|
||||
assert query["input"] == "> You ask Mara about the brass dial."
|
||||
assert "The Crooked Lantern" in query["context"] and "Mara" in query["context"]
|
||||
assert "glances up at the shelf" in query["context"]
|
||||
assert "brass" in query["input_terms"] and "dial" in query["input_terms"]
|
||||
|
||||
|
||||
def test_a_continue_turn_has_no_input_and_searches_by_the_scene(db, settings):
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("do", "> You sit down."),
|
||||
("ai", "The fire burns low in the grate."),
|
||||
])
|
||||
query = memorybank.retrieval_query(adventure)
|
||||
assert query["input"] == "" and query["input_terms"] == []
|
||||
assert "fire burns low" in query["context"]
|
||||
|
||||
|
||||
def test_a_retry_searches_with_the_input_it_is_retrying(db, settings):
|
||||
adventure, placed = make_adventure(db, settings, [
|
||||
("ai", "The market is quiet."),
|
||||
("do", "> You ask about the sundial."),
|
||||
("ai", "A discarded attempt about lanterns."),
|
||||
])
|
||||
query = memorybank.retrieval_query(adventure, exclude_action_id=placed[-1].id)
|
||||
assert query["input"] == "> You ask about the sundial."
|
||||
assert "lanterns" not in query["context"]
|
||||
assert "market is quiet" in query["context"]
|
||||
|
||||
|
||||
def test_the_query_is_bounded_however_long_the_story(db, settings):
|
||||
long = "The travellers walked the long grey road north past the salt market. " * 400
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("ai", long), ("story", long)], state=SCENE_STATE)
|
||||
query = memorybank.retrieval_query(adventure)
|
||||
assert builder.count_tokens(query["input"]) <= memorybank.QUERY_INPUT_TOKENS
|
||||
assert builder.count_tokens(query["context"]) <= (
|
||||
memorybank.QUERY_SCENE_TOKENS + memorybank.QUERY_NARRATION_TOKENS + 2)
|
||||
|
||||
|
||||
# ------------------------------------------------ retrieval, end to end
|
||||
|
||||
|
||||
class SameVector:
|
||||
"""Every text embeds the same, so only the lexical term separates memories."""
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.0, 0.0] for _ in texts]
|
||||
|
||||
|
||||
def add_memory(db, adventure, text, **kwargs):
|
||||
memory = models.Memory(adventure_id=adventure.id, text=text, **kwargs)
|
||||
db.add(memory)
|
||||
db.flush()
|
||||
memorybank.set_vector(memory, [1.0, 0.0, 0.0])
|
||||
db.commit()
|
||||
return memory
|
||||
|
||||
|
||||
def retrieve(adventure, settings, **kwargs):
|
||||
memorybank.embedding_provider = lambda s: SameVector()
|
||||
return asyncio.run(memorybank.retrieve_memories(adventure, settings, **kwargs))
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def played(db, settings):
|
||||
adventure, _ = make_adventure(db, settings, [
|
||||
("ai", "The tavern is warm."),
|
||||
("do", "> You ask Mara where the amber sundial went."),
|
||||
], state=SCENE_STATE)
|
||||
bank = {
|
||||
"road": add_memory(db, adventure, "Aldric walked the north road."),
|
||||
"sundial": add_memory(db, adventure, "Mara hid the amber sundial in the teapot."),
|
||||
"gate": add_memory(db, adventure, "The gate guard asked for a toll."),
|
||||
}
|
||||
return adventure, bank
|
||||
|
||||
|
||||
def test_every_used_memory_reports_the_parts_of_its_score(db, settings, played):
|
||||
adventure, bank = played
|
||||
result = retrieve(adventure, settings)
|
||||
first = result["used"][0]
|
||||
assert first["id"] == bank["sundial"].id
|
||||
assert first["similarity"] == first["semantic_score"]
|
||||
assert first["lexical_score"] > 0
|
||||
assert first["final_score"] == pytest.approx(
|
||||
first["semantic_score"] + memorybank.LEXICAL_WEIGHT * first["lexical_score"], abs=2e-4)
|
||||
assert result["query"]["input"] == "> You ask Mara where the amber sundial went."
|
||||
assert result["query"]["lexical_weight"] == memorybank.LEXICAL_WEIGHT
|
||||
assert result["query"]["input_weight"] == memorybank.INPUT_WEIGHT
|
||||
|
||||
|
||||
def test_a_pin_is_still_always_used_and_counts_toward_top_k(db, settings, played):
|
||||
adventure, bank = played
|
||||
settings.memory_top_k = 1
|
||||
bank["gate"].pinned = True
|
||||
db.commit()
|
||||
used = retrieve(adventure, settings)["used"]
|
||||
assert [m["id"] for m in used] == [bank["gate"].id]
|
||||
assert used[0]["pinned"] is True
|
||||
|
||||
|
||||
def memory_text_reads(statements):
|
||||
return [s for s in statements
|
||||
if s.lstrip().upper().startswith("SELECT") and "FROM memories" in s
|
||||
and "memories.text" in s]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def sql_log():
|
||||
statements: list[str] = []
|
||||
|
||||
def record(conn, cursor, statement, parameters, context, executemany):
|
||||
statements.append(statement)
|
||||
|
||||
event.listen(engine, "before_cursor_execute", record)
|
||||
try:
|
||||
yield statements
|
||||
finally:
|
||||
event.remove(engine, "before_cursor_execute", record)
|
||||
|
||||
|
||||
def test_memory_text_is_read_once_and_then_held(db, settings, played, sql_log):
|
||||
adventure, _ = played
|
||||
retrieve(adventure, settings)
|
||||
sql_log.clear()
|
||||
result = retrieve(adventure, settings)
|
||||
reads = memory_text_reads(sql_log)
|
||||
# Only the detail read of the memories chosen remains. (Every memory here
|
||||
# embeds identically, so redundancy suppression keeps just one of them.)
|
||||
assert len(reads) == 1 and reads[0].count("?") == len(result["used"])
|
||||
|
||||
|
||||
def test_a_continue_turn_reads_no_memory_text_to_rank(db, settings, sql_log):
|
||||
adventure, _ = make_adventure(db, settings, [("do", "> You wait."), ("ai", "Night falls.")])
|
||||
for text in ("one", "two", "three"):
|
||||
add_memory(db, adventure, f"memory {text}")
|
||||
sql_log.clear()
|
||||
result = retrieve(adventure, settings)
|
||||
assert all(m["lexical_score"] == 0.0 for m in result["used"])
|
||||
assert len(memory_text_reads(sql_log)) == 1 # the top-k detail read only
|
||||
|
||||
|
||||
def test_an_edited_memory_is_matched_on_its_new_text(db, settings, played):
|
||||
adventure, bank = played
|
||||
assert retrieve(adventure, settings)["used"][0]["id"] == bank["sundial"].id
|
||||
# An edit clears the vector (the route calls set_vector(None)); re-embedding
|
||||
# sets it again. Both go through set_vector, which drops the held terms.
|
||||
bank["road"].text = "The amber sundial was traded for the road toll."
|
||||
memorybank.set_vector(bank["road"], None)
|
||||
memorybank.set_vector(bank["road"], [1.0, 0.0, 0.0])
|
||||
bank["sundial"].text = "Mara hid a bottle in the cellar."
|
||||
memorybank.set_vector(bank["sundial"], None)
|
||||
memorybank.set_vector(bank["sundial"], [1.0, 0.0, 0.0])
|
||||
db.commit()
|
||||
assert retrieve(adventure, settings)["used"][0]["id"] == bank["road"].id
|
||||
@@ -0,0 +1,225 @@
|
||||
"""v1.1 WP-B.2: the memory summariser, after the rejected B2.4 prompt experiment.
|
||||
|
||||
B2.4 tried a memory prompt instructing the model to keep named facts and objects.
|
||||
Measured against the reference model, it did not correct the creation failure it
|
||||
was for, and it was not shipped (`V1.1-WP-B2-REPORT.md` §T). The shipped prompt
|
||||
is v1.0.0's.
|
||||
|
||||
This file keeps two kinds of test apart.
|
||||
|
||||
**Acceptance tests** gate the tree:
|
||||
- the shipped memory prompt is exactly v1.0.0's, so the experiment is gone;
|
||||
- every fidelity fixture reaches the summariser whole, through the application's
|
||||
own prompt assembly;
|
||||
- a long memory is stored as the model wrote it, never cut;
|
||||
- the memory the attempt-2 block should have produced ranks first under B2.1.
|
||||
|
||||
**Diagnostic-measurement tests** check only that `tools/memory_fidelity.py`
|
||||
measures correctly: fact retention, attribution, invention, word count, a leading
|
||||
"Memory:", second person and promise retention, on hand-written memories whose
|
||||
answers are known. What a real model scores on those measurements is
|
||||
nondeterministic, is taken with inference, and is reported. It is never a gate
|
||||
here.
|
||||
|
||||
python -m pytest tests/test_v11_b2_summarizer_fidelity.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import re
|
||||
import subprocess
|
||||
|
||||
import pytest
|
||||
|
||||
from app import memorybank, models, tree
|
||||
from app.database import Base, SessionLocal, engine
|
||||
from tools import memory_diagnostic as md
|
||||
from tools import memory_fidelity as mf
|
||||
|
||||
# ==================================================================== acceptance
|
||||
|
||||
|
||||
def test_the_shipped_memory_prompt_is_v1_0_0s():
|
||||
"""The B2.4 experiment is reverted: production sends the prompt v1.0.0 and
|
||||
WP-B.1 shipped, unchanged."""
|
||||
try:
|
||||
source = subprocess.run(["git", "show", "beb17ad:backend/app/memorybank.py"],
|
||||
capture_output=True, text=True, check=True).stdout
|
||||
except (OSError, subprocess.CalledProcessError):
|
||||
pytest.skip("git history not available")
|
||||
block = re.search(r"^MEMORY_SYSTEM_PROMPT = \((.*?)^\)$", source, re.S | re.M).group(1)
|
||||
shipped = eval(f"({block})", {"MEMORY_MAX_WORDS": 50}) # noqa: S307 - our own source
|
||||
assert memorybank.MEMORY_SYSTEM_PROMPT == shipped
|
||||
assert memorybank.MEMORY_MAX_WORDS == 50
|
||||
|
||||
|
||||
def test_the_rejected_experiment_is_not_what_ships():
|
||||
assert mf.B24_EXPERIMENT_PROMPT != memorybank.MEMORY_SYSTEM_PROMPT
|
||||
assert "Keep each fact with the person it belongs to" not in memorybank.MEMORY_SYSTEM_PROMPT
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
|
||||
def test_every_fixture_reaches_the_summariser_whole(fixture):
|
||||
"""Creation can only fail at the model if the fact was sent. Each fixture
|
||||
fits the excerpt budget, so the whole block is the excerpt."""
|
||||
user = mf.user_prompt_for(fixture)
|
||||
assert memorybank.count_tokens(fixture.raw) <= memorybank.MEMORY_EXCERPT_TOKENS
|
||||
assert f"Story excerpt:\n\n{fixture.raw}\n\nMemory:" in user
|
||||
assert user.startswith("Cast:\n- " + fixture.protagonist + " — the protagonist.")
|
||||
|
||||
|
||||
class Scripted:
|
||||
def __init__(self, reply):
|
||||
self.reply = reply
|
||||
self.calls: list[tuple[str, str]] = []
|
||||
|
||||
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
|
||||
self.calls.append((system, user))
|
||||
return self.reply
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def db():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
session = SessionLocal()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
session.close()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def campaign(db, fixture):
|
||||
user = models.User(is_guest=False, email="b2-fidelity@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(user_id=user.id, model="m", embedding_model=""))
|
||||
adventure = models.Adventure(user_id=user.id, title="Fidelity", script_state={}, auto_summarize=True,
|
||||
persona_name=fixture.protagonist)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
nodes = []
|
||||
for kind, text in fixture.actions + (("ai", "The story moves on."),):
|
||||
action = models.Action(adventure_id=adventure.id, type=kind, text=text)
|
||||
tree.place_action(db, adventure, action)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
nodes.append(action)
|
||||
db.commit()
|
||||
return adventure, nodes
|
||||
|
||||
|
||||
def write_memory(db, adventure, monkeypatch, provider):
|
||||
monkeypatch.setattr(memorybank, "MEMORY_START", 0)
|
||||
monkeypatch.setattr(memorybank, "MAX_MEMORIES_PER_RUN", 1)
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: provider)
|
||||
settings = db.query(models.Settings).first()
|
||||
asyncio.run(memorybank._create_due_memories(adventure, settings, db))
|
||||
return db.query(models.Memory).filter_by(adventure_id=adventure.id).all()
|
||||
|
||||
|
||||
def test_the_application_sends_the_shipped_prompt_and_the_whole_planting_block(db, monkeypatch):
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
adventure, nodes = campaign(db, fixture)
|
||||
provider = Scripted(fixture.faithful)
|
||||
[memory] = write_memory(db, adventure, monkeypatch, provider)
|
||||
system, user = provider.calls[0]
|
||||
assert system == memorybank.MEMORY_SYSTEM_PROMPT
|
||||
assert "> I watch Mara slip the amber sundial inside the cracked teapot" in user
|
||||
assert (memory.source_start, memory.source_end) == (nodes[0].depth, nodes[5].depth)
|
||||
|
||||
|
||||
def test_an_over_long_memory_is_stored_as_written_never_cut(db, monkeypatch):
|
||||
"""The word target is an instruction, not a truncation: cutting a memory
|
||||
after the fact can split or drop exactly the fact it was written to keep."""
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
adventure, _ = campaign(db, fixture)
|
||||
long_reply = fixture.faithful + " " + " ".join(["They advanced cautiously through the dark."] * 12)
|
||||
[memory] = write_memory(db, adventure, monkeypatch, Scripted(long_reply))
|
||||
assert memory.text == long_reply
|
||||
assert len(memory.text.split()) > 2 * memorybank.MEMORY_MAX_WORDS
|
||||
|
||||
|
||||
def test_a_faithful_regression_memory_ranks_first_for_its_question():
|
||||
"""If the summariser keeps the fact, B2.1 finds it: the memory the attempt-2
|
||||
block should have produced, among the memories its bank really held for that
|
||||
stretch, under production scoring."""
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
stored, _ = fixture.unfaithful[0]
|
||||
bank = {
|
||||
1: fixture.faithful,
|
||||
2: stored,
|
||||
3: "Aldric, Mara and Edrin advanced through the cold crypt, the silver key heavy in Aldric's hands.",
|
||||
4: "Aldric told Mara the silver key opens the crypt beneath the Old Abbey.",
|
||||
5: "Rain kept falling on Westhaven as the travellers walked toward the abbey grounds.",
|
||||
}
|
||||
embed = md.ConceptEmbedder.vector
|
||||
query = {"input": "> I ask Mara quietly where she hid the amber sundial.",
|
||||
"context": "Aldric and Mara in the Crooked Lantern, rain outside."}
|
||||
held = {i: embed(t) for i, t in bank.items()}
|
||||
terms = {i: memorybank.lexical_terms(t) for i, t in bank.items()}
|
||||
rows = memorybank.score_candidates(list(bank), held, terms, embed(query["input"]),
|
||||
embed(query["context"]),
|
||||
sorted(memorybank.lexical_terms(query["input"])))
|
||||
assert rows[0][1] == 1
|
||||
assert rows[0][3] > 0
|
||||
|
||||
|
||||
# ======================================================= diagnostic measurements
|
||||
# These prove the measuring instrument. They say nothing about any model.
|
||||
|
||||
|
||||
def test_the_fixtures_cover_each_measurement_in_more_than_one_genre():
|
||||
requirements = {f.requirement for f in mf.FIXTURES}
|
||||
assert {"distinctive object and place", "player-established concrete fact", "promise / commitment",
|
||||
"attribution", "clutter pressure", "no invention", "multiple concrete facts",
|
||||
"the actual failed-run block"} <= requirements
|
||||
assert {"office", "contemporary", "science-fiction-neutral"} <= {f.genre for f in mf.FIXTURES}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture", mf.FIXTURES, ids=lambda f: f.fixture_id)
|
||||
def test_the_checker_passes_a_faithful_memory(fixture):
|
||||
result = mf.evaluate(fixture, fixture.faithful)
|
||||
assert result["passed"], result
|
||||
assert not result["over_target"] and not result["memory_prefix"] and not result["second_person"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture, memory, reason", [
|
||||
(f, memory, reason) for f in mf.FIXTURES for memory, reason in f.unfaithful
|
||||
], ids=lambda v: v.fixture_id if isinstance(v, mf.Fixture) else None)
|
||||
def test_the_checker_fails_each_failure_shape(fixture, memory, reason):
|
||||
result = mf.evaluate(fixture, memory)
|
||||
assert not result["passed"], result
|
||||
if reason == "not retained":
|
||||
assert not result["retained"]
|
||||
elif reason == "misattributed":
|
||||
assert result["misattributed"]
|
||||
elif reason == "invented":
|
||||
assert result["inventions"]
|
||||
|
||||
|
||||
def test_the_checker_reads_the_stored_attempt_2_memory_as_the_real_failure():
|
||||
fixture = mf.FIXTURES_BY_ID["regression_attempt_2"]
|
||||
stored, _ = fixture.unfaithful[0]
|
||||
result = mf.evaluate(fixture, stored)
|
||||
assert result["retained"] is False and result["words"] == 102 and result["over_target"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("memory, prefix, you", [
|
||||
("Memory: Dana promised Marcus the lease by Friday.", True, False),
|
||||
(" memory: Dana promised the lease.", True, False),
|
||||
("You thanked Marcus and left.", False, True),
|
||||
("Dana thanked Marcus; your lease is due.", False, True),
|
||||
("Dana promised Marcus she would bring the signed lease by Friday.", False, False),
|
||||
])
|
||||
def test_the_checker_measures_framing(memory, prefix, you):
|
||||
result = mf.evaluate(mf.FIXTURES_BY_ID["promise_contemporary"], memory)
|
||||
assert result["memory_prefix"] is prefix
|
||||
assert result["second_person"] is you
|
||||
|
||||
|
||||
def test_the_checker_measures_promise_retention():
|
||||
fixture = mf.FIXTURES_BY_ID["promise_contemporary"]
|
||||
kept = mf.evaluate(fixture, "Dana promised to bring Marcus the signed lease by Friday.")
|
||||
scenery = mf.evaluate(fixture, "Memory: Dana looked around the empty living room while a dog barked.")
|
||||
assert kept["facts"]["lease by Friday"]["kept"] and kept["passed"]
|
||||
assert not scenery["facts"]["lease by Friday"]["kept"] and scenery["memory_prefix"]
|
||||
@@ -0,0 +1,91 @@
|
||||
"""v1.1 WP-C: the browser harness's download helpers, without a browser.
|
||||
|
||||
`tools/m11_webdriver.wait_for_download` is what decides that an export actually
|
||||
left the browser as a file. It must never call a download finished because a
|
||||
file appeared, because it is still being written, or because it is empty — each
|
||||
of those would make "the export works" a claim the evidence does not support.
|
||||
|
||||
python -m pytest tests/test_v11_c_browser_helpers.py -v
|
||||
"""
|
||||
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from tools import m11_webdriver as wd
|
||||
|
||||
|
||||
def later(seconds, action):
|
||||
timer = threading.Timer(seconds, action)
|
||||
timer.start()
|
||||
return timer
|
||||
|
||||
|
||||
def test_a_finished_file_is_returned(tmp_path):
|
||||
later(0.2, lambda: (tmp_path / "campaign.json").write_text('{"format": "x"}'))
|
||||
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05)
|
||||
assert found == tmp_path / "campaign.json"
|
||||
|
||||
|
||||
def test_a_file_that_was_already_there_is_not_the_download(tmp_path):
|
||||
(tmp_path / "old.json").write_text("{}")
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.wait_for_download(tmp_path, {"old.json"}, timeout=0.6, poll=0.05)
|
||||
|
||||
|
||||
def test_an_empty_file_never_counts(tmp_path):
|
||||
(tmp_path / "empty.json").write_bytes(b"")
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
|
||||
|
||||
|
||||
def test_nothing_counts_while_firefox_is_still_writing(tmp_path):
|
||||
"""Firefox writes `<name>.part` beside the final name until it is done."""
|
||||
(tmp_path / "campaign.json").write_text('{"format": "x"}')
|
||||
(tmp_path / "campaign.json.part").write_text("")
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.wait_for_download(tmp_path, set(), timeout=0.6, poll=0.05)
|
||||
later(0.1, lambda: (tmp_path / "campaign.json.part").unlink())
|
||||
assert wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05).name == "campaign.json"
|
||||
|
||||
|
||||
def test_a_file_that_is_still_growing_is_not_finished(tmp_path):
|
||||
target = tmp_path / "big.json"
|
||||
target.write_text("{")
|
||||
stop = threading.Event()
|
||||
|
||||
def grow():
|
||||
for _ in range(8):
|
||||
if stop.is_set():
|
||||
return
|
||||
with target.open("a") as fh:
|
||||
fh.write("x" * 100)
|
||||
time.sleep(0.05)
|
||||
|
||||
writer = threading.Thread(target=grow)
|
||||
started = time.monotonic()
|
||||
writer.start()
|
||||
found = wd.wait_for_download(tmp_path, set(), timeout=5, poll=0.05, stable_polls=3)
|
||||
writer.join()
|
||||
# Returned only once the size held still, so after the last write.
|
||||
assert found == target
|
||||
assert target.stat().st_size == 1 + 8 * 100
|
||||
assert time.monotonic() - started >= 0.4
|
||||
|
||||
|
||||
def test_the_prefs_save_downloads_unasked_to_the_folder_given(tmp_path):
|
||||
prefs = wd.firefox_download_prefs(tmp_path)
|
||||
assert prefs["browser.download.folderList"] == 2
|
||||
assert prefs["browser.download.dir"] == str(tmp_path)
|
||||
assert prefs["browser.download.useDownloadDir"] is True
|
||||
assert prefs["browser.download.always_ask_before_handling_new_types"] is False
|
||||
assert "application/json" in prefs["browser.helperApps.neverAsk.saveToDisk"]
|
||||
|
||||
|
||||
def test_a_download_folder_must_be_under_home():
|
||||
with pytest.raises(wd.WebDriverError):
|
||||
wd.require_under_home(Path("/tmp/wp-c-downloads"))
|
||||
inside = Path.home() / "v11-evidence" / "wp-c" / "downloads"
|
||||
assert wd.require_under_home(inside) == inside.resolve()
|
||||
@@ -0,0 +1,348 @@
|
||||
"""v1.1 WP-A1 corrective: a cold model is loaded, not guessed about.
|
||||
|
||||
A1's accounting caught a real cold-model turn: `/api/ps` knew nothing because the
|
||||
model was not resident, `/api/show` found no `num_ctx`, the window was therefore
|
||||
unverified, and the prompt was built to the configured 16,384. Ollama loaded the
|
||||
model at its own 4,096 default, read 2,050 of the 13,875 tokens and answered 200.
|
||||
|
||||
Detection was right. The case is also preventable: once the model is loaded its
|
||||
window is readable. So before an unverified turn is assembled, the application
|
||||
asks the configured server, once, to load the model (`POST /api/generate` with a
|
||||
model and no prompt, which Ollama answers with `"done_reason": "load"` and no
|
||||
text), probes again, and builds the turn to whatever that probe says. A window
|
||||
still unverified afterwards changes nothing: the configured budget stands and
|
||||
the post-response accounting still watches for a cut prompt.
|
||||
|
||||
python -m pytest tests/test_v11_cold_window.py -v
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy import text
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, contextwindow, limits, models
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.providers.base import ProviderError
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
ENDPOINT = "http://127.0.0.1:11434/v1"
|
||||
MODEL = "qwen2.5:3b-instruct"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_window_cache():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
class ColdOllama:
|
||||
"""The shapes a real Ollama 0.33 returned, with a model that starts cold.
|
||||
|
||||
`/api/ps` lists only loaded models. `/api/show` carries no `num_ctx`.
|
||||
`/api/generate` with no prompt loads the model at `load_window`, exactly as
|
||||
the real server answered: HTTP 200, `"response": ""`, `"done_reason": "load"`.
|
||||
"""
|
||||
|
||||
def __init__(self, *, loaded=None, load_window=4096, generate_status=200,
|
||||
report_after_load=True):
|
||||
self.loaded = dict(loaded or {})
|
||||
self.load_window = load_window
|
||||
self.generate_status = generate_status
|
||||
self.report_after_load = report_after_load
|
||||
self.requests: list[tuple[str, str, dict | None]] = []
|
||||
|
||||
def handler(self, request: httpx.Request) -> httpx.Response:
|
||||
body = None
|
||||
if request.content:
|
||||
import json
|
||||
body = json.loads(request.content)
|
||||
self.requests.append((request.method, str(request.url), body))
|
||||
path = request.url.path
|
||||
if path == "/api/ps":
|
||||
return httpx.Response(200, json={"models": [
|
||||
{"name": name, "model": name, "context_length": tokens}
|
||||
for name, tokens in self.loaded.items()
|
||||
]})
|
||||
if path == "/api/show":
|
||||
return httpx.Response(200, json={
|
||||
"model_info": {"qwen2.context_length": 32768}, "parameters": ""})
|
||||
if path == "/api/generate":
|
||||
if self.generate_status != 200:
|
||||
return httpx.Response(self.generate_status, json={"error": "model not found"})
|
||||
if self.report_after_load:
|
||||
self.loaded[body["model"]] = self.load_window
|
||||
return httpx.Response(200, json={
|
||||
"model": body["model"], "response": "", "done": True, "done_reason": "load"})
|
||||
return httpx.Response(404)
|
||||
|
||||
def paths(self):
|
||||
return [httpx.URL(url).path for _method, url, _body in self.requests]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def server(monkeypatch):
|
||||
def install(fake: ColdOllama):
|
||||
original = httpx.AsyncClient
|
||||
|
||||
def build(*args, **kwargs):
|
||||
kwargs.pop("verify", None)
|
||||
return original(*args, transport=httpx.MockTransport(fake.handler), **kwargs)
|
||||
|
||||
monkeypatch.setattr(contextwindow.httpx, "AsyncClient", build)
|
||||
return fake
|
||||
|
||||
return install
|
||||
|
||||
|
||||
# ------------------------------------------------------------ ensure_window
|
||||
|
||||
def test_a_cold_model_is_loaded_once_and_its_window_verified(server):
|
||||
fake = server(ColdOllama(loaded={}, load_window=4096))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert (window.tokens, window.source, window.verified) == (4096, contextwindow.LOADED, True)
|
||||
assert preflight == {"attempted": True, "loaded": True, "verified_before": False,
|
||||
"verified_after": True,
|
||||
"detail": "the server loaded the model (load)"}
|
||||
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate", "/api/ps"]
|
||||
# One load request, naming the model and nothing else: no prompt, so no text.
|
||||
warms = [body for _m, url, body in fake.requests if url.endswith("/api/generate")]
|
||||
assert warms == [{"model": MODEL}]
|
||||
|
||||
|
||||
def test_an_already_loaded_model_is_not_warmed(server):
|
||||
fake = server(ColdOllama(loaded={MODEL: 16384}))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert window.verified and window.tokens == 16384
|
||||
assert preflight["attempted"] is False
|
||||
assert "/api/generate" not in fake.paths()
|
||||
|
||||
|
||||
def test_a_model_that_loads_but_still_cannot_be_read_stays_unverified(server):
|
||||
"""A server that loads the model but whose `/api/ps` still cannot say. The
|
||||
existing unknown path stands: no guessed window, the configured budget kept."""
|
||||
fake = server(ColdOllama(loaded={}, report_after_load=False))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert not window.verified and window.tokens is None
|
||||
assert preflight["attempted"] is True and preflight["loaded"] is True
|
||||
assert preflight["verified_after"] is False
|
||||
assert fake.paths().count("/api/generate") == 1
|
||||
assert contextwindow.effective_budget(16384, window) == 16384
|
||||
|
||||
|
||||
@pytest.mark.parametrize("status", [404, 500])
|
||||
def test_a_failed_load_is_recorded_and_leaves_the_window_unverified(server, status):
|
||||
fake = server(ColdOllama(loaded={}, generate_status=status))
|
||||
window, preflight = asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert not window.verified
|
||||
assert preflight["attempted"] is True and preflight["loaded"] is False
|
||||
assert f"HTTP {status}" in preflight["detail"]
|
||||
# Bounded: one attempt, and no second probe after a failed load.
|
||||
assert fake.paths() == ["/api/ps", "/api/show", "/api/generate"]
|
||||
|
||||
|
||||
def test_a_declared_window_does_not_stop_the_server_being_asked(server):
|
||||
"""A declaration fills a hole the server leaves. Loading the model can close
|
||||
the hole, and a verified answer always wins over a declaration."""
|
||||
server(ColdOllama(loaded={}, load_window=4096))
|
||||
window, _preflight = asyncio.run(
|
||||
contextwindow.ensure_window(ENDPOINT, MODEL, declared=8192))
|
||||
assert (window.tokens, window.source) == (4096, contextwindow.LOADED)
|
||||
|
||||
|
||||
def test_an_unreachable_server_is_not_asked_to_load_anything():
|
||||
"""Nothing listens here. No load is attempted against a server that did not
|
||||
answer the probe, so an offline turn costs no second timeout."""
|
||||
window, preflight = asyncio.run(
|
||||
contextwindow.ensure_window("http://127.0.0.1:1/v1", MODEL))
|
||||
assert not window.verified
|
||||
assert preflight["attempted"] is False
|
||||
assert "did not answer" in preflight["detail"]
|
||||
|
||||
|
||||
def test_the_load_request_obeys_the_endpoint_policy():
|
||||
"""ADR 011. No transport is installed: a load that ignored the policy would
|
||||
try to reach a public address for real."""
|
||||
for url in ("https://api.openai.com/v1", "http://8.8.8.8:11434/v1"):
|
||||
loaded, detail = asyncio.run(contextwindow.warm(url, MODEL, timeout=2))
|
||||
assert loaded is False
|
||||
assert "not allowed" in detail
|
||||
|
||||
|
||||
def test_the_load_request_goes_only_to_the_configured_host(server):
|
||||
fake = server(ColdOllama(loaded={}))
|
||||
asyncio.run(contextwindow.ensure_window("http://192.168.0.50:11434/v1", MODEL))
|
||||
hosts = {httpx.URL(url).host for _m, url, _b in fake.requests}
|
||||
ports = {httpx.URL(url).port for _m, url, _b in fake.requests}
|
||||
assert hosts == {"192.168.0.50"} and ports == {11434}
|
||||
|
||||
|
||||
# ------------------------------------------------------------- end to end
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="v11cold@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
|
||||
context_token_budget=16384, max_output_tokens=500,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Cold",
|
||||
campaign_canon={"rules": ["The sealed crypt is named CANON-SENTINEL-COLD-2050."]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(adventure_id=adventure.id, type="start", text="Rain."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _long_story(adv_id, turns=120):
|
||||
from app import tree
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
for i in range(turns):
|
||||
for kind, body in (
|
||||
("do", f"I search the {i}th chamber of the undercroft."),
|
||||
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
|
||||
):
|
||||
action = models.Action(adventure_id=adv_id, type=kind, text=body)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
tree.place_action(db, adventure, action)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _counts():
|
||||
with SessionLocal() as db:
|
||||
return {table: db.execute(text(f"SELECT COUNT(*) FROM {table}")).scalar()
|
||||
for table in ("actions", "state_events", "state_proposals", "memories",
|
||||
"summaries")}
|
||||
|
||||
|
||||
def _latest_ai(adv_id):
|
||||
with SessionLocal() as db:
|
||||
return (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
|
||||
|
||||
def test_a_cold_turn_is_built_to_the_window_the_loaded_model_reports(client, server):
|
||||
"""The observed failure, prevented. Without the load this turn would be built
|
||||
to the configured 16,384 against a 4,096 server."""
|
||||
_long_story(client.adv_id)
|
||||
fake = server(ColdOllama(loaded={}, load_window=4096))
|
||||
ScriptedProvider.replies = ["The seal holds."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
assert snapshot["window"]["verified"] is True
|
||||
assert snapshot["tokens"]["budget"] == 4096
|
||||
assert snapshot["window"]["preflight"]["attempted"] is True
|
||||
assert snapshot["window"]["preflight"]["verified_after"] is True
|
||||
system, story = ScriptedProvider.prompts[-1]
|
||||
sent = builder.count_tokens(system) + builder.count_tokens(story)
|
||||
assert sent + snapshot["tokens"]["transport"] + 500 + 256 <= 4096
|
||||
assert "CANON-SENTINEL-COLD-2050" in system
|
||||
assert fake.paths().count("/api/generate") == 1
|
||||
|
||||
|
||||
def test_without_the_load_the_same_cold_turn_would_have_been_built_too_large(client, server,
|
||||
monkeypatch):
|
||||
"""The negative control: v1.0.0 and the first A1 tree probed only."""
|
||||
_long_story(client.adv_id)
|
||||
server(ColdOllama(loaded={}, load_window=4096))
|
||||
|
||||
async def probe_only(endpoint_url, model, *, declared=None, warm_timeout=300.0):
|
||||
window = await contextwindow.probe(endpoint_url, model, declared=declared)
|
||||
return window, {"attempted": False}
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "ensure_window", probe_only)
|
||||
ScriptedProvider.replies = ["The seal holds."]
|
||||
client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
assert snapshot["window"]["verified"] is False
|
||||
assert snapshot["tokens"]["budget"] == 16384
|
||||
system, story = ScriptedProvider.prompts[-1]
|
||||
assert builder.count_tokens(system) + builder.count_tokens(story) > 4096 * 2
|
||||
|
||||
|
||||
def test_the_load_itself_writes_nothing(client, server):
|
||||
"""No action, narration, state event, proposal, memory or summary comes from
|
||||
the preflight: it is a request to the server and nothing else."""
|
||||
fake = server(ColdOllama(loaded={}))
|
||||
before = _counts()
|
||||
asyncio.run(contextwindow.ensure_window(ENDPOINT, MODEL))
|
||||
assert _counts() == before
|
||||
assert fake.paths().count("/api/generate") == 1
|
||||
|
||||
|
||||
def test_a_failed_load_then_a_failed_model_call_leaves_the_story_safe(client, server):
|
||||
"""The ordinary failure semantics: the error is reported, no narration is
|
||||
accepted, and nothing about the state changes."""
|
||||
server(ColdOllama(loaded={}, generate_status=404))
|
||||
before = _counts()
|
||||
ScriptedProvider.replies = [ProviderError("Endpoint or model not found (HTTP 404).")]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "open the door"})
|
||||
assert response.status_code == 200
|
||||
assert '"type": "error"' in response.text or '"error"' in response.text
|
||||
after = _counts()
|
||||
assert after["state_events"] == before["state_events"]
|
||||
assert after["state_proposals"] == before["state_proposals"]
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.Action).filter_by(adventure_id=client.adv_id,
|
||||
type="ai").count() == 0
|
||||
|
||||
|
||||
def test_a_failed_load_does_not_stop_a_turn_the_model_can_still_answer(client, server):
|
||||
server(ColdOllama(loaded={}, generate_status=500))
|
||||
ScriptedProvider.replies = ["The door opens."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "open the door"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
assert snapshot["window"]["verified"] is False
|
||||
assert snapshot["window"]["preflight"]["loaded"] is False
|
||||
assert snapshot["tokens"]["budget"] == 16384
|
||||
assert snapshot["accounting"]["status"] == contextwindow.UNKNOWN
|
||||
|
||||
|
||||
def test_the_context_dry_run_never_loads_a_model(client, server):
|
||||
fake = server(ColdOllama(loaded={}))
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/context")
|
||||
assert response.status_code == 200
|
||||
assert "/api/generate" not in fake.paths()
|
||||
@@ -0,0 +1,495 @@
|
||||
"""v1.1 WP-A1: a deliberate safety reserve, and a turn the server cut is not silent.
|
||||
|
||||
M11 made the verified window a ceiling. It did not make the application's count
|
||||
the server's count. The application counts with `cl100k_base`, the narrator with
|
||||
its own tokenizer, and the v1 evidence left 23-42 real tokens between the largest
|
||||
prompt and the edge of a 16,384 window. Past that edge Ollama does not refuse.
|
||||
Measured against the reference CPU host (Ollama 0.33, a 4,096 window), a
|
||||
6,316-token prompt came back 200 with `prompt_tokens` 2,050: the front of the
|
||||
prompt, which in this design is the narrator's rules and the canon, was gone.
|
||||
|
||||
So the tests below are in three halves.
|
||||
|
||||
**The reserve.** `max(256, ceil(5% of the effective window))`, taken from the
|
||||
budget before any history is chosen, on top of an exact reply allocation.
|
||||
|
||||
**The arithmetic.** The assembled prompt, plus the application text the provider
|
||||
adds to every request, plus the reply allocation, plus the reserve, fits the
|
||||
effective window. Protected context that cannot fit that way fails before the
|
||||
model is called.
|
||||
|
||||
**The accounting.** Where the server reports how many prompt tokens it read, the
|
||||
turn records `fits`, `exceeded` or `truncation_suspected`. Where it reports
|
||||
nothing, the turn says `unknown`, never `fits`. A discrepancy found after the
|
||||
reply is recorded and shown; it never costs the reader an accepted turn.
|
||||
|
||||
python -m pytest tests/test_v11_context_reserve.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, contextwindow, limits, models
|
||||
from app.context import builder
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.providers.openai_compatible import CHAT_CONTINUE_HINT, OpenAICompatibleProvider
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
ENDPOINT = "http://127.0.0.1:11434/v1"
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clear_window_cache():
|
||||
contextwindow.cache_clear()
|
||||
yield
|
||||
contextwindow.cache_clear()
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the reserve
|
||||
|
||||
@pytest.mark.parametrize("window, reserve", [
|
||||
(1024, 256),
|
||||
(4096, 256), # 5% is 204.8, so the floor holds
|
||||
(5120, 256), # exactly 5% is the floor
|
||||
(5121, 257), # 256.05 rounds up
|
||||
(8192, 410), # 409.6 rounds up
|
||||
(16384, 820), # 819.2 rounds up
|
||||
(32768, 1639), # 1638.4 rounds up
|
||||
])
|
||||
def test_the_reserve_is_the_larger_of_the_floor_and_five_percent_rounded_up(window, reserve):
|
||||
assert contextwindow.safety_reserve(window) == reserve
|
||||
|
||||
|
||||
def test_the_reserve_is_far_larger_than_the_v1_margin_at_the_evidence_window():
|
||||
"""The v1 evidence left 23-42 tokens at 16,384. 64 tokens of slack was all
|
||||
the arithmetic kept for drift and separators together."""
|
||||
assert contextwindow.safety_reserve(16384) >= 10 * 64
|
||||
|
||||
|
||||
# ---------------------------------------------------------- the arithmetic
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="v11reserve@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="qwen2.5:3b-instruct", endpoint_url=ENDPOINT,
|
||||
embedding_model="", context_token_budget=16384, max_output_tokens=500,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Reserved",
|
||||
campaign_canon={"rules": [
|
||||
"The abbey seal has never been broken.",
|
||||
"The sealed crypt is named CANON-SENTINEL-RESERVE-5120.",
|
||||
]},
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text="Rain over Westhaven."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def _long_story(adv_id, turns=120):
|
||||
from app import tree
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
for i in range(turns):
|
||||
for kind, text in (
|
||||
("do", f"I search the {i}th chamber of the undercroft."),
|
||||
("ai", "The lantern gutters. " + ("Cold stone, and older dust. " * 40)),
|
||||
):
|
||||
action = models.Action(adventure_id=adv_id, type=kind, text=text)
|
||||
db.add(action)
|
||||
db.flush()
|
||||
tree.place_action(db, adventure, action)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _settings(**changes):
|
||||
with SessionLocal() as db:
|
||||
settings = db.query(models.Settings).first()
|
||||
for key, value in changes.items():
|
||||
setattr(settings, key, value)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _build(client, window):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
settings = db.query(models.Settings).first()
|
||||
return builder.build_context(adventure, settings, window=window)
|
||||
|
||||
|
||||
def _sent(system, story) -> int:
|
||||
"""What the provider actually sends in chat mode, by the application's count."""
|
||||
return (builder.count_tokens(system) + builder.count_tokens(story)
|
||||
+ builder.count_tokens(CHAT_CONTINUE_HINT))
|
||||
|
||||
|
||||
CONFIGURATIONS = {
|
||||
"verified 4,096": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(4096, contextwindow.LOADED)),
|
||||
"verified 8,192": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(8192, contextwindow.PARAMETERS)),
|
||||
"verified 16,384": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(16384, contextwindow.LOADED)),
|
||||
"declared 6,000": (dict(context_token_budget=16384),
|
||||
contextwindow.Window(6000, contextwindow.DECLARED)),
|
||||
"unverified, configured 12,000": (dict(context_token_budget=12000),
|
||||
contextwindow.UNVERIFIED),
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", list(CONFIGURATIONS))
|
||||
def test_the_prompt_leaves_the_reply_and_the_reserve_free(client, name):
|
||||
"""A1-2, on the assembled text rather than the builder's own arithmetic."""
|
||||
changes, window = CONFIGURATIONS[name]
|
||||
_settings(**changes)
|
||||
_long_story(client.adv_id, turns=120)
|
||||
system, story, report = _build(client, window)
|
||||
tokens = report["tokens"]
|
||||
budget = tokens["budget"]
|
||||
|
||||
assert tokens["safety_reserve"] == contextwindow.safety_reserve(budget)
|
||||
assert tokens["output_reserve"] == 500
|
||||
sent = _sent(system, story)
|
||||
assert sent + tokens["output_reserve"] + tokens["safety_reserve"] <= budget, (
|
||||
name, sent, tokens)
|
||||
# The history is what gave way, not the canon.
|
||||
assert "CANON-SENTINEL-RESERVE-5120" in system
|
||||
assert report["history"]["included"] < report["history"]["total"]
|
||||
|
||||
|
||||
def test_the_report_prices_the_text_the_provider_adds(client):
|
||||
"""The chat hint rides on every request and was never counted."""
|
||||
_, _, report = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
tokens = report["tokens"]
|
||||
assert tokens["transport"] >= builder.count_tokens(CHAT_CONTINUE_HINT)
|
||||
assert tokens["estimate"] == tokens["total"] + builder.count_tokens(CHAT_CONTINUE_HINT)
|
||||
|
||||
|
||||
def test_the_reserve_follows_the_effective_window_not_the_setting(client):
|
||||
"""5% of a 4,096 server, not 5% of a 16,384 setting it will never read."""
|
||||
_, _, capped = _build(client, contextwindow.Window(4096, contextwindow.LOADED))
|
||||
_, _, full = _build(client, contextwindow.Window(16384, contextwindow.LOADED))
|
||||
assert capped["tokens"]["safety_reserve"] == 256
|
||||
assert full["tokens"]["safety_reserve"] == 820
|
||||
|
||||
|
||||
def test_protected_context_that_only_fits_without_the_reserve_fails_explicitly(client):
|
||||
"""A1-3 in the builder. Before v1.1 this prompt would have been built.
|
||||
|
||||
The canon is sized so that protected text plus the reply fits a 4,096 window
|
||||
with room to spare, and does not fit once the 256-token reserve is taken.
|
||||
"""
|
||||
small = contextwindow.Window(4096, contextwindow.LOADED)
|
||||
# Measured with a window large enough never to overflow, because repeated
|
||||
# text merges tokens at its seams and cannot be priced by multiplication.
|
||||
roomy = contextwindow.Window(32768, contextwindow.LOADED)
|
||||
rules = None
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
settings = db.query(models.Settings).first()
|
||||
base_rules = list(adventure.campaign_canon["rules"])
|
||||
filler = "The bell tolls once for every name in the ledger."
|
||||
copies = 1
|
||||
while True:
|
||||
candidate = base_rules + [" ".join([filler] * copies)]
|
||||
adventure.campaign_canon = {"rules": candidate}
|
||||
_, _, measured = builder.build_context(adventure, settings, window=roomy)
|
||||
t = measured["tokens"]
|
||||
# What protected context costs at 4,096, without the reserve.
|
||||
without_reserve = t["protected"] + t["transport"] + t["output_reserve"]
|
||||
if without_reserve + 64 >= 4096 - 60:
|
||||
break
|
||||
copies += 1
|
||||
rules = candidate
|
||||
adventure.campaign_canon = {"rules": rules}
|
||||
db.commit()
|
||||
|
||||
# The case this test is about: v1's arithmetic, with its 64-token margin,
|
||||
# would have built this prompt. v1.1's reserve does not fit.
|
||||
assert without_reserve + 64 < 4096
|
||||
assert without_reserve + contextwindow.safety_reserve(4096) >= 4096
|
||||
|
||||
with pytest.raises(builder.ContextOverflow) as caught:
|
||||
builder.build_context(adventure, settings, window=small)
|
||||
message = str(caught.value)
|
||||
assert "safety" in message
|
||||
assert "load the model with a larger window" in message
|
||||
|
||||
|
||||
def test_an_overflowing_turn_never_reaches_the_model(client, monkeypatch):
|
||||
"""A1-3 end to end: the refusal happens before the provider is called."""
|
||||
async def verified(endpoint, model, declared=None, use_cache=True):
|
||||
return contextwindow.Window(1024, contextwindow.LOADED)
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
|
||||
ScriptedProvider.replies = ["This must never be generated."]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "open the crypt"})
|
||||
assert response.status_code == 200
|
||||
assert "safety" in response.text
|
||||
assert ScriptedProvider.calls == 0
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.Action).filter_by(
|
||||
adventure_id=client.adv_id, type="ai").count() == 0
|
||||
|
||||
|
||||
# ---------------------------------------------------------- the accounting
|
||||
|
||||
def _classify(prompt_tokens=None, *, usage=None, estimate=3500, budget=4096,
|
||||
output=500, verified=True):
|
||||
if usage is None and prompt_tokens is not None:
|
||||
usage = {"prompt_tokens": prompt_tokens, "completion_tokens": 40}
|
||||
return contextwindow.classify_usage(
|
||||
usage, estimate=estimate, budget=budget, max_output_tokens=output,
|
||||
window_verified=verified,
|
||||
)
|
||||
|
||||
|
||||
def test_a_prompt_the_server_read_in_full_fits():
|
||||
# The 13-token chat-template overhead measured against the real server.
|
||||
result = _classify(3513)
|
||||
assert result["status"] == contextwindow.FITS
|
||||
assert result["server_prompt_tokens"] == 3513
|
||||
assert result["difference"] == 13
|
||||
assert result["safety_reserve"] == 256
|
||||
assert result["observed_margin"] == 4096 - 500 - 3513
|
||||
|
||||
|
||||
def test_a_server_that_counts_more_than_the_reserve_allows_is_exceeded():
|
||||
"""The prompt plus the reply allocation no longer fits the window."""
|
||||
result = _classify(3700)
|
||||
assert result["status"] == contextwindow.EXCEEDED
|
||||
assert result["observed_margin"] < 0
|
||||
|
||||
|
||||
def test_a_server_that_read_far_less_than_was_sent_is_suspected_of_truncating():
|
||||
"""The real shape: 6,316 sent, 2,050 read, HTTP 200, no error."""
|
||||
result = _classify(2050, estimate=6316)
|
||||
assert result["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
assert result["difference"] == 2050 - 6316
|
||||
|
||||
|
||||
def test_a_small_undercount_is_tokenizer_drift_not_truncation():
|
||||
"""A tokenizer thriftier than `cl100k_base` reads fewer tokens honestly. Only
|
||||
a shortfall larger than the reserve is called truncation."""
|
||||
assert _classify(3500 - 255)["status"] == contextwindow.FITS
|
||||
assert _classify(3500 - 257)["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
|
||||
|
||||
@pytest.mark.parametrize("usage", [
|
||||
None,
|
||||
{},
|
||||
{"completion_tokens": 40},
|
||||
{"prompt_tokens": 0},
|
||||
{"prompt_tokens": "3500"},
|
||||
{"prompt_tokens": -1},
|
||||
])
|
||||
def test_no_usable_count_is_unknown_never_fits(usage):
|
||||
result = _classify(usage=usage)
|
||||
assert result["status"] == contextwindow.UNKNOWN
|
||||
assert result["server_prompt_tokens"] is None
|
||||
assert result["observed_margin"] is None
|
||||
|
||||
|
||||
def test_the_accounting_says_when_the_window_itself_was_not_verified():
|
||||
result = _classify(3513, verified=False)
|
||||
assert result["status"] == contextwindow.FITS
|
||||
assert result["window_verified"] is False
|
||||
assert "not verified" in result["detail"]
|
||||
|
||||
|
||||
def test_the_stream_asks_the_server_to_report_its_usage():
|
||||
"""Measured: Ollama 0.33 sends no usage in a stream unless asked."""
|
||||
provider = OpenAICompatibleProvider(ENDPOINT, "m")
|
||||
from app.providers.base import PromptParts
|
||||
|
||||
for mode in ("chat", "completion"):
|
||||
provider.api_mode = mode
|
||||
_url, body = provider._request(PromptParts(system="s", story="t"), 0.7, 50)
|
||||
assert body["stream"] is True
|
||||
assert body["stream_options"] == {"include_usage": True}
|
||||
|
||||
|
||||
def _latest_ai(adv_id):
|
||||
with SessionLocal() as db:
|
||||
return (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv_id, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
|
||||
|
||||
def _play(client, monkeypatch, usage, window=4096, reply="The crypt is still sealed."):
|
||||
async def verified(endpoint, model, declared=None, use_cache=True):
|
||||
return contextwindow.Window(window, contextwindow.LOADED, 32768, "fake")
|
||||
|
||||
monkeypatch.setattr(adventures.turns.contextwindow, "probe", verified)
|
||||
monkeypatch.setattr(ScriptedProvider, "last_usage", usage)
|
||||
ScriptedProvider.replies = [reply]
|
||||
return client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": "look at the seal"})
|
||||
|
||||
|
||||
def test_a_turn_records_what_the_server_read(client, monkeypatch):
|
||||
response = _play(client, monkeypatch, None)
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
|
||||
|
||||
response = _play(client, monkeypatch,
|
||||
{"prompt_tokens": estimate + 13, "completion_tokens": 9})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
snapshot = _latest_ai(client.adv_id).context_snapshot
|
||||
accounting = snapshot["accounting"]
|
||||
assert accounting["status"] == contextwindow.FITS
|
||||
assert accounting["server_prompt_tokens"] == estimate + 13
|
||||
assert accounting["estimate"] == snapshot["tokens"]["estimate"]
|
||||
assert '"accounting"' in response.text
|
||||
assert contextwindow.FITS in response.text
|
||||
|
||||
|
||||
def test_a_turn_with_no_reported_usage_is_unknown(client, monkeypatch):
|
||||
response = _play(client, monkeypatch, None)
|
||||
assert response.status_code == 200
|
||||
assert _latest_ai(client.adv_id).context_snapshot["accounting"]["status"] == (
|
||||
contextwindow.UNKNOWN)
|
||||
|
||||
|
||||
def test_a_suspected_truncation_keeps_the_turn_and_says_so(client, monkeypatch, caplog):
|
||||
"""A1-7 and A1-8. The reader watched the narration arrive; it stays."""
|
||||
response = _play(client, monkeypatch, {"prompt_tokens": 12, "completion_tokens": 9},
|
||||
reply="The seal holds, and the rain goes on.")
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
action = _latest_ai(client.adv_id)
|
||||
assert action is not None
|
||||
assert action.text == "The seal holds, and the rain goes on."
|
||||
accounting = action.context_snapshot["accounting"]
|
||||
assert accounting["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
assert contextwindow.TRUNCATION_SUSPECTED in response.text
|
||||
assert any(contextwindow.TRUNCATION_SUSPECTED in r.getMessage() for r in caplog.records)
|
||||
|
||||
# Inspectable afterwards through the same route the context panel reads.
|
||||
context = client.get(
|
||||
f"/api/adventures/{client.adv_id}/actions/{action.id}/context")
|
||||
assert context.status_code == 200
|
||||
assert context.json()["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
|
||||
|
||||
def test_each_attempt_keeps_its_own_accounting_when_the_live_flag_moves():
|
||||
"""Found by the A2 long run. Accounting belongs to one API call, not to the
|
||||
turn's shared prompt. A retry demotes the old attempt, and a take selection
|
||||
hands the prompt from one attempt to another. Neither may drop an attempt's
|
||||
accounting or give it another attempt's."""
|
||||
from app import attempts
|
||||
|
||||
class Node:
|
||||
def __init__(self, snapshot):
|
||||
self.context_snapshot = snapshot
|
||||
|
||||
shared = {"tokens": {"estimate": 3000}, "sections": [], "window": {"verified": True}}
|
||||
first = Node(shared | {"raw_output": "one", "usage": {"prompt_tokens": 3015},
|
||||
"accounting": {"status": contextwindow.FITS, "server_prompt_tokens": 3015}})
|
||||
second = Node({"raw_output": "two", "usage": {"prompt_tokens": 12},
|
||||
"accounting": {"status": contextwindow.TRUNCATION_SUSPECTED,
|
||||
"server_prompt_tokens": 12}})
|
||||
|
||||
# Superseded by a retry: the old attempt keeps only its own slices.
|
||||
attempts.keep_own_slices(Node(dict(first.context_snapshot)))
|
||||
demoted = Node(dict(first.context_snapshot))
|
||||
attempts.keep_own_slices(demoted)
|
||||
assert demoted.context_snapshot["accounting"]["server_prompt_tokens"] == 3015
|
||||
assert "tokens" not in demoted.context_snapshot
|
||||
|
||||
# The prompt moves to the second attempt; each keeps its own accounting.
|
||||
attempts.hand_over_the_prompt(first, second)
|
||||
assert second.context_snapshot["tokens"] == {"estimate": 3000}
|
||||
assert second.context_snapshot["accounting"]["status"] == contextwindow.TRUNCATION_SUSPECTED
|
||||
assert second.context_snapshot["accounting"]["server_prompt_tokens"] == 12
|
||||
assert first.context_snapshot["accounting"]["status"] == contextwindow.FITS
|
||||
assert "tokens" not in first.context_snapshot
|
||||
|
||||
|
||||
def test_a_retry_leaves_each_take_with_its_own_accounting(client, monkeypatch):
|
||||
"""End to end, through the real retry route. Before the fix the live take
|
||||
inherited the superseded take's accounting, so the inspector could show one
|
||||
call's server count as another's."""
|
||||
response = _play(client, monkeypatch, None, reply="The first take.")
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
estimate = _latest_ai(client.adv_id).context_snapshot["tokens"]["estimate"]
|
||||
# Replay the first take with a real count, so it has accounting of its own.
|
||||
with SessionLocal() as db:
|
||||
first = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == client.adv_id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot)).first())
|
||||
snapshot = dict(first.context_snapshot)
|
||||
snapshot["accounting"] = contextwindow.classify_usage(
|
||||
{"prompt_tokens": estimate + 15}, estimate=estimate, budget=4096,
|
||||
max_output_tokens=500, window_verified=True)
|
||||
first.context_snapshot = snapshot
|
||||
db.commit()
|
||||
first_id = first.id
|
||||
|
||||
monkeypatch.setattr(ScriptedProvider, "last_usage",
|
||||
{"prompt_tokens": 12, "completion_tokens": 9})
|
||||
ScriptedProvider.replies = ["The second take."]
|
||||
retried = client.post(f"/api/adventures/{client.adv_id}/retry")
|
||||
assert retried.status_code == 200, retried.text[:300]
|
||||
|
||||
with SessionLocal() as db:
|
||||
rows = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == client.adv_id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id).all())
|
||||
by_id = {row.id: row for row in rows}
|
||||
old = by_id[first_id]
|
||||
new = [row for row in rows if row.id != first_id][-1]
|
||||
assert new.text == "The second take."
|
||||
assert new.live and not old.live
|
||||
# The superseded take keeps its own accounting and gives up the prompt.
|
||||
assert old.context_snapshot["accounting"]["status"] == contextwindow.FITS
|
||||
assert old.context_snapshot["accounting"]["server_prompt_tokens"] == estimate + 15
|
||||
assert "tokens" not in old.context_snapshot
|
||||
# The live take carries the prompt and its own accounting, not the old one's.
|
||||
assert "tokens" in new.context_snapshot
|
||||
assert new.context_snapshot["accounting"]["status"] == (
|
||||
contextwindow.TRUNCATION_SUSPECTED)
|
||||
assert new.context_snapshot["accounting"]["server_prompt_tokens"] == 12
|
||||
|
||||
|
||||
def test_an_exceeded_turn_is_also_kept(client, monkeypatch):
|
||||
response = _play(client, monkeypatch, {"prompt_tokens": 3900, "completion_tokens": 9})
|
||||
assert response.status_code == 200
|
||||
action = _latest_ai(client.adv_id)
|
||||
assert action.text == "The crypt is still sealed."
|
||||
assert action.context_snapshot["accounting"]["status"] == contextwindow.EXCEEDED
|
||||
@@ -0,0 +1,294 @@
|
||||
"""v1.1 WP-D: a backup that was really checked, and an export that says what it is.
|
||||
|
||||
Two recovery-path claims, each of which was true only in the small before this
|
||||
package:
|
||||
|
||||
- **A backup is verified.** M9 ran `PRAGMA quick_check` on the finished copy.
|
||||
That reads every page and every record, and skips the cross-check between a
|
||||
table and its indexes — so a copy whose index disagrees with its table passed.
|
||||
`test_the_fixture_is_the_difference_between_the_two_checks` builds exactly that
|
||||
damage and shows the two pragmas disagreeing about it, before anything here
|
||||
uses it as evidence.
|
||||
- **An export is importable.** Nothing compared the bundle with
|
||||
`limits.MAX_IMPORT_BODY_BYTES`, so a campaign could be exported and then
|
||||
refused by its own importer, with the reader finding out at the moment they
|
||||
needed it. The export still succeeds — the file is complete, and a version
|
||||
that refused to write it would destroy the copy someone was trying to make —
|
||||
and now it says so.
|
||||
|
||||
python -m pytest tests/test_v11_d_recovery.py -v
|
||||
"""
|
||||
|
||||
import json
|
||||
import sqlite3
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, backup, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client():
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="wp-d@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(user_id=user.id, model="test-model"))
|
||||
adventure = models.Adventure(user_id=user.id, title="Recovery")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(adventure_id=adventure.id, type="start", text="The story opens."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id))
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
test_client.user_id = user_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
# ------------------------------------------------------------ the fixture
|
||||
|
||||
def build_corrupt_copy(path: Path) -> None:
|
||||
"""A database whose index disagrees with its table, and nothing else.
|
||||
|
||||
One digit inside one index leaf page is changed, so that entry names a key
|
||||
no row holds and one row's key is in no index entry. Every page is still
|
||||
structurally sound and every record still parses, which is the whole point:
|
||||
this is the damage `quick_check` is not looking for.
|
||||
"""
|
||||
path.unlink(missing_ok=True)
|
||||
connection = sqlite3.connect(path)
|
||||
connection.execute("PRAGMA page_size=4096")
|
||||
connection.execute("CREATE TABLE t (id INTEGER PRIMARY KEY, k TEXT NOT NULL, filler TEXT)")
|
||||
connection.execute("CREATE INDEX i_t_k ON t(k)")
|
||||
connection.executemany("INSERT INTO t (k, filler) VALUES (?, ?)",
|
||||
[(f"k{n:06d}", "x" * 40) for n in range(400)])
|
||||
connection.commit()
|
||||
page_size = connection.execute("PRAGMA page_size").fetchone()[0]
|
||||
leaves = [row[0] for row in connection.execute(
|
||||
"SELECT pageno FROM dbstat WHERE name='i_t_k' AND pagetype='leaf' ORDER BY pageno")]
|
||||
connection.close()
|
||||
assert leaves, "the index must have a leaf page to damage"
|
||||
|
||||
raw = bytearray(path.read_bytes())
|
||||
start = (leaves[0] - 1) * page_size
|
||||
page = raw[start:start + page_size]
|
||||
at = page.find(b"k000")
|
||||
assert at != -1, "expected an indexed key on the index's first leaf page"
|
||||
page[at + 4] = ord("9") # k000144 -> k000944: a key no row has
|
||||
raw[start:start + page_size] = page
|
||||
path.write_bytes(bytes(raw))
|
||||
|
||||
|
||||
def check(path: Path, pragma: str) -> str:
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
return ", ".join(str(row[0]) for row in connection.execute(f"PRAGMA {pragma}").fetchall())
|
||||
finally:
|
||||
connection.close()
|
||||
|
||||
|
||||
def test_the_fixture_is_the_difference_between_the_two_checks(tmp_path):
|
||||
"""Before using it as evidence: quick_check calls this database fine."""
|
||||
damaged = tmp_path / "damaged.db"
|
||||
build_corrupt_copy(damaged)
|
||||
assert check(damaged, "quick_check") == "ok"
|
||||
integrity = check(damaged, "integrity_check")
|
||||
assert integrity != "ok"
|
||||
assert "i_t_k" in integrity # it names the index that disagrees
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- backups
|
||||
|
||||
def test_a_healthy_backup_passes_the_full_check_and_is_kept(client, tmp_path):
|
||||
source = tmp_path / "campaign.db"
|
||||
source.write_bytes(Path(str(engine.url.database)).read_bytes())
|
||||
result = backup.create(source)
|
||||
assert result.integrity == "ok"
|
||||
assert result.path.exists() and result.bytes > 0
|
||||
assert check(result.path, "integrity_check") == "ok"
|
||||
assert result.path.parent == backup.directory(source)
|
||||
|
||||
|
||||
def test_the_backup_runs_the_full_check_not_the_quick_one(client, tmp_path, monkeypatch):
|
||||
"""The pragma itself, named. SQLite traces every statement it executes, so
|
||||
this reads what the backup actually asked the copy rather than inferring it."""
|
||||
asked: list[str] = []
|
||||
real_connect = sqlite3.connect
|
||||
|
||||
def tracing(*args, **kwargs):
|
||||
connection = real_connect(*args, **kwargs)
|
||||
connection.set_trace_callback(asked.append)
|
||||
return connection
|
||||
|
||||
monkeypatch.setattr(backup.sqlite3, "connect", tracing)
|
||||
source = tmp_path / "campaign.db"
|
||||
source.write_bytes(Path(str(engine.url.database)).read_bytes())
|
||||
backup.create(source).path.unlink()
|
||||
assert any("integrity_check" in sql for sql in asked), asked
|
||||
assert not any("quick_check" in sql for sql in asked), asked
|
||||
|
||||
|
||||
def test_a_copy_the_full_check_rejects_is_not_kept(client, tmp_path, monkeypatch):
|
||||
"""The copy is damaged after it is written and before it is verified, which
|
||||
is where a real page-level fault would appear: between the copy and the
|
||||
rename. Nothing wearing a backup's name may be left behind."""
|
||||
source = tmp_path / "campaign.db"
|
||||
source.write_bytes(Path(str(engine.url.database)).read_bytes())
|
||||
real_copy = backup._copy
|
||||
|
||||
def damage(source_path, working):
|
||||
pages = real_copy(source_path, working)
|
||||
build_corrupt_copy(working)
|
||||
return pages
|
||||
|
||||
monkeypatch.setattr(backup, "_copy", damage)
|
||||
with pytest.raises(backup.BackupError) as refused:
|
||||
backup.create(source)
|
||||
assert "did not verify" in str(refused.value)
|
||||
assert "i_t_k" in str(refused.value) # it says what was wrong
|
||||
kept = list(backup.directory(source).glob("*"))
|
||||
assert kept == [], f"a rejected backup was left behind: {kept}"
|
||||
|
||||
|
||||
def test_a_rejected_backup_leaves_an_earlier_good_one_alone(client, tmp_path, monkeypatch):
|
||||
source = tmp_path / "campaign.db"
|
||||
source.write_bytes(Path(str(engine.url.database)).read_bytes())
|
||||
good = backup.create(source)
|
||||
before = good.path.read_bytes()
|
||||
|
||||
real_copy = backup._copy
|
||||
|
||||
def damage(source_path, working):
|
||||
pages = real_copy(source_path, working)
|
||||
build_corrupt_copy(working)
|
||||
return pages
|
||||
|
||||
monkeypatch.setattr(backup, "_copy", damage)
|
||||
with pytest.raises(backup.BackupError):
|
||||
backup.create(source)
|
||||
assert good.path.exists()
|
||||
assert good.path.read_bytes() == before
|
||||
assert check(good.path, "integrity_check") == "ok"
|
||||
assert [p.name for p in backup.directory(source).glob("*")] == [good.path.name]
|
||||
|
||||
|
||||
def test_the_backup_file_semantics_are_unchanged(client, tmp_path):
|
||||
"""Same directory, same stamped name, same reported fields: WP-D changed the
|
||||
check, not the file."""
|
||||
source = tmp_path / "campaign.db"
|
||||
source.write_bytes(Path(str(engine.url.database)).read_bytes())
|
||||
first = backup.create(source)
|
||||
second = backup.create(source)
|
||||
assert first.path.name.startswith(backup.PREFIX) and first.path.suffix == ".db"
|
||||
assert first.path != second.path, "an existing backup is never overwritten"
|
||||
assert set(first.as_dict()) == {"filename", "bytes", "pages", "seconds", "integrity"}
|
||||
listed = [row["filename"] for row in backup.existing(source)]
|
||||
assert sorted(listed) == sorted([first.path.name, second.path.name])
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- exports
|
||||
|
||||
def export(client, adv_id):
|
||||
response = client.get(f"/api/adventures/{adv_id}/export")
|
||||
assert response.status_code == 200, response.text[:200]
|
||||
return response
|
||||
|
||||
|
||||
def test_a_normal_export_carries_no_warning(client):
|
||||
response = export(client, client.adv_id)
|
||||
assert "X-Export-Warning" not in response.headers
|
||||
assert response.headers["X-Importable-By-This-Version"] == "true"
|
||||
assert int(response.headers["X-Import-Limit-Bytes"]) == limits.MAX_IMPORT_BODY_BYTES
|
||||
assert int(response.headers["X-Export-Bytes"]) == len(response.content)
|
||||
assert response.json()["format"] == "ai-dnd-adventure-v3"
|
||||
|
||||
|
||||
def fill_past_the_limit(adv_id: int) -> int:
|
||||
"""Real rows, until the campaign's bundle is genuinely over the ceiling.
|
||||
|
||||
Not a mocked size: the export below serialises all of it.
|
||||
"""
|
||||
chunk = "The rain kept on over the harbour road, and nobody came. " * 900 # ~50 kB
|
||||
written = 0
|
||||
with SessionLocal() as db:
|
||||
while written < limits.MAX_IMPORT_BODY_BYTES + 2 * 1024 * 1024:
|
||||
db.add_all([models.Action(adventure_id=adv_id, type="ai", text=chunk)
|
||||
for _ in range(40)])
|
||||
db.commit()
|
||||
written += 40 * len(chunk)
|
||||
return written
|
||||
|
||||
|
||||
def test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back(client):
|
||||
fill_past_the_limit(client.adv_id)
|
||||
response = export(client, client.adv_id)
|
||||
|
||||
# Delivered, whole, and still the same format.
|
||||
body = response.content
|
||||
assert len(body) > limits.MAX_IMPORT_BODY_BYTES
|
||||
parsed = json.loads(body)
|
||||
assert parsed["format"] == "ai-dnd-adventure-v3"
|
||||
assert len(parsed["actions"]) > 40
|
||||
|
||||
# And honest about what this version can do with it.
|
||||
assert response.headers["X-Importable-By-This-Version"] == "false"
|
||||
warning = response.headers["X-Export-Warning"]
|
||||
assert limits.import_limit_label() in warning
|
||||
assert "exported successfully" in warning
|
||||
assert "cannot import" in warning
|
||||
assert int(response.headers["X-Export-Bytes"]) == len(body)
|
||||
|
||||
|
||||
def test_the_warning_follows_the_configured_limit(monkeypatch):
|
||||
"""The text is generated from the constant, so changing the constant changes
|
||||
the sentence rather than leaving a stale number in it."""
|
||||
assert "20 MB" in limits.oversized_export_warning(21_000_000)
|
||||
monkeypatch.setattr(limits, "MAX_IMPORT_BODY_BYTES", 50 * 1024 * 1024)
|
||||
assert limits.import_limit_label() == "50 MB"
|
||||
assert "50 MB" in limits.oversized_export_warning(60_000_000)
|
||||
assert "20 MB" not in limits.oversized_export_warning(60_000_000)
|
||||
|
||||
|
||||
def test_the_bundle_itself_never_carries_the_warning(client):
|
||||
"""The warning is about the export, not part of the portable story file."""
|
||||
fill_past_the_limit(client.adv_id)
|
||||
parsed = json.loads(export(client, client.adv_id).content)
|
||||
flat = json.dumps(parsed).lower()
|
||||
assert "import limit" not in flat
|
||||
assert "cannot import" not in flat
|
||||
for key in parsed:
|
||||
assert "warning" not in key.lower()
|
||||
|
||||
|
||||
def test_that_same_bundle_is_refused_by_import_naming_the_limit(client):
|
||||
fill_past_the_limit(client.adv_id)
|
||||
body = export(client, client.adv_id).content
|
||||
response = client.post("/api/adventures/import", content=body,
|
||||
headers={"Content-Type": "application/json"})
|
||||
assert response.status_code == 413
|
||||
detail = response.json()["detail"]
|
||||
assert "too large" in detail.lower()
|
||||
assert limits.import_limit_label() in detail
|
||||
|
||||
|
||||
def test_a_bundle_under_the_limit_still_imports(client):
|
||||
"""The refusal is about size alone: the ordinary path is untouched."""
|
||||
body = export(client, client.adv_id).content
|
||||
assert len(body) < limits.MAX_IMPORT_BODY_BYTES
|
||||
response = client.post("/api/adventures/import", content=body,
|
||||
headers={"Content-Type": "application/json"})
|
||||
assert response.status_code == 201, response.text[:200]
|
||||
@@ -0,0 +1,152 @@
|
||||
"""v1.1 WP-E: the contrast audit is a gate, not a report.
|
||||
|
||||
Before this package `tools/contrast_audit.py` measured control boundaries,
|
||||
printed that two of them were below 3:1, and exited 0 anyway — on the argument
|
||||
that a control is identified by its label rather than its edge. A check that
|
||||
cannot fail is not a check, and these tests are what make it one: the threshold
|
||||
is exercised from both sides, on a real tokens file, so a future palette change
|
||||
that dims a control's edge stops the run instead of adding a line to it.
|
||||
|
||||
python -m pytest tests/test_v11_e_contrast.py -v
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tools"))
|
||||
|
||||
import contrast_audit as audit # noqa: E402
|
||||
|
||||
|
||||
# --------------------------------------------------------------- the maths
|
||||
|
||||
def test_the_ratio_is_the_wcag_ratio():
|
||||
"""Anchored on values with known answers, so a broken formula is visible."""
|
||||
assert audit.ratio("#ffffff", "#000000") == pytest.approx(21.0, abs=0.01)
|
||||
assert audit.ratio("#ffffff", "#ffffff") == pytest.approx(1.0, abs=0.001)
|
||||
# Order does not matter: contrast is symmetric.
|
||||
assert audit.ratio("#131320", "#676792") == pytest.approx(
|
||||
audit.ratio("#676792", "#131320"), abs=1e-9)
|
||||
|
||||
|
||||
# ------------------------------------------------------- the gate itself
|
||||
|
||||
def tokens_file(tmp_path: Path, **overrides: str) -> Path:
|
||||
"""A real tokens file with named colours replaced."""
|
||||
source = audit.TOKENS.read_text()
|
||||
for name, value in overrides.items():
|
||||
token = "--" + name.replace("_", "-")
|
||||
start = source.index(f"{token}: ")
|
||||
end = source.index(";", start)
|
||||
source = source[:start] + f"{token}: {value}" + source[end:]
|
||||
written = tmp_path / "tokens.css"
|
||||
written.write_text(source)
|
||||
return written
|
||||
|
||||
|
||||
def run_against(path: Path, monkeypatch) -> int:
|
||||
monkeypatch.setattr(audit, "TOKENS", path)
|
||||
return audit.main()
|
||||
|
||||
|
||||
def test_a_boundary_below_three_to_one_fails_the_run(tmp_path, monkeypatch, capsys):
|
||||
"""2.99:1 against --bg-input — the wrong side of the line by one hundredth.
|
||||
|
||||
This value clears 3:1 against --bg-panel (3.21:1), so it would have passed
|
||||
the audit as M11 wrote it. It fails now because the floor is taken against
|
||||
the background the control is actually drawn on.
|
||||
"""
|
||||
below = tokens_file(tmp_path, border="#58639a")
|
||||
assert run_against(below, monkeypatch) == 1
|
||||
printed = capsys.readouterr().out
|
||||
assert "2.99:1" in printed
|
||||
assert "FAIL" in printed
|
||||
assert "below 3:1 (WCAG 1.4.11)" in printed
|
||||
|
||||
|
||||
def test_a_boundary_at_three_to_one_passes(tmp_path, monkeypatch, capsys):
|
||||
"""3.00:1 — the right side of the same line, one hundredth from the value
|
||||
above, and not failed for arithmetic the reader cannot see."""
|
||||
at = tokens_file(tmp_path, border="#58639b")
|
||||
assert run_against(at, monkeypatch) == 0
|
||||
printed = capsys.readouterr().out
|
||||
assert "3.00:1" in printed
|
||||
assert "every control boundary clears 3:1" in printed
|
||||
|
||||
|
||||
def test_the_v1_0_0_boundary_would_now_fail(tmp_path, monkeypatch, capsys):
|
||||
"""The value v1.0.0 shipped. This is the defect WP-E closes, and the gate
|
||||
has to be the thing that would have caught it."""
|
||||
shipped = tokens_file(tmp_path, border="#2b2b3d")
|
||||
assert run_against(shipped, monkeypatch) == 1
|
||||
assert "1.24:1" in capsys.readouterr().out # against --bg-input, the worst case
|
||||
|
||||
|
||||
def test_a_text_pair_below_four_point_five_still_fails(tmp_path, monkeypatch, capsys):
|
||||
"""WP-E raised the boundary floor and must not have lowered the text one."""
|
||||
dimmed = tokens_file(tmp_path, text_dim="#5a5750")
|
||||
assert run_against(dimmed, monkeypatch) == 1
|
||||
assert "below WCAG AA (1.4.3)" in capsys.readouterr().out
|
||||
|
||||
|
||||
def test_a_missing_token_is_a_failure_not_a_skip(tmp_path, monkeypatch, capsys):
|
||||
source = audit.TOKENS.read_text().replace("--border-bright:", "--border-was-renamed:")
|
||||
written = tmp_path / "tokens.css"
|
||||
written.write_text(source)
|
||||
assert run_against(written, monkeypatch) == 1
|
||||
assert "MISSING TOKEN" in capsys.readouterr().out
|
||||
|
||||
|
||||
# ------------------------------------------------- the shipped palette
|
||||
|
||||
def test_the_real_tokens_pass_both_criteria(capsys):
|
||||
"""The palette as it stands, through the same gate CI would run."""
|
||||
assert audit.main() == 0
|
||||
printed = capsys.readouterr().out
|
||||
assert "every text pair clears WCAG AA (1.4.3)" in printed
|
||||
assert "every control boundary clears 3:1 (1.4.11)" in printed
|
||||
assert "FAIL" not in printed
|
||||
|
||||
|
||||
def test_every_boundary_pair_is_measured_against_the_background_it_is_drawn_on():
|
||||
"""The audit used to check borders only against --bg-panel, which is not
|
||||
where the bordered controls are: inputs and buttons sit on --bg-input
|
||||
(styles/forms.css), which is lighter and therefore harder. Checking only the
|
||||
easier background would let a token pass while the real control failed."""
|
||||
boundary = [(fg, bg) for kind, fg, bg, _, _ in audit.PAIRS if kind == "boundary"]
|
||||
for background in ("--bg-input", "--bg-panel", "--bg"):
|
||||
assert ("--border", background) in boundary, background
|
||||
assert ("--border-bright", "--bg-input") in boundary
|
||||
|
||||
|
||||
def test_the_text_baselines_are_unchanged_by_wp_e():
|
||||
"""WP-E changed only boundary tokens. These are the M11 text measurements,
|
||||
and they have to still be exactly what the earlier reports recorded."""
|
||||
tokens = audit.read_tokens(audit.TOKENS)
|
||||
measured = {
|
||||
"body text on the page": audit.ratio(tokens["--text"], tokens["--bg"]),
|
||||
"body text in a panel": audit.ratio(tokens["--text"], tokens["--bg-panel"]),
|
||||
"secondary text in a panel": audit.ratio(tokens["--text-dim"], tokens["--bg-panel"]),
|
||||
"secondary text on the page": audit.ratio(tokens["--text-dim"], tokens["--bg"]),
|
||||
}
|
||||
assert measured["body text on the page"] == pytest.approx(14.57, abs=0.01)
|
||||
assert measured["body text in a panel"] == pytest.approx(13.57, abs=0.01)
|
||||
assert measured["secondary text in a panel"] == pytest.approx(5.48, abs=0.01)
|
||||
assert measured["secondary text on the page"] == pytest.approx(5.88, abs=0.01)
|
||||
|
||||
|
||||
def test_the_hover_edge_stays_brighter_than_the_resting_edge():
|
||||
"""Rest and hover have to remain distinguishable from each other, not merely
|
||||
each clear the floor against the background."""
|
||||
tokens = audit.read_tokens(audit.TOKENS)
|
||||
panel = tokens["--bg-panel"]
|
||||
assert audit.ratio(tokens["--border-bright"], panel) > audit.ratio(tokens["--border"], panel)
|
||||
|
||||
|
||||
def test_a_boundary_never_becomes_as_loud_as_body_text():
|
||||
"""A control's edge that outshines the words inside it is its own defect."""
|
||||
tokens = audit.read_tokens(audit.TOKENS)
|
||||
panel = tokens["--bg-panel"]
|
||||
assert audit.ratio(tokens["--border-bright"], panel) < audit.ratio(tokens["--text"], panel)
|
||||
@@ -0,0 +1,324 @@
|
||||
"""v1.1 WP-A2: the protocol a narrator copies stays out of the story, and nothing else does.
|
||||
|
||||
The M11 closeout's identity run (an office meeting, a 3B narrator, a 4,096
|
||||
window) stored four turns carrying text the application wrote, not the story:
|
||||
|
||||
- `> Create_entity(new_person, "john", …` — the event vocabulary as the prompt
|
||||
printed it, `name(field, …)`, copied as if it were a call;
|
||||
- `> set_possession(silver-key, "alice") Adds the silver key to Alice's
|
||||
possession.` — the call again, naming the fantasy example slug from the fixed
|
||||
state rule, in a meeting room;
|
||||
- `Scene: Bill, Alice, … (at The meeting room)` — the renderer's own scene line;
|
||||
- `[Hard limit: your next turn must not exceed 180 words, … append the state
|
||||
block well inside the limit.]` — the length hint, reworded at the front and
|
||||
verbatim at the end.
|
||||
|
||||
v1.0.0 removed none of them. The rule this module is held to is unchanged from
|
||||
M5: **removing story is worse than leaving protocol.** Every removal below is
|
||||
anchored to a string or a vocabulary the application owns, and every one has
|
||||
story beside it that must survive.
|
||||
|
||||
python -m pytest tests/test_v11_protocol_echo.py -v
|
||||
"""
|
||||
|
||||
import json
|
||||
import re
|
||||
|
||||
import pytest
|
||||
|
||||
from app.context import builder
|
||||
from app.narrative import events, extract, render
|
||||
|
||||
# ---------------------------------------------------------------- the prompt
|
||||
|
||||
#: The identifiers the v1 state rule taught every campaign, from the fantasy
|
||||
#: acceptance fixture. None may come back into a fixed instruction.
|
||||
FANTASY_IDENTIFIERS = ("mara", "silver-key", "silver key", "old-abbey", "abbey",
|
||||
"aldric", "westhaven", "crypt", "edrin")
|
||||
#: And nothing from the science-fiction fixture either: neutral means neutral,
|
||||
#: not "the other genre".
|
||||
SCIFI_IDENTIFIERS = ("persephone", "imani", "data-crystal", "data crystal", "airlock")
|
||||
|
||||
|
||||
def _fixed_instructions() -> str:
|
||||
return "\n".join([
|
||||
extract.EMIT_RULE,
|
||||
extract.EMIT_REMINDER,
|
||||
events.vocabulary_for_prompt(),
|
||||
builder.length_hint(500),
|
||||
builder.length_hint(500, "brief"),
|
||||
builder.length_hint(500, "long"),
|
||||
builder.length_hint(120),
|
||||
]).lower()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("identifier", FANTASY_IDENTIFIERS + SCIFI_IDENTIFIERS)
|
||||
def test_the_fixed_state_instructions_name_no_fixture_identifier(identifier):
|
||||
"""A2-1. The example slug the office run copied cannot come back."""
|
||||
assert not re.search(rf"\b{re.escape(identifier)}\b", _fixed_instructions())
|
||||
|
||||
|
||||
def test_the_example_uses_neutral_identifiers():
|
||||
for neutral in ("character-1", "item-1", "location-1"):
|
||||
assert neutral in extract.EMIT_RULE
|
||||
|
||||
|
||||
def test_the_worked_example_is_a_block_this_extractor_accepts():
|
||||
"""The example is the wire format, byte for byte, not an illustration of it."""
|
||||
example = extract.EMIT_RULE[extract.EMIT_RULE.index("```state"):]
|
||||
prose, parsed, _raw = extract.split("The door opens.\n\n" + example)
|
||||
assert prose == "The door opens."
|
||||
assert isinstance(parsed, dict)
|
||||
assert [e["type"] for e in parsed["events"]] == [
|
||||
"set_possession", "set_current_location"]
|
||||
for event in parsed["events"]:
|
||||
assert events.is_allowed(event["type"])
|
||||
|
||||
|
||||
def test_the_vocabulary_is_not_written_as_function_calls():
|
||||
"""A2-2. `set_possession(item, owner)` is the notation the narrator copied."""
|
||||
vocabulary = events.vocabulary_for_prompt()
|
||||
for name in events.SPECS:
|
||||
assert not re.search(rf"\b{name}\s*\(", vocabulary), name
|
||||
|
||||
|
||||
def test_every_event_is_described_in_the_shape_the_model_must_send():
|
||||
lines = events.vocabulary_for_prompt().splitlines()
|
||||
assert len(lines) == len(events.SPECS)
|
||||
for name, line in zip(events.SPECS, lines):
|
||||
shape = line.strip().split(" — ", 1)[0]
|
||||
obj = json.loads(shape)
|
||||
assert obj["type"] == name
|
||||
assert set(obj) - {"type"} == set(events.SPECS[name]["required"])
|
||||
for optional in events.SPECS[name]["optional"]:
|
||||
assert optional in line
|
||||
|
||||
|
||||
def test_the_length_hint_carries_the_phrases_the_extractor_recognises():
|
||||
"""One source for the words, so the builder and the extractor cannot drift."""
|
||||
for narration_length in ("", "brief", "medium", "long"):
|
||||
hint = builder.length_hint(500, narration_length)
|
||||
assert hint.startswith(extract.LENGTH_HINT_OPENING)
|
||||
assert extract.LENGTH_HINT_TAIL in hint
|
||||
|
||||
|
||||
# ---------------------------------------------------------- observed shapes
|
||||
|
||||
STORY = (
|
||||
"Alice looks at John, the tension in the room palpable.\n\n"
|
||||
"John nods. \"I'm ready to contribute.\""
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("leak", [
|
||||
# Depth 20: the call, the fantasy slug, and a gloss on the same line.
|
||||
'> set_possession(silver-key, "alice") Adds the silver key to Alice\'s possession.',
|
||||
# Depth 10: cut off by the output limit mid-call.
|
||||
'> Create_entity(new_person, "john", "character", "A determined team member", ["john',
|
||||
# Unquoted, and a vocabulary name in any case.
|
||||
'SET_CURRENT_LOCATION(bill, office)',
|
||||
'add_fact(predicate="knows the plan", subject="alice")',
|
||||
])
|
||||
def test_an_event_call_line_at_the_end_leaves_the_story(leak):
|
||||
prose, parsed, _raw = extract.split(f"{STORY}\n\n{leak}")
|
||||
assert prose == STORY
|
||||
assert parsed is None
|
||||
|
||||
|
||||
def test_an_event_call_line_in_the_middle_leaves_and_the_story_after_it_stays():
|
||||
"""Depth 12: the call, then more narration."""
|
||||
reply = (
|
||||
f"{STORY}\n\n"
|
||||
'> Create_entity(new_person, "mike", "character", "A new team member.", ["mike"])\n\n'
|
||||
"Mike takes the empty chair by the window."
|
||||
)
|
||||
prose, _parsed, _raw = extract.split(reply)
|
||||
assert prose == f"{STORY}\n\nMike takes the empty chair by the window."
|
||||
|
||||
|
||||
def test_the_depth_fourteen_tail_leaves_entirely():
|
||||
"""A call, a rendered scene line, and a reworded length hint, in that order."""
|
||||
reply = (
|
||||
f"{STORY}\n\n"
|
||||
'> Create_entity(mike, "character", "A new team member.", ["mike"])\n\n'
|
||||
"Scene: Bill, Alice, Roger, John, and Mike at the table. (at The meeting room)\n\n"
|
||||
"[Hard limit: your next turn must not exceed 180 words, and it should not stop "
|
||||
"short of about 70. Prefer the lower end of that range unless the scene genuinely "
|
||||
"needs more. Finish the narration and append the state block well inside the limit.]"
|
||||
)
|
||||
prose, parsed, _raw = extract.split(reply)
|
||||
assert prose == STORY
|
||||
assert parsed is None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("hint", [
|
||||
builder.length_hint(500),
|
||||
builder.length_hint(500, "brief"),
|
||||
# Cut off by the output limit before the tail.
|
||||
"[Hard limit: this turn must not exceed 180 words, and it should not stop short",
|
||||
# Reworded at the front, as the 3B narrator did.
|
||||
"[Hard limit: your next turn must not exceed 506 words. Write only as much as the "
|
||||
"moment needs — a typical turn is much shorter. Finish the narration and append "
|
||||
"the state block well inside the limit.]",
|
||||
])
|
||||
def test_a_parroted_length_hint_at_the_end_leaves_the_story(hint):
|
||||
prose, _parsed, _raw = extract.split(f"{STORY}\n\n{hint}")
|
||||
assert prose == STORY
|
||||
|
||||
|
||||
def test_a_rendered_scene_line_at_the_end_leaves_the_story():
|
||||
prose, _parsed, _raw = extract.split(
|
||||
f"{STORY}\n\nScene: A tense budget meeting. (at The meeting room)")
|
||||
assert prose == STORY
|
||||
|
||||
|
||||
def test_a_fenced_block_with_a_call_line_above_it_still_parses_and_applies():
|
||||
"""A2-6. The proposal is still read when protocol litter surrounds it."""
|
||||
reply = (
|
||||
f"{STORY}\n\n"
|
||||
'> set_current_location(john, office)\n\n'
|
||||
'```state\n{"events": [{"type": "set_current_location", '
|
||||
'"entity": "john", "location": "office"}]}\n```'
|
||||
)
|
||||
prose, parsed, raw = extract.split(reply)
|
||||
assert prose == STORY
|
||||
assert parsed["events"][0]["entity"] == "john"
|
||||
assert raw.startswith("{")
|
||||
|
||||
|
||||
# ------------------------------------------------ adversarial story that stays
|
||||
|
||||
@pytest.mark.parametrize("reply", [
|
||||
# The owner's cases.
|
||||
'The engineer writes "set_power(core, 80)" on the whiteboard.',
|
||||
'She says, "Create_entity is a terrible name for a company."',
|
||||
'The old manual contains a heading labeled "Scene:"',
|
||||
'He reads aloud: "[Hard limit: 500 words]" and laughs.',
|
||||
# A vocabulary name, written into a story, not at the start of a line.
|
||||
'Nadia squints at the log: the last command was set_possession(badge, guard).',
|
||||
# Call-shaped, at the start of a line, but not an event this protocol has.
|
||||
"The terminal scrolls.\n\n> open_door(north)\n\nNothing happens.",
|
||||
# A vocabulary call inside the story's own code block is the story's code.
|
||||
"She types:\n\n```python\ncreate_entity(ship)\nset_possession(key, captain)\n```\n\n"
|
||||
"The console beeps twice.",
|
||||
# A bracket at the very end, in-world, that is not the application's hint.
|
||||
"The warning light blinks.\n\n[Hard limit of the reactor: three hours]",
|
||||
"The contract ends with a clause.\n\n[Hard limit: forty days, no extensions]",
|
||||
# A scene heading in a screenplay the characters are writing, mid-story.
|
||||
"Scene: a kitchen, late.\n\nShe crosses it out and starts again.",
|
||||
# A last line that starts like the renderer's but is not its shape.
|
||||
"The director calls it.\n\nScene: take two, and nobody moves.",
|
||||
# A fact restated inside a sentence.
|
||||
"Alice knew the badge opened the server room, and said nothing.",
|
||||
"Memory: she remembered the bells.",
|
||||
])
|
||||
def test_story_that_resembles_the_new_rules_is_kept(reply):
|
||||
"""A2-5."""
|
||||
prose, parsed, _raw = extract.split(reply)
|
||||
assert prose == reply
|
||||
assert parsed is None
|
||||
|
||||
|
||||
# ------------------------------------------------------ the replay attribution
|
||||
|
||||
@pytest.mark.parametrize("line, rule", [
|
||||
('> set_possession(silver-key, "alice") Adds the key.', extract.RULE_EVENT_CALL),
|
||||
("Create_entity(new_person", extract.RULE_EVENT_CALL),
|
||||
("[Hard limit: this turn must not exceed 90 words. Finish the narration and append "
|
||||
"the state block well inside the limit.]", extract.RULE_LENGTH_HINT),
|
||||
("Scene: A meeting. (at The meeting room)", extract.RULE_SCENE_LINE),
|
||||
("John nods.", None),
|
||||
('He reads aloud: "[Hard limit: 500 words]" and laughs.', None),
|
||||
])
|
||||
def test_a_removed_line_is_attributed_to_the_rule_that_removes_it(line, rule):
|
||||
assert extract.explain_removed_line(line) == rule
|
||||
|
||||
|
||||
# ------------------------------------- corrective: the depth-16 instruction tail
|
||||
|
||||
#: Cut down from the v1.1 identity diagnostic's depth-16 turn, whose stored text
|
||||
#: was exactly the extractor's output. The two story paragraphs are shortened;
|
||||
#: the four trailing lines are verbatim.
|
||||
DEPTH_16_STORY = (
|
||||
"John's initial ideas are thoughtful and insightful, and the room fills with a "
|
||||
"sense of optimism.\n\n"
|
||||
"John's enthusiasm is contagious, and the meeting room is electric with the "
|
||||
"excitement of a fruitful collaboration ahead."
|
||||
)
|
||||
DEPTH_16_TAIL = (
|
||||
"Scene: Bill, Alice and Roger at the table; John not yet arrived.\n\n"
|
||||
"[Hard limit: this is now 180 words.]\n\n"
|
||||
"[Reminder: end your reply with a `state` block listing the events your narration "
|
||||
"made true, with absolute values.]\n\n"
|
||||
"[You don't need to continue; your turn must now be about John entering the room. "
|
||||
"Continue the story here, directly. Output only story text.]"
|
||||
)
|
||||
|
||||
|
||||
def test_the_depth_sixteen_instruction_tail_leaves_entirely():
|
||||
"""The corrective's positive regression. v1.1's first A2 left all four lines:
|
||||
the last bracket was a reworded continue hint nothing recognised, so nothing
|
||||
above it was ever at the end."""
|
||||
prose, parsed, _raw = extract.split(f"{DEPTH_16_STORY}\n\n{DEPTH_16_TAIL}")
|
||||
assert prose == DEPTH_16_STORY
|
||||
assert parsed is None
|
||||
|
||||
|
||||
def test_the_continue_hint_phrase_is_the_providers_own_sentence():
|
||||
from app.providers.openai_compatible import CHAT_CONTINUE_HINT
|
||||
|
||||
assert extract.CONTINUE_HINT_PHRASE in CHAT_CONTINUE_HINT
|
||||
|
||||
|
||||
def test_an_echoed_continue_hint_alone_at_the_end_leaves():
|
||||
prose, _p, _r = extract.split(
|
||||
f"{STORY}\n\n[Keep going. Continue the story here, directly. Output only story text.]")
|
||||
assert prose == STORY
|
||||
|
||||
|
||||
@pytest.mark.parametrize("reply", [
|
||||
# A hint-opened bracket with no echoed instruction below it is in-world.
|
||||
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
|
||||
# Nor does a state block below it make it an instruction.
|
||||
f"{STORY}\n\n[Hard limit: forty days, no extensions]",
|
||||
# The phrase in the middle of a story is prose, not a trailing echo.
|
||||
'She wrote "output only story text" on the card, then crossed it out.\n\nThe rain went on.',
|
||||
# A trailing in-world bracket that only resembles a continuation.
|
||||
f"{STORY}\n\n[To be continued]",
|
||||
])
|
||||
def test_story_brackets_near_the_corrective_rule_are_kept(reply):
|
||||
prose, _p, _r = extract.split(reply)
|
||||
assert prose == reply
|
||||
|
||||
|
||||
def test_a_hint_opened_bracket_above_a_state_block_is_kept():
|
||||
reply = (f"{STORY}\n\n[Hard limit: forty days, no extensions]\n\n"
|
||||
'```state\n{"events": []}\n```')
|
||||
prose, parsed, _r = extract.split(reply)
|
||||
assert prose == f"{STORY}\n\n[Hard limit: forty days, no extensions]"
|
||||
assert parsed == {"events": []}
|
||||
|
||||
|
||||
def test_an_in_world_bracket_above_an_echoed_hint_is_kept():
|
||||
"""Only a bracket opening the way the application's hint opens is taken with
|
||||
the echo. Any other bracket above it is the story's."""
|
||||
reply = (f"{STORY}\n\n[The sign on the door reads: Closed]\n\n"
|
||||
"[Continue the story here, directly. Output only story text.]")
|
||||
prose, _p, _r = extract.split(reply)
|
||||
assert prose == f"{STORY}\n\n[The sign on the door reads: Closed]"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("line, rule", [
|
||||
("[Hard limit: this is now 180 words.]", extract.RULE_INSTRUCTION_TAIL),
|
||||
("[Reminder: end your reply with a `state` block listing the events.]",
|
||||
extract.RULE_INSTRUCTION_TAIL),
|
||||
("[You don't need to continue. Output only story text.]", extract.RULE_INSTRUCTION_TAIL),
|
||||
])
|
||||
def test_the_corrective_rule_is_attributed(line, rule):
|
||||
assert extract.explain_removed_line(line) == rule
|
||||
|
||||
|
||||
def test_the_new_rules_do_not_disturb_the_section_headings_they_share_a_module_with():
|
||||
"""The renderer's headings are the M11 rules' anchor. A2 adds none."""
|
||||
assert render.HEADING_SCENE == "Scene:"
|
||||
assert "Scene:" not in render.SECTION_HEADINGS
|
||||
@@ -26,16 +26,24 @@ TOKENS = Path(__file__).resolve().parent.parent.parent / "frontend/src/styles/to
|
||||
#: rather than combinatorial, because "every colour against every other" reports
|
||||
#: pairs that never meet on screen.
|
||||
#:
|
||||
#: The `kind` matters and is not a way of grading on a curve. **text** pairs are
|
||||
#: WCAG 1.4.3 Contrast (Minimum) and are what §21 of the M11 brief asks about;
|
||||
#: they are pass/fail. **boundary** pairs are WCAG 1.4.11 Non-text Contrast,
|
||||
#: which applies to "visual information required to identify user interface
|
||||
#: components" — and in this design a control is identified by its *label*,
|
||||
#: which is measured above and passes, not by its edge. So a boundary below 3:1
|
||||
#: is reported with its number and does not fail the run; what it would take to
|
||||
#: turn it into a real failure is a control with no visible label, and there is
|
||||
#: no such control (`tools/m11_browser.py` asserts every visible control has an
|
||||
#: accessible name, and the story controls are text buttons).
|
||||
#: The `kind` records which success criterion a pair is measured against —
|
||||
#: **text** is WCAG 1.4.3 Contrast (Minimum), **boundary** is WCAG 1.4.11
|
||||
#: Non-text Contrast — and **both are pass/fail**.
|
||||
#:
|
||||
#: v1.1 WP-E overturned the earlier position here, which was that a boundary
|
||||
#: below 3:1 could be recorded rather than failed because "a control is
|
||||
#: identified by its label, not by its edge". That argument understates what
|
||||
#: 1.4.11 asks: the criterion covers the visual information needed to identify
|
||||
#: a component *and its boundary*, and a reader who cannot see where a text box
|
||||
#: ends cannot see that there is a text box to type into, label or no label.
|
||||
#: The edges were at 1.33:1 and 1.75:1 — the tokens were raised instead.
|
||||
#:
|
||||
#: Borders are measured against **every background they are drawn on**, and the
|
||||
#: floor is the worst of them. Inputs and buttons sit on --bg-input, which is
|
||||
#: lighter than --bg-panel and so the harder case; checking only --bg-panel
|
||||
#: would have let a token pass the audit while the real control failed.
|
||||
#: `--bg-panel-glass` is translucent and cannot be resolved from tokens alone;
|
||||
#: that edge is measured on the rendered page by `tools/m11_browser.py`.
|
||||
PAIRS = [
|
||||
("text", "--text", "--bg", 4.5, "body text on the page"),
|
||||
("text", "--text", "--bg-panel", 4.5, "body text in a panel"),
|
||||
@@ -48,8 +56,14 @@ PAIRS = [
|
||||
("text", "--danger", "--bg-panel", 4.5, "an error message"),
|
||||
("text", "--warning", "--bg-panel", 4.5, "a caution message"),
|
||||
("text", "--player", "--bg", 4.5, "the player's own words"),
|
||||
("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge"),
|
||||
("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge"),
|
||||
("boundary", "--border", "--bg-input", 3.0, "a field or button's resting edge"),
|
||||
("boundary", "--border-bright", "--bg-input", 3.0, "a field or button's hover edge"),
|
||||
("boundary", "--border", "--bg-panel", 3.0, "a control's resting edge in a panel"),
|
||||
("boundary", "--border-bright", "--bg-panel", 3.0, "a control's hover edge in a panel"),
|
||||
("boundary", "--border", "--bg", 3.0, "a divider on the page"),
|
||||
("boundary", "--border-bright", "--bg", 3.0, "the scrollbar thumb"),
|
||||
("boundary", "--accent-dim", "--bg-input", 3.0, "a focused field's edge"),
|
||||
("boundary", "--accent-dim", "--bg-panel", 3.0, "a focused control's edge in a panel"),
|
||||
("boundary", "--chart-1", "--bg-panel", 3.0, "a chart bar"),
|
||||
("boundary", "--chart-2", "--bg-panel", 3.0, "a chart bar"),
|
||||
("boundary", "--chart-3", "--bg-panel", 3.0, "a chart bar"),
|
||||
@@ -81,37 +95,44 @@ def ratio(a: str, b: str) -> float:
|
||||
|
||||
def main() -> int:
|
||||
tokens = read_tokens(TOKENS)
|
||||
print(f"{TOKENS.relative_to(TOKENS.parents[3])}: {len(tokens)} colour tokens\n")
|
||||
# Named defensively: the tests run this against a temporary tokens file,
|
||||
# which need not sit four directories deep the way the real one does.
|
||||
label = TOKENS.name
|
||||
if len(TOKENS.parents) > 3:
|
||||
label = TOKENS.relative_to(TOKENS.parents[3])
|
||||
print(f"{label}: {len(tokens)} colour tokens\n")
|
||||
print(f"{'pair':44} {'kind':9} {'ratio':>7} {'floor':>6} verdict")
|
||||
print("-" * 82)
|
||||
failures, advisories = 0, 0
|
||||
text_failures, boundary_failures = 0, 0
|
||||
for kind, foreground, background, floor, description in PAIRS:
|
||||
if foreground not in tokens or background not in tokens:
|
||||
print(f"{description:44} {kind:9} {'—':>7} {floor:>6.1f} MISSING TOKEN")
|
||||
failures += 1
|
||||
text_failures += 1
|
||||
continue
|
||||
measured = ratio(tokens[foreground], tokens[background])
|
||||
ok = measured >= floor
|
||||
if not ok:
|
||||
if kind == "text":
|
||||
failures += 1
|
||||
verdict = "FAIL"
|
||||
else:
|
||||
advisories += 1
|
||||
verdict = "below 1.4.11 (label carries it)"
|
||||
else:
|
||||
# Rounded to the two decimals printed, so the verdict matches what the
|
||||
# reader is shown: a pair reported as 3.00:1 is not failed for arithmetic
|
||||
# the output does not display.
|
||||
if round(measured, 2) >= floor:
|
||||
verdict = "pass"
|
||||
else:
|
||||
verdict = "FAIL"
|
||||
if kind == "text":
|
||||
text_failures += 1
|
||||
else:
|
||||
boundary_failures += 1
|
||||
print(f"{description:44} {kind:9} {measured:>6.2f}:1 {floor:>6.1f} {verdict}")
|
||||
print()
|
||||
if failures:
|
||||
print(f"{failures} text pair(s) below WCAG AA — this is a defect")
|
||||
if text_failures:
|
||||
print(f"{text_failures} text pair(s) below WCAG AA (1.4.3) — this is a defect")
|
||||
else:
|
||||
print("every text pair clears WCAG AA (1.4.3)")
|
||||
if advisories:
|
||||
print(f"{advisories} boundary pair(s) below 3:1 (1.4.11). Recorded rather "
|
||||
"than failed: every control in this design carries a visible text "
|
||||
"label, which is measured above and passes.")
|
||||
return 1 if failures else 0
|
||||
if boundary_failures:
|
||||
print(f"{boundary_failures} boundary pair(s) below 3:1 (WCAG 1.4.11) — "
|
||||
"this is a defect")
|
||||
else:
|
||||
print("every control boundary clears 3:1 (1.4.11)")
|
||||
return 1 if (text_failures or boundary_failures) else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
+1111
-137
File diff suppressed because it is too large
Load Diff
@@ -90,6 +90,11 @@ from app.routers import adventures # noqa: E402
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
#: v1.1 release Gate 7 asks for this diagnostic with **memory on**. It shipped
|
||||
#: with no embedding model and the bank switched off, so a release run of it
|
||||
#: would have reported a clean identity result without memory ever taking part.
|
||||
#: Empty keeps the old behaviour, which is what `--scripted` wants.
|
||||
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
||||
|
||||
#: The cast the finding describes: a protagonist and three others, all on stage.
|
||||
CAST = [
|
||||
@@ -140,7 +145,7 @@ def _setup(scripted: bool):
|
||||
db.add(models.Settings(
|
||||
user_id=user.id,
|
||||
model=MODEL or "scripted", endpoint_url=ENDPOINT or "http://127.0.0.1:11434/v1",
|
||||
embedding_model="", context_token_budget=16384, max_output_tokens=500,
|
||||
embedding_model=EMBED_MODEL, context_token_budget=16384, max_output_tokens=500,
|
||||
model_timeout_seconds=300,
|
||||
))
|
||||
db.commit()
|
||||
@@ -167,6 +172,12 @@ def _campaign(client) -> int:
|
||||
})
|
||||
created.raise_for_status()
|
||||
adv = created.json()["id"]
|
||||
# Memory and summaries are per-campaign switches defaulting to off. Gate 7
|
||||
# asks for this diagnostic with memory on, and the ten beats below write
|
||||
# twenty actions — past `MEMORY_START` — so the bank has something to do.
|
||||
client.patch(f"/api/adventures/{adv}",
|
||||
json={"memory_bank_enabled": True, "auto_summarize": True}
|
||||
).raise_for_status()
|
||||
answer = client.post(f"/api/adventures/{adv}/state/corrections", json={
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": key, "entity_type": kind, "name": name}
|
||||
|
||||
@@ -190,6 +190,31 @@ CLUE_FACT = {
|
||||
"fact_id": "silver-key-opens-crypt",
|
||||
}
|
||||
|
||||
#: v1.1 WP-B.1: a second planted fact, established in the **story only**.
|
||||
#:
|
||||
#: The M04 clue above is planted as accepted state, and memories are written
|
||||
#: from story text, so no memory could ever carry it on its own. That is why
|
||||
#: every M04 recovery so far ran through state. This fact is told to the reader
|
||||
#: in narration and never corrected into state, so memory is the only layer that
|
||||
#: is meant to carry it. `--independent-fact` plants it and reports
|
||||
#: `recovered_through_memory_independent` only when every other layer is proven
|
||||
#: not to carry it. The words are copied from `tools/memory_diagnostic.FACT_F`,
|
||||
#: for the reason `HISTORY_LABELS` is copied.
|
||||
INDEPENDENT_FACT_TEXT = ("I watch Mara slip the amber sundial inside the cracked teapot on "
|
||||
"the tavern's top shelf, and she makes me promise to tell no one.")
|
||||
INDEPENDENT_FACT_TERMS = ("sundial", "teapot")
|
||||
INDEPENDENT_RECALL_TEXT = "I ask Mara quietly where she hid the amber sundial."
|
||||
#: How far past the planting turn its memory block can reach. Narration inside
|
||||
#: that block may repeat the fact; narration after it may not.
|
||||
INDEPENDENT_BLOCK_SLACK = 6
|
||||
INDEPENDENT_PRECONDITIONS = (
|
||||
"planted_turn_outside_history",
|
||||
"absent_from_state",
|
||||
"absent_from_summary",
|
||||
"absent_from_knowledge",
|
||||
"absent_from_later_narration",
|
||||
)
|
||||
|
||||
CANON = [
|
||||
"The dead do not return. No rite, relic or bargain has ever returned anyone.",
|
||||
"The abbey crypt has been sealed since the founding.",
|
||||
@@ -365,6 +390,14 @@ class Run:
|
||||
#: The depth of the player turn that planted the clue. M04's
|
||||
#: precondition is that this turn has left the history window.
|
||||
self.planted_depth: int | None = None
|
||||
#: v1.1 WP-B.1, with --independent-fact: where the story-only fact was
|
||||
#: planted, and the accepted-turn count at which each isolation
|
||||
#: precondition first failed.
|
||||
self.independent_fact = False
|
||||
self.independent_depth: int | None = None
|
||||
self.independent_violations: dict[str, int] = {}
|
||||
self.last_done: dict = {}
|
||||
self.last_report: dict = {}
|
||||
|
||||
# ------------------------------------------------------------ recording
|
||||
|
||||
@@ -393,6 +426,8 @@ class Run:
|
||||
"turns_target": self.turns_target,
|
||||
"log_offset": self.log_offset,
|
||||
"planted_depth": self.planted_depth,
|
||||
"independent_depth": self.independent_depth,
|
||||
"independent_violations": self.independent_violations,
|
||||
"written": datetime.now().isoformat(timespec="seconds"),
|
||||
}
|
||||
tmp = self.out / (RESUME_FILE + ".tmp")
|
||||
@@ -414,6 +449,8 @@ class Run:
|
||||
self.elapsed_before = prior.get("elapsed_seconds", 0)
|
||||
self.log_offset = prior.get("log_offset", 0)
|
||||
self.planted_depth = prior.get("planted_depth")
|
||||
self.independent_depth = prior.get("independent_depth")
|
||||
self.independent_violations = dict(prior.get("independent_violations") or {})
|
||||
self.resumed = True
|
||||
|
||||
def reattach(self) -> None:
|
||||
@@ -559,15 +596,55 @@ class Run:
|
||||
self.accepted += 1
|
||||
seconds = time.monotonic() - started
|
||||
sample = self.measure()
|
||||
# v1.1 WP-A1: what the server said it read for the turn just played, from
|
||||
# the `done` event. `.get` because a build before v1.1 sends none.
|
||||
done = next((e for e in events if e.get("type") == "done"), {})
|
||||
accounting = done.get("accounting") or {}
|
||||
sample.update({
|
||||
"accounting_status": accounting.get("status"),
|
||||
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
|
||||
"app_prompt_estimate": accounting.get("estimate"),
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"safety_reserve": accounting.get("safety_reserve"),
|
||||
})
|
||||
self.last_done = done
|
||||
if self.independent_fact and self.independent_depth is not None:
|
||||
self._check_independent_isolation(done, sample)
|
||||
self.note("turn", text=text, seconds=round(seconds, 1), **sample)
|
||||
return {"accepted": True, "seconds": seconds, **sample}
|
||||
|
||||
def _check_independent_isolation(self, done: dict, sample: dict) -> None:
|
||||
"""v1.1 WP-B.1: does anything but memory carry the story-only fact yet?
|
||||
|
||||
Checked on every accepted turn, so a run knows the first turn at which
|
||||
the experiment stopped being about memory, instead of finding out at
|
||||
recall. Each precondition records only its first failure.
|
||||
"""
|
||||
depth = sample.get("total_actions", 0) - 1
|
||||
text = (done.get("action") or {}).get("text") or ""
|
||||
found = {}
|
||||
if depth > self.independent_depth + INDEPENDENT_BLOCK_SLACK and _mentions_fact(text):
|
||||
found["absent_from_later_narration"] = f"narration at depth {depth}"
|
||||
document = self.state().get("document") or {}
|
||||
if _mentions_fact(json.dumps(document)):
|
||||
found["absent_from_state"] = "the narrative state names the fact"
|
||||
summary = next((sec.get("text", "") for sec in (self.last_report.get("sections") or [])
|
||||
if sec.get("label") == SUMMARY_LABEL), "")
|
||||
if _mentions_fact(summary):
|
||||
found["absent_from_summary"] = "the active summary names the fact"
|
||||
for name, detail in found.items():
|
||||
if name not in self.independent_violations:
|
||||
self.independent_violations[name] = self.accepted
|
||||
self.note("independent_precondition_failed", precondition=name, detail=detail)
|
||||
sample["independent_violations"] = dict(self.independent_violations)
|
||||
|
||||
def count_actions(self) -> int:
|
||||
return self.server.call("GET", f"/adventures/{self.adv}/actions?limit=1")["total"]
|
||||
|
||||
def measure(self) -> dict:
|
||||
"""M03's numbers, read from the prompt the app would send right now."""
|
||||
report = self.server.call("GET", f"/adventures/{self.adv}/context")
|
||||
self.last_report = report
|
||||
tokens = report["tokens"]
|
||||
sections = {s["label"]: s["tokens"] for s in report["sections"]}
|
||||
window = report.get("window") or {}
|
||||
@@ -703,6 +780,10 @@ def main() -> int:
|
||||
"--max-consecutive-failures", type=int,
|
||||
default=DEFAULT_MAX_CONSECUTIVE_FAILURES,
|
||||
help="stop and write the evidence after this many unaccepted turns")
|
||||
parser.add_argument(
|
||||
"--independent-fact", action="store_true",
|
||||
help=("v1.1 WP-B.1: also plant a story-only fact at depth 3 and report "
|
||||
"whether memory alone recovers it"))
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL and EMBED_MODEL):
|
||||
@@ -735,6 +816,7 @@ def main() -> int:
|
||||
server.start()
|
||||
run = Run(server, out, turns_target=args.turns,
|
||||
turn_timeout=args.turn_timeout)
|
||||
run.independent_fact = args.independent_fact
|
||||
if prior:
|
||||
run.adopt(prior)
|
||||
|
||||
@@ -771,6 +853,18 @@ def main() -> int:
|
||||
"the planted clue is not in accepted state, so M04 cannot "
|
||||
"be measured from this run. Stopping before the campaign "
|
||||
"starts rather than reporting a recall failure later.")
|
||||
if args.independent_fact:
|
||||
# v1.1 WP-B.1: the story-only fact, told in the next turn and
|
||||
# never corrected into state. Depth 3: the opening, the clue turn
|
||||
# and its reply come first.
|
||||
if any(_mentions_fact(md) for md in (CANON_MD, REFERENCE_MD, INSPIRATION_MD)):
|
||||
raise SystemExit("the imported knowledge names the independent fact")
|
||||
planting_f = run.turn(INDEPENDENT_FACT_TEXT)
|
||||
if not planting_f.get("accepted"):
|
||||
raise SystemExit("the turn that plants the independent fact was not accepted")
|
||||
run.independent_depth = planting_f["total_actions"] - 2
|
||||
run.note("independent_fact_planted", depth=run.independent_depth,
|
||||
terms=list(INDEPENDENT_FACT_TERMS))
|
||||
# The first checkpoint, and the point from which --resume works: the
|
||||
# campaign exists and its clue is planted.
|
||||
run.save_resume()
|
||||
@@ -835,10 +929,15 @@ def main() -> int:
|
||||
# Skipped on an aborted run: it asks the narrator a question, and the
|
||||
# reason the run stopped is that the narrator does not answer.
|
||||
recall = None
|
||||
independent = None
|
||||
if aborted is None:
|
||||
run.note("recall_begin")
|
||||
recall = _recall(run)
|
||||
(out / "recall.json").write_text(json.dumps(recall, indent=2))
|
||||
if args.independent_fact and run.independent_depth is not None:
|
||||
independent = _independent_recall(run, out / "campaign.db")
|
||||
(out / "recall-independent.json").write_text(json.dumps(independent, indent=2))
|
||||
run.note("independent_recall", verdict=independent["verdict"])
|
||||
|
||||
# ---- Export whatever exists, for the recovery evidence. ----
|
||||
# Attempted even for an aborted run: the recovery check and the storage
|
||||
@@ -886,6 +985,7 @@ def main() -> int:
|
||||
"elapsed_seconds": run.elapsed(),
|
||||
"turn_timeout_seconds": args.turn_timeout,
|
||||
"recall": recall,
|
||||
"independent_recall": independent,
|
||||
"final_state": _or_none(lambda: run.state()["document"]),
|
||||
"final_measurement": _or_none(run.measure),
|
||||
"db_bytes": db_path.stat().st_size,
|
||||
@@ -1213,6 +1313,123 @@ def _m04_verdict(recall: dict) -> str:
|
||||
return "not_recovered"
|
||||
|
||||
|
||||
def _mentions_fact(text: str | None) -> bool:
|
||||
"""v1.1 WP-B.1: whether `text` names the independent fact, as a whole word."""
|
||||
low = (text or "").lower()
|
||||
return any(re.search(rf"(?<![a-z]){term}(?![a-z])", low) for term in INDEPENDENT_FACT_TERMS)
|
||||
|
||||
|
||||
def _independent_memory_verdict(check: dict) -> str:
|
||||
"""v1.1 WP-B.1: whether memory alone recovered the story-only fact.
|
||||
|
||||
`recovered_through_memory_independent` requires every precondition, so no
|
||||
other layer could have carried the fact. It also requires that a memory
|
||||
covering the planting turn carries the fact and was injected into the recall
|
||||
turn. A failed precondition is named and is never a recovery, and the M04
|
||||
verdicts above are untouched.
|
||||
"""
|
||||
if check.get("independent_planted_depth") is None:
|
||||
return "precondition_unknown:planted_depth"
|
||||
for name in INDEPENDENT_PRECONDITIONS:
|
||||
value = check.get(name)
|
||||
if value is None:
|
||||
return f"precondition_unknown:{name}"
|
||||
if not value:
|
||||
return f"precondition_failed:{name}"
|
||||
if not check.get("memory_covering_planting_carries_fact"):
|
||||
return "not_recovered:not_created"
|
||||
if check.get("memory_forgotten"):
|
||||
return "not_recovered:evicted"
|
||||
if not check.get("memory_injected"):
|
||||
return "not_recovered:not_injected"
|
||||
return "recovered_through_memory_independent"
|
||||
|
||||
|
||||
def _independent_recall(run: "Run", db_path: Path) -> dict:
|
||||
"""v1.1 WP-B.1: ask for the story-only fact, and find out which layer answered.
|
||||
|
||||
The prompt-level facts come from the recall turn's own stored context. The
|
||||
memory rows come from the campaign database, read-only. Ranking is not
|
||||
recomputed here, because that needs the embedding model;
|
||||
`tools/v11_b1_memory.py diagnose` does it afterwards against a copy of the
|
||||
database.
|
||||
"""
|
||||
import sqlite3
|
||||
import zlib
|
||||
|
||||
result = run.turn(INDEPENDENT_RECALL_TEXT)
|
||||
action_id = (run.last_done.get("action") or {}).get("id")
|
||||
snapshot = (run.server.call("GET", f"/adventures/{run.adv}/actions/{action_id}/context")
|
||||
if result.get("accepted") and action_id else {}) or {}
|
||||
sections = {}
|
||||
for sec in snapshot.get("sections") or []:
|
||||
sections.setdefault(sec.get("label"), []).append(sec.get("text", ""))
|
||||
text_of = {label: "\n".join(parts) for label, parts in sections.items()}
|
||||
floor = (snapshot.get("history") or {}).get("floor_depth")
|
||||
depth = run.independent_depth
|
||||
used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
|
||||
|
||||
covering = []
|
||||
connection = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute(
|
||||
"SELECT id, text, source_start, source_end, forgotten, pinned, use_count, "
|
||||
"branch_id, depth FROM memories WHERE adventure_id = ? AND source_start <= ? "
|
||||
"AND source_end >= ? ORDER BY id", (run.adv, depth, depth)).fetchall()
|
||||
blob = connection.execute(
|
||||
"SELECT context_snapshot FROM actions WHERE id = ?", (action_id or -1,)).fetchone()
|
||||
finally:
|
||||
connection.close()
|
||||
for row in rows:
|
||||
memory_id, text, start, end, forgotten, pinned, use_count, branch_id, node_depth = row
|
||||
covering.append({
|
||||
"memory_id": memory_id, "text": text, "source_start": start, "source_end": end,
|
||||
"forgotten": bool(forgotten), "pinned": bool(pinned), "use_count": use_count,
|
||||
"branch_id": branch_id, "depth": node_depth,
|
||||
"carries_fact": all(re.search(rf"(?<![a-z]){t}(?![a-z])", (text or "").lower())
|
||||
for t in INDEPENDENT_FACT_TERMS),
|
||||
"injected": memory_id in used,
|
||||
})
|
||||
carrying = [c for c in covering if c["carries_fact"]]
|
||||
best = next((c for c in carrying if c["injected"]), carrying[0] if carrying else None)
|
||||
stored_snapshot_readable = blob is not None and blob[0] is not None
|
||||
if stored_snapshot_readable:
|
||||
try:
|
||||
json.loads(zlib.decompress(blob[0]))
|
||||
except Exception: # noqa: BLE001
|
||||
stored_snapshot_readable = False
|
||||
|
||||
document = run.state().get("document") or {}
|
||||
violations = dict(run.independent_violations)
|
||||
check = {
|
||||
"independent_planted_depth": depth,
|
||||
"recall_accepted": bool(result.get("accepted")),
|
||||
"history_floor_depth": floor,
|
||||
"planted_turn_outside_history": (None if not snapshot else
|
||||
floor is not None and depth < floor),
|
||||
"absent_from_state": ("absent_from_state" not in violations
|
||||
and not _mentions_fact(json.dumps(document))
|
||||
and not _mentions_fact(text_of.get(STATE_LABEL))),
|
||||
"absent_from_summary": ("absent_from_summary" not in violations
|
||||
and not _mentions_fact(text_of.get(SUMMARY_LABEL))),
|
||||
"absent_from_knowledge": not any(
|
||||
_mentions_fact(text_of.get(label)) for label in IMPORTED_KNOWLEDGE_LABELS),
|
||||
"absent_from_later_narration": "absent_from_later_narration" not in violations,
|
||||
"violations_first_turn": violations,
|
||||
"covering_memories": covering,
|
||||
"memory_covering_planting_carries_fact": bool(carrying),
|
||||
"memory_forgotten": bool(best and best["forgotten"]),
|
||||
"memory_injected": bool(best and best["injected"]),
|
||||
"memory_text_in_memories_section": bool(
|
||||
best and best["text"] and best["text"] in (text_of.get(MEMORIES_LABEL) or "")),
|
||||
"memory_ids_used": used,
|
||||
"recall_action_id": action_id,
|
||||
"stored_snapshot_readable": stored_snapshot_readable,
|
||||
}
|
||||
check["verdict"] = _independent_memory_verdict(check)
|
||||
return check
|
||||
|
||||
|
||||
#: Signs the application stored protocol as story. The first is a state-section
|
||||
#: heading with an indented entry under it, in any markdown, because
|
||||
#: `## Established:` got past a plain substring match and the count read 1
|
||||
@@ -1225,10 +1442,26 @@ PROTOCOL_LEAK_HEADING_RE = re.compile(
|
||||
re.MULTILINE,
|
||||
)
|
||||
PROTOCOL_LEAK_EVENTS = '"events"'
|
||||
#: v1.1 WP-A2: the two shapes the M11 closeout's identity run stored that the
|
||||
#: two signs above cannot see — a line opening with a call to an event, and the
|
||||
#: length hint echoed with the application's own wording. Copied, as above.
|
||||
PROTOCOL_LEAK_CALL_RE = re.compile(
|
||||
r"^[ \t]*(?:>[ \t]*)?(?:create_entity|set_entity_status|set_entity_attribute"
|
||||
r"|set_entity_conditions|set_current_location|set_possession|clear_possession"
|
||||
r"|add_fact|invalidate_fact|add_relationship|end_relationship"
|
||||
r"|open_story_thread|resolve_story_thread|set_scene)[ \t]*\(",
|
||||
re.IGNORECASE | re.MULTILINE,
|
||||
)
|
||||
PROTOCOL_LEAK_HINT_RE = re.compile(
|
||||
r"\[Hard limit:[^\]]*(?:append the state block|turn must not exceed \d+ words)",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
|
||||
|
||||
def _leaks_protocol(text: str) -> bool:
|
||||
return bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
|
||||
return (bool(PROTOCOL_LEAK_HEADING_RE.search(text)) or PROTOCOL_LEAK_EVENTS in text
|
||||
or bool(PROTOCOL_LEAK_CALL_RE.search(text))
|
||||
or bool(PROTOCOL_LEAK_HINT_RE.search(text)))
|
||||
|
||||
|
||||
def _protocol_leaks(bundle: dict) -> dict:
|
||||
|
||||
@@ -15,10 +15,17 @@ what M9 recorded as "this machine cannot drive a file into the browser". The
|
||||
narrower and more useful statement is that it refuses `/tmp`: a path under the
|
||||
user's home works. `stage()` exists to put evidence files there, so knowledge
|
||||
import can be exercised through the real file input rather than in two halves.
|
||||
|
||||
**Downloads (v1.1 WP-C).** The same snap Firefox saves a download without any
|
||||
dialog when its profile says where to, and the folder is under `$HOME`.
|
||||
`firefox_download_prefs` is that profile, `require_under_home` refuses a folder
|
||||
the sandbox would not let it write, and `wait_for_download` decides when a file
|
||||
has actually finished arriving — never the click that started it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
@@ -33,6 +40,10 @@ GECKODRIVER = shutil.which("geckodriver") or "/snap/bin/geckodriver"
|
||||
#: Where files the browser must open are staged. Under $HOME because the snap
|
||||
#: sandbox denies /tmp; see the module docstring.
|
||||
STAGE = Path.home() / "m11-evidence"
|
||||
#: The W3C key an element reference is returned under.
|
||||
ELEMENT_KEY = "element-6066-11e4-a52e-4f735466cecf"
|
||||
#: What Firefox names a download while it is still arriving.
|
||||
PARTIAL_SUFFIXES = (".part",)
|
||||
|
||||
|
||||
def stage(name: str, body: str | bytes) -> str:
|
||||
@@ -51,14 +62,86 @@ def free_port() -> int:
|
||||
return s.getsockname()[1]
|
||||
|
||||
|
||||
def geckodriver_version() -> str:
|
||||
try:
|
||||
out = subprocess.run([GECKODRIVER, "--version"], capture_output=True, text=True, timeout=30)
|
||||
return (out.stdout.splitlines() or ["?"])[0].strip()
|
||||
except (OSError, subprocess.SubprocessError):
|
||||
return "?"
|
||||
|
||||
|
||||
class WebDriverError(RuntimeError):
|
||||
pass
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ downloads
|
||||
|
||||
def require_under_home(path: Path) -> Path:
|
||||
"""`path`, resolved, if it is inside the user's home; otherwise refuse.
|
||||
|
||||
The snap sandbox will not write elsewhere, and a download folder under
|
||||
`/tmp` would also put evidence where a reboot deletes it.
|
||||
"""
|
||||
resolved = Path(path).expanduser().resolve()
|
||||
home = Path.home().resolve()
|
||||
if resolved != home and home not in resolved.parents:
|
||||
raise WebDriverError(f"{resolved} is not under {home}; the browser cannot write there")
|
||||
return resolved
|
||||
|
||||
|
||||
def firefox_download_prefs(directory: Path) -> dict:
|
||||
"""Profile preferences that save every download to `directory`, unasked."""
|
||||
return {
|
||||
"browser.download.folderList": 2, # 2 = the folder named below
|
||||
"browser.download.dir": str(directory),
|
||||
"browser.download.useDownloadDir": True,
|
||||
"browser.download.start_downloads_in_tmp_dir": False,
|
||||
"browser.download.always_ask_before_handling_new_types": False,
|
||||
"browser.helperApps.neverAsk.saveToDisk": "application/json,application/octet-stream",
|
||||
"browser.download.manager.showWhenStarting": False,
|
||||
"browser.download.alwaysOpenPanel": False,
|
||||
"browser.download.panel.shown": True,
|
||||
}
|
||||
|
||||
|
||||
def wait_for_download(directory: Path, before: set[str], *, timeout: float = 60,
|
||||
poll: float = 0.2, stable_polls: int = 3) -> Path:
|
||||
"""The file a download wrote into `directory`, once it has finished.
|
||||
|
||||
Finished means all of these, at once:
|
||||
- a name that was not in `before` (the listing taken before the click);
|
||||
- no in-progress file (`*.part`) left in the folder;
|
||||
- more than zero bytes;
|
||||
- the same size for `stable_polls` consecutive polls.
|
||||
|
||||
A first appearance is not a finished download, and a zero-byte or partial
|
||||
file never counts. Raises `WebDriverError` when nothing finishes in time.
|
||||
"""
|
||||
deadline = time.monotonic() + timeout
|
||||
last: dict[str, int] = {}
|
||||
steady: dict[str, int] = {}
|
||||
while time.monotonic() < deadline:
|
||||
names = {p.name for p in directory.iterdir()} if directory.exists() else set()
|
||||
partial = any(n.endswith(PARTIAL_SUFFIXES) for n in names)
|
||||
fresh = sorted(n for n in names - before if not n.endswith(PARTIAL_SUFFIXES))
|
||||
for name in fresh:
|
||||
size = (directory / name).stat().st_size
|
||||
steady[name] = steady.get(name, 0) + 1 if last.get(name) == size else 1
|
||||
last[name] = size
|
||||
if not partial and size > 0 and steady[name] >= stable_polls:
|
||||
return directory / name
|
||||
time.sleep(poll)
|
||||
listing = sorted(p.name for p in directory.iterdir()) if directory.exists() else []
|
||||
raise WebDriverError(f"no finished download in {directory} within {timeout}s; saw {listing}")
|
||||
|
||||
|
||||
# -------------------------------------------------------------------- browser
|
||||
|
||||
class Browser:
|
||||
"""One headless Firefox, driven over the wire protocol."""
|
||||
|
||||
def __init__(self, *, headless: bool = True, log: Path | None = None):
|
||||
def __init__(self, *, headless: bool = True, log: Path | None = None,
|
||||
download_dir: Path | None = None):
|
||||
self.port = free_port()
|
||||
handle = open(log, "ab") if log else subprocess.DEVNULL
|
||||
self.proc = subprocess.Popen(
|
||||
@@ -68,9 +151,15 @@ class Browser:
|
||||
self.base = f"http://127.0.0.1:{self.port}"
|
||||
self._wait_for_driver()
|
||||
args = ["-headless"] if headless else []
|
||||
options: dict = {"args": args}
|
||||
self.download_dir = None
|
||||
if download_dir is not None:
|
||||
self.download_dir = require_under_home(download_dir)
|
||||
self.download_dir.mkdir(parents=True, exist_ok=True)
|
||||
options["prefs"] = firefox_download_prefs(self.download_dir)
|
||||
answer = self._call("POST", "/session", {"capabilities": {"alwaysMatch": {
|
||||
"browserName": "firefox",
|
||||
"moz:firefoxOptions": {"args": args},
|
||||
"moz:firefoxOptions": options,
|
||||
# Never silently accept a bad certificate: the endpoint policy and
|
||||
# the TLS trust union are release claims (H12, A06), and a browser
|
||||
# that ignored certificates would hide a failure of either.
|
||||
@@ -78,6 +167,7 @@ class Browser:
|
||||
}}})["value"]
|
||||
self.session = answer["sessionId"]
|
||||
self.version = answer["capabilities"].get("browserVersion", "?")
|
||||
self.capabilities = answer["capabilities"]
|
||||
|
||||
# ------------------------------------------------------------- plumbing
|
||||
|
||||
@@ -123,6 +213,9 @@ class Browser:
|
||||
def go(self, url: str) -> None:
|
||||
self._call("POST", self._s("/url"), {"url": url})
|
||||
|
||||
def reload(self) -> None:
|
||||
self._call("POST", self._s("/refresh"), {})
|
||||
|
||||
@property
|
||||
def url(self) -> str:
|
||||
return self._call("GET", self._s("/url"))["value"]
|
||||
@@ -134,10 +227,56 @@ class Browser:
|
||||
def source(self) -> str:
|
||||
return self._call("GET", self._s("/source"))["value"]
|
||||
|
||||
def screenshot(self, path) -> Path:
|
||||
"""The viewport as a PNG, written where you ask (v1.1 WP-E).
|
||||
|
||||
Evidence for a change a reader judges by looking at it: a contrast ratio
|
||||
says a boundary is measurable, and a picture says what it looks like.
|
||||
"""
|
||||
encoded = self._call("GET", self._s("/screenshot"))["value"]
|
||||
target = Path(path)
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
target.write_bytes(base64.b64decode(encoded))
|
||||
return target
|
||||
|
||||
def hover(self, element: str) -> None:
|
||||
"""A real pointer over an element, so `:hover` actually applies.
|
||||
|
||||
Dispatching a mouseover event from JavaScript does not do this: CSS
|
||||
`:hover` follows the browser's own pointer state, not a synthetic event,
|
||||
so a measurement taken after `dispatchEvent` reads the resting style and
|
||||
reports it as the hover style. This moves the pointer (v1.1 WP-E).
|
||||
"""
|
||||
self._call("POST", self._s("/execute/sync"), {
|
||||
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
|
||||
"args": [{ELEMENT_KEY: element}]})
|
||||
self._call("POST", self._s("/actions"), {"actions": [{
|
||||
"type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
|
||||
"actions": [{"type": "pointerMove", "duration": 60,
|
||||
"origin": {ELEMENT_KEY: element}, "x": 0, "y": 0}]}]})
|
||||
|
||||
def unhover(self) -> None:
|
||||
"""Move the pointer off whatever it was over, and forget the input state."""
|
||||
self._call("POST", self._s("/actions"), {"actions": [{
|
||||
"type": "pointer", "id": "mouse", "parameters": {"pointerType": "mouse"},
|
||||
"actions": [{"type": "pointerMove", "duration": 30,
|
||||
"origin": "viewport", "x": 0, "y": 0}]}]})
|
||||
try:
|
||||
self._call("DELETE", self._s("/actions"))
|
||||
except WebDriverError:
|
||||
pass
|
||||
|
||||
def js(self, script: str, *args):
|
||||
return self._call("POST", self._s("/execute/sync"),
|
||||
{"script": script, "args": list(args)})["value"]
|
||||
|
||||
def element_by_js(self, script: str, *args):
|
||||
"""An element a script returns, as a reference `click` can use, or None."""
|
||||
value = self.js(script, *args)
|
||||
if isinstance(value, dict) and ELEMENT_KEY in value:
|
||||
return value[ELEMENT_KEY]
|
||||
return None
|
||||
|
||||
def find(self, css: str, *, required=True):
|
||||
try:
|
||||
answer = self._call("POST", self._s("/element"),
|
||||
@@ -163,6 +302,17 @@ class Browser:
|
||||
return self._call("GET", self._s(f"/element/{element}/property/{name}"))["value"]
|
||||
|
||||
def click(self, element: str) -> None:
|
||||
"""A real click, on an element first scrolled to the middle of the view.
|
||||
|
||||
WebDriver scrolls a target only as far as its edge, and the play page's
|
||||
composer is fixed to the bottom of the window: a control just under it
|
||||
(a failure notice's details, a turn's Inspect button) is then covered,
|
||||
and the click is intercepted. A reader scrolls it clear first; so does
|
||||
this (v1.1 WP-C).
|
||||
"""
|
||||
self._call("POST", self._s("/execute/sync"), {
|
||||
"script": "arguments[0].scrollIntoView({block: 'center', inline: 'nearest'})",
|
||||
"args": [{ELEMENT_KEY: element}]})
|
||||
self._call("POST", self._s(f"/element/{element}/click"), {})
|
||||
|
||||
def clear(self, element: str) -> None:
|
||||
@@ -183,6 +333,21 @@ class Browser:
|
||||
answer = self._call("GET", self._s("/element/active"))
|
||||
return list(answer["value"].values())[0]
|
||||
|
||||
# -------------------------------------------------------------- windows
|
||||
|
||||
@property
|
||||
def window(self) -> str:
|
||||
return self._call("GET", self._s("/window"))["value"]
|
||||
|
||||
def new_tab(self) -> str:
|
||||
return self._call("POST", self._s("/window/new"), {"type": "tab"})["value"]["handle"]
|
||||
|
||||
def switch_to(self, handle: str) -> None:
|
||||
self._call("POST", self._s("/window"), {"handle": handle})
|
||||
|
||||
def close_window(self) -> None:
|
||||
self._call("DELETE", self._s("/window"))
|
||||
|
||||
# ------------------------------------------------------------- waiting
|
||||
|
||||
def wait_for(self, css: str, *, timeout=90, gone=False):
|
||||
@@ -196,12 +361,19 @@ class Browser:
|
||||
f"{'still present' if gone else 'never appeared'}: {css}")
|
||||
|
||||
def wait_until(self, script: str, *, timeout=90, what=""):
|
||||
if self.wait_js(script, timeout=timeout):
|
||||
return True
|
||||
raise WebDriverError(f"condition never held: {what or script}")
|
||||
|
||||
def wait_js(self, script: str, *, timeout=90) -> bool:
|
||||
"""Whether `script` became true within `timeout`. For a check to record,
|
||||
where `wait_until` is for a precondition that must hold."""
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
if self.js(f"return ({script})"):
|
||||
return True
|
||||
time.sleep(0.25)
|
||||
raise WebDriverError(f"condition never held: {what or script}")
|
||||
return False
|
||||
|
||||
|
||||
class Site:
|
||||
|
||||
@@ -0,0 +1,960 @@
|
||||
"""v1.1 WP-B.1: where an early story fact is lost on its way to the narrator.
|
||||
|
||||
One planted fact **F** has four stages to survive before the narrator can use
|
||||
it from memory, and this module reports each one separately:
|
||||
|
||||
created a memory whose `source_start`..`source_end` covers the planting
|
||||
depth carries F
|
||||
retained that memory is not `forgotten`
|
||||
ranked it is eligible on the active lineage and embedded, and where it
|
||||
scores for the recall query against `memory_top_k`
|
||||
injected the recall turn's own stored `memories.used` names it, and its text
|
||||
is in that turn's `used_memories` section
|
||||
|
||||
A fact is only evidence about memory if memory is the **only** thing carrying it.
|
||||
`isolation()` checks every other layer: the authoritative document, per-node state
|
||||
snapshots, the active summary, imported knowledge, the narration after the
|
||||
planting block, and the recent-history window. A run where any of those carries F
|
||||
is reported as a failed precondition, never as a memory result.
|
||||
|
||||
**Nothing here changes behaviour.**
|
||||
- It reads rows.
|
||||
- It reuses production's own pure helpers (`memorybank._drop_redundant`,
|
||||
`memorybank.classify_authority`, `vectors.cosine`, `lineage.path_of`), so its
|
||||
ranking is production's ranking, not a second opinion.
|
||||
- It checks itself against what the recall turn actually recorded.
|
||||
- The only computed fields are ephemeral report data. No column or table is
|
||||
added.
|
||||
|
||||
The deterministic stubs at the bottom stand in for the models when a test needs a
|
||||
fixed answer. **Read what they model before reading any result they produce:**
|
||||
|
||||
- `BestCaseSummariser` keeps F if and only if F is in the excerpt it is given.
|
||||
It is the ideal summariser, so a creation failure under it is the
|
||||
application's, not the model's.
|
||||
- `ConceptEmbedder` maps words to a small concept table, so that "the brass dial
|
||||
that tells the hour" lands near "sundial". It models what an embedding is
|
||||
supposed to do. It says nothing about how well `nomic-embed-text` does it,
|
||||
which is what the real-model run is for.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import math
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from app import memorybank, models, summaries, vectors
|
||||
from app.context import builder, history, lineage
|
||||
from app.knowledge import classes as knowledge_classes
|
||||
|
||||
VERDICTS = (
|
||||
"not_created",
|
||||
"created_but_evicted",
|
||||
"retained_but_not_ranked",
|
||||
"ranked_but_not_selected",
|
||||
"selected_but_not_injected",
|
||||
"injected",
|
||||
)
|
||||
|
||||
#: Section labels in a stored context snapshot. Copied from the builder's
|
||||
#: vocabulary so a renamed section fails loudly here.
|
||||
HISTORY_LABELS = ("history", "recent_history")
|
||||
SUMMARY_LABEL = "story_summary"
|
||||
MEMORIES_LABEL = "used_memories"
|
||||
STATE_LABEL = "narrative_state"
|
||||
KNOWLEDGE_LABELS = (
|
||||
knowledge_classes.SECTION_CANON,
|
||||
knowledge_classes.SECTION_REFERENCE,
|
||||
knowledge_classes.SECTION_INSPIRATION,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Fact:
|
||||
"""A planted fact, and how to recognise it in a text.
|
||||
|
||||
`carry_groups`: a text carries the fact when every group matches, where a
|
||||
group matches when any one of its terms appears as a whole word. A memory has
|
||||
to name both the thing and where it is to carry "where the thing is".
|
||||
|
||||
`leak_terms`: any one of these in another layer means that layer carries the
|
||||
fact. This is deliberately looser than `carry_groups`. For isolation, a
|
||||
mention is enough to disqualify.
|
||||
"""
|
||||
|
||||
fact_id: str
|
||||
sentence: str
|
||||
carry_groups: tuple[tuple[str, ...], ...]
|
||||
leak_terms: tuple[str, ...]
|
||||
|
||||
def carried_by(self, text: str | None) -> bool:
|
||||
low = (text or "").lower()
|
||||
return all(any(_has_word(low, term) for term in group) for group in self.carry_groups)
|
||||
|
||||
def mentioned_by(self, text: str | None) -> bool:
|
||||
low = (text or "").lower()
|
||||
return any(_has_word(low, term) for term in self.leak_terms)
|
||||
|
||||
|
||||
def _has_word(low: str, term: str) -> bool:
|
||||
return re.search(rf"(?<![a-z]){re.escape(term.lower())}(?![a-z])", low) is not None
|
||||
|
||||
|
||||
#: The fixture's planted fact. Chosen to be natural in a tavern scene and absent
|
||||
#: from every existing fixture: no "sundial" or "teapot" appears anywhere in the
|
||||
#: Westhaven campaign, its knowledge files or its beats.
|
||||
FACT_F = Fact(
|
||||
fact_id="F-amber-sundial",
|
||||
sentence="Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf.",
|
||||
carry_groups=(("sundial",), ("teapot",)),
|
||||
leak_terms=("sundial", "teapot"),
|
||||
)
|
||||
#: The abandoned-line control fact.
|
||||
FACT_G = Fact(
|
||||
fact_id="G-iron-weathervane",
|
||||
sentence="Edrin buried the iron weathervane beneath the mill's broken waterwheel.",
|
||||
carry_groups=(("weathervane",), ("waterwheel",)),
|
||||
leak_terms=("weathervane", "waterwheel"),
|
||||
)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ reading
|
||||
|
||||
def _lineage_actions(db, adventure):
|
||||
path = lineage.path_of(db, adventure)
|
||||
return (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adventure.id, path.clause(models.Action))
|
||||
.order_by(models.Action.depth, models.Action.id)
|
||||
.all()
|
||||
)
|
||||
|
||||
|
||||
def covering_memories(db, adventure, depth: int, *, any_branch: bool = False):
|
||||
"""Memories whose source range covers `depth`, oldest first."""
|
||||
query = select(models.Memory).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.source_start <= depth,
|
||||
models.Memory.source_end >= depth,
|
||||
)
|
||||
if not any_branch:
|
||||
query = query.where(lineage.path_of(db, adventure).clause(models.Memory))
|
||||
return db.execute(query.order_by(models.Memory.id)).scalars().all()
|
||||
|
||||
|
||||
def planting_block_end(db, adventure, plant_depth: int) -> int:
|
||||
"""The last depth of the memory block holding the planted turn.
|
||||
|
||||
Taken from the memory that covers it where one exists. Before one exists it
|
||||
is the furthest a block could reach, so a later-narration check never counts
|
||||
a turn inside the planting block as a repetition.
|
||||
"""
|
||||
rows = covering_memories(db, adventure, plant_depth)
|
||||
if rows:
|
||||
return max(row.source_end for row in rows)
|
||||
return plant_depth + memorybank.MEMORY_INTERVAL
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- isolation
|
||||
|
||||
def isolation(db, adventure, fact: Fact, plant_depth: int, *,
|
||||
recall_snapshot: dict | None = None,
|
||||
recall_depth: int | None = None) -> dict:
|
||||
"""Every layer other than memory that could carry F, checked.
|
||||
|
||||
Returns `{check: {"ok": bool, "detail": str}}` and `ok` over all of them.
|
||||
With `recall_snapshot`, the recall turn's stored context, the prompt-level
|
||||
checks (history window, summary section, knowledge sections) are made
|
||||
against what the narrator was actually given.
|
||||
"""
|
||||
checks: dict[str, dict] = {}
|
||||
|
||||
document = adventure.narrative_state or {}
|
||||
hits = [key for key in ("entities", "facts", "relationships", "threads", "scene",
|
||||
"possessions")
|
||||
if fact.mentioned_by(json.dumps(document.get(key), default=str))]
|
||||
checks["state_document"] = {
|
||||
"ok": not hits and not fact.mentioned_by(json.dumps(document, default=str)),
|
||||
"detail": f"mentioned in {hits}" if hits else "absent",
|
||||
}
|
||||
|
||||
snapshot_hits = []
|
||||
later_hits = []
|
||||
block_end = planting_block_end(db, adventure, plant_depth)
|
||||
for action in _lineage_actions(db, adventure):
|
||||
if fact.mentioned_by(json.dumps(action.narrative_state_after, default=str)):
|
||||
snapshot_hits.append(action.depth)
|
||||
if (action.type == "ai" and action.depth is not None and action.depth > block_end
|
||||
and (recall_depth is None or action.depth < recall_depth)
|
||||
and fact.mentioned_by(action.text)):
|
||||
later_hits.append(action.depth)
|
||||
checks["state_snapshots"] = {
|
||||
"ok": not snapshot_hits,
|
||||
"detail": f"mentioned in snapshots at depths {snapshot_hits[:10]}" if snapshot_hits
|
||||
else "absent from every node's narrative_state_after on the active lineage",
|
||||
}
|
||||
checks["later_narration"] = {
|
||||
"ok": not later_hits,
|
||||
"detail": (f"narration after the planting block (ends at depth {block_end}) "
|
||||
f"mentions the fact at depths {later_hits[:10]}") if later_hits
|
||||
else f"no narrator turn after depth {block_end} mentions the fact",
|
||||
}
|
||||
|
||||
active = summaries.current(db, adventure)
|
||||
summary_text = active.text if active is not None else ""
|
||||
if recall_snapshot is not None:
|
||||
summary_text += "\n" + _section(recall_snapshot, SUMMARY_LABEL)
|
||||
checks["summary"] = {
|
||||
"ok": not fact.mentioned_by(summary_text),
|
||||
"detail": "the active summary mentions the fact" if fact.mentioned_by(summary_text)
|
||||
else ("absent from the active summary" if active is not None else "no summary yet"),
|
||||
}
|
||||
|
||||
sources = db.execute(
|
||||
select(models.KnowledgeSource.content).where(
|
||||
models.KnowledgeSource.adventure_id == adventure.id)
|
||||
).scalars().all()
|
||||
knowledge_text = "\n".join(s or "" for s in sources)
|
||||
if recall_snapshot is not None:
|
||||
knowledge_text += "\n" + "\n".join(_section(recall_snapshot, l) for l in KNOWLEDGE_LABELS)
|
||||
checks["knowledge"] = {
|
||||
"ok": not fact.mentioned_by(knowledge_text),
|
||||
"detail": "imported knowledge mentions the fact" if fact.mentioned_by(knowledge_text)
|
||||
else f"absent from {len(sources)} imported source(s)",
|
||||
}
|
||||
|
||||
if recall_snapshot is not None:
|
||||
hist = recall_snapshot.get("history") or {}
|
||||
floor = hist.get("floor_depth")
|
||||
history_text = "\n".join(_section(recall_snapshot, l) for l in HISTORY_LABELS)
|
||||
outside = floor is not None and plant_depth < floor
|
||||
checks["recent_history"] = {
|
||||
"ok": outside and not fact.carried_by(history_text),
|
||||
"detail": (f"history window starts at depth {floor}; planted at {plant_depth}; "
|
||||
f"fact text in history sections: {fact.carried_by(history_text)}"),
|
||||
}
|
||||
checks["state_section"] = {
|
||||
"ok": not fact.mentioned_by(_section(recall_snapshot, STATE_LABEL)),
|
||||
"detail": "the recall prompt's narrative_state section "
|
||||
+ ("mentions the fact" if fact.mentioned_by(_section(recall_snapshot, STATE_LABEL))
|
||||
else "does not mention the fact"),
|
||||
}
|
||||
|
||||
return {"ok": all(c["ok"] for c in checks.values()), "checks": checks}
|
||||
|
||||
|
||||
def _section(snapshot: dict, label: str) -> str:
|
||||
return "\n".join(s.get("text", "") for s in (snapshot.get("sections") or [])
|
||||
if s.get("label") == label)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- stages
|
||||
|
||||
async def rank_bank(db, adventure, settings, query: dict, embed) -> dict:
|
||||
"""Production's ranking, recomputed for `query`, for every eligible memory.
|
||||
|
||||
`query` is a `memorybank.retrieval_query` dict. The catalogue clause, the
|
||||
scoring (`memorybank.score_candidates`) and the selection with its pins and
|
||||
redundancy rule (`memorybank.select_memories`) are production's own
|
||||
functions, so this is production's ranking, not a second opinion. Returns
|
||||
every scored row, not just the top-k, because "where did F rank" is the
|
||||
question.
|
||||
"""
|
||||
catalogue = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.authority,
|
||||
models.Memory.embedding_blob, models.Memory.text).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory),
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.embedded.is_(True),
|
||||
)
|
||||
).all()
|
||||
top_k = max(1, settings.memory_top_k)
|
||||
texts = [t for t in (query["input"], query["context"]) if t.strip()]
|
||||
if not catalogue or not texts:
|
||||
return {"query": query, "scored": [], "selected": [], "top_k": top_k}
|
||||
vectors_by_text = dict(zip(texts, await embed(texts)))
|
||||
input_vec = vectors_by_text.get(query["input"]) if query["input"].strip() else None
|
||||
context_vec = vectors_by_text.get(query["context"]) if query["context"].strip() else None
|
||||
held = {row.id: vectors.unpack(row.embedding_blob) for row in catalogue if row.embedding_blob}
|
||||
terms_of = ({row.id: memorybank.lexical_terms(row.text or "") for row in catalogue}
|
||||
if query["input_terms"] else {})
|
||||
authority_of = {row.id: row.authority for row in catalogue}
|
||||
pinned_of = {row.id: row.pinned for row in catalogue}
|
||||
scored = memorybank.score_candidates(
|
||||
[row.id for row in catalogue if row.id in held], held, terms_of,
|
||||
input_vec, context_vec, query["input_terms"])
|
||||
used, suppressed = memorybank.select_memories(scored, pinned_of, held, authority_of, top_k)
|
||||
selected = {row[1] for row in used}
|
||||
suppressed_by = dict(suppressed)
|
||||
return {
|
||||
"query": query,
|
||||
"top_k": top_k,
|
||||
"scored": [
|
||||
{"rank": i + 1, "memory_id": memory_id, "similarity": round(semantic, 4),
|
||||
"semantic_score": round(semantic, 4), "lexical_score": round(lexical, 4),
|
||||
"final_score": round(final, 4),
|
||||
"pinned": pinned_of[memory_id], "selected": memory_id in selected,
|
||||
"suppressed_as_duplicate_of": suppressed_by.get(memory_id)}
|
||||
for i, (final, memory_id, semantic, lexical) in enumerate(scored)
|
||||
],
|
||||
"selected": sorted(selected),
|
||||
}
|
||||
|
||||
|
||||
def production_query(adventure, exclude_action_id: int | None) -> dict:
|
||||
"""The retrieval query a turn used, built by production's own `retrieval_query`."""
|
||||
return memorybank.retrieval_query(adventure, exclude_action_id)
|
||||
|
||||
|
||||
def variant_query(base: dict, player_input: str) -> dict:
|
||||
"""`base` with a different player input: "what if the player had asked this
|
||||
here", with the scene and narration context the recall turn really had."""
|
||||
return {"input": player_input, "context": base["context"],
|
||||
"input_terms": sorted(memorybank.lexical_terms(player_input))}
|
||||
|
||||
|
||||
def eviction_order(db, adventure) -> list[int]:
|
||||
"""The order `_evict_over_capacity` would take unpinned active memories in.
|
||||
|
||||
Production's own `memorybank.eviction_order`, run to the end of the bank."""
|
||||
rows = db.execute(
|
||||
select(models.Memory.id, models.Memory.pinned, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.last_used_at,
|
||||
models.Memory.created_at, models.Memory.use_count).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False),
|
||||
)
|
||||
).all()
|
||||
return memorybank.eviction_order(rows, len(rows))
|
||||
|
||||
|
||||
def coverage(ranges: list[tuple[int, int]], tip: int | None) -> dict:
|
||||
"""How much of the story `ranges` (active memories' source ranges) describe.
|
||||
|
||||
`largest_gap` is the longest run of depths, between the first memory's start
|
||||
and `tip`, that no memory covers. It is how the eviction rule is judged in
|
||||
general, not only for the planted fact."""
|
||||
if not ranges:
|
||||
return {"first_start": None, "last_end": None, "largest_gap": None}
|
||||
ordered = sorted(ranges)
|
||||
gaps = []
|
||||
reach = ordered[0][1]
|
||||
for start, end in ordered[1:]:
|
||||
gaps.append(max(0, start - reach - 1))
|
||||
reach = max(reach, end)
|
||||
return {"first_start": ordered[0][0], "last_end": reach,
|
||||
"largest_gap": max(gaps, default=0)}
|
||||
|
||||
|
||||
async def diagnose(db, adventure, settings, fact: Fact, plant_depth: int, *,
|
||||
recall_action: models.Action, embed) -> dict:
|
||||
"""The four stages for `fact`, judged at `recall_action`, the recall turn's AI node.
|
||||
|
||||
Ranking is recomputed with the query that turn used, and checked against the
|
||||
turn's own stored `memories.used`. Injection is read from that snapshot, so
|
||||
it reports what the narrator was actually given, not a re-run.
|
||||
"""
|
||||
snapshot = recall_action.context_snapshot or {}
|
||||
out: dict = {"fact_id": fact.fact_id, "plant_depth": plant_depth,
|
||||
"recall_depth": recall_action.depth}
|
||||
|
||||
covering = covering_memories(db, adventure, plant_depth)
|
||||
carrying = [m for m in covering if fact.carried_by(m.text)]
|
||||
elsewhere = [m for m in db.execute(select(models.Memory).where(
|
||||
models.Memory.adventure_id == adventure.id)).scalars().all()
|
||||
if fact.carried_by(m.text) and m not in carrying]
|
||||
creation_input = []
|
||||
for memory in covering:
|
||||
block = memorybank.source_block(db, memory)
|
||||
raw = "\n\n".join(a.text for a in block)
|
||||
excerpt = memorybank.memory_excerpt(raw) # what `summarize_block` sends
|
||||
creation_input.append({
|
||||
"memory_id": memory.id, "source_start": memory.source_start,
|
||||
"source_end": memory.source_end, "block_tokens": builder.count_tokens(raw),
|
||||
"fact_in_block": fact.carried_by(raw),
|
||||
"fact_in_summariser_excerpt": fact.carried_by(excerpt),
|
||||
"memory_text": memory.text,
|
||||
})
|
||||
memory = carrying[0] if carrying else None
|
||||
out["created"] = {
|
||||
"yes": memory is not None,
|
||||
"memory_id": getattr(memory, "id", None),
|
||||
"source_start": getattr(memory, "source_start", None),
|
||||
"source_end": getattr(memory, "source_end", None),
|
||||
"memory_text": getattr(memory, "text", None),
|
||||
"covering_memories": creation_input,
|
||||
"no_covering_memory": not covering,
|
||||
"carried_by_other_memories": [
|
||||
{"memory_id": m.id, "source_start": m.source_start, "source_end": m.source_end}
|
||||
for m in elsewhere],
|
||||
}
|
||||
|
||||
if memory is None:
|
||||
out["verdict"] = "not_created"
|
||||
return out
|
||||
|
||||
order = eviction_order(db, adventure)
|
||||
active = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False))).scalars().all()
|
||||
on_lineage = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id == memory.id,
|
||||
lineage.path_of(db, adventure).clause(models.Memory))).scalar() is not None
|
||||
out["retained"] = {
|
||||
"yes": not memory.forgotten,
|
||||
"forgotten": memory.forgotten,
|
||||
"pinned": memory.pinned,
|
||||
"embedded": memory.embedded,
|
||||
"on_active_lineage": on_lineage,
|
||||
"use_count": memory.use_count,
|
||||
"last_used_at": str(memory.last_used_at) if memory.last_used_at else None,
|
||||
"created_at": str(memory.created_at),
|
||||
"active_memories": len(active),
|
||||
"memory_bank_capacity": settings.memory_bank_capacity,
|
||||
"eviction_position": (order.index(memory.id) + 1) if memory.id in order else None,
|
||||
"reason": ("evicted: marked forgotten by capacity eviction" if memory.forgotten
|
||||
else "active"),
|
||||
}
|
||||
if memory.forgotten:
|
||||
out["verdict"] = "created_but_evicted"
|
||||
return out
|
||||
|
||||
# The recall turn's AI node is excluded, so the newest action is the recall
|
||||
# player action, exactly as the turn saw it before it wrote its reply.
|
||||
query = production_query(adventure, recall_action.id)
|
||||
ranking = await rank_bank(db, adventure, settings, query, embed)
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory.id), None)
|
||||
stored_used = [m.get("id") for m in (snapshot.get("memories") or {}).get("used") or []]
|
||||
out["ranked"] = {
|
||||
"yes": row is not None and row["rank"] <= ranking["top_k"],
|
||||
"eligible": row is not None,
|
||||
"lexical_score": row["lexical_score"] if row else None,
|
||||
"semantic_score": row["semantic_score"] if row else None,
|
||||
"final_score": row["final_score"] if row else None,
|
||||
"selected_top_k": ranking["selected"],
|
||||
"pin_effect": "always selected" if memory.pinned else "none",
|
||||
"rank": row["rank"] if row else None,
|
||||
"of": len(ranking["scored"]),
|
||||
"top_k_cutoff": ranking["top_k"],
|
||||
"selected": bool(row and row["selected"]),
|
||||
"suppressed_as_duplicate_of": row["suppressed_as_duplicate_of"] if row else None,
|
||||
"query": query,
|
||||
"replica_matches_stored_selection": sorted(stored_used) == ranking["selected"],
|
||||
}
|
||||
if row is None or row["rank"] > ranking["top_k"] and not row["selected"]:
|
||||
out["verdict"] = "retained_but_not_ranked"
|
||||
return out
|
||||
if not row["selected"]:
|
||||
out["verdict"] = "ranked_but_not_selected"
|
||||
return out
|
||||
|
||||
section = _section(snapshot, MEMORIES_LABEL)
|
||||
injected = memory.id in stored_used and memory.text in section
|
||||
out["injected"] = {
|
||||
"yes": injected,
|
||||
"context_component": MEMORIES_LABEL,
|
||||
"in_stored_memories_used": memory.id in stored_used,
|
||||
"text_in_section": memory.text in section,
|
||||
"token_count": builder.count_tokens(section) if section else 0,
|
||||
}
|
||||
out["verdict"] = "injected" if injected else "selected_but_not_injected"
|
||||
return out
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- the stubs
|
||||
|
||||
@dataclass
|
||||
class BestCaseSummariser:
|
||||
"""The ideal memory writer: F survives if, and only if, F reached it.
|
||||
|
||||
A memory keeps every sentence of the excerpt that carries a planted fact, and
|
||||
adds one sentence naming the block's own distinct detail so memories differ.
|
||||
Summary updates never repeat a planted fact, so the summary layer stays out of
|
||||
the experiment. Every excerpt it was given is kept, for the creation-window
|
||||
diagnostic.
|
||||
"""
|
||||
|
||||
facts: tuple[Fact, ...] = (FACT_F, FACT_G)
|
||||
excerpts: list = field(default_factory=list)
|
||||
|
||||
async def complete(self, system, user, *, temperature=0.3, max_tokens=400):
|
||||
if "Current story summary:" in user:
|
||||
return "The travellers kept moving through the country around Westhaven."
|
||||
excerpt = user.split("Story excerpt:\n\n", 1)[-1].rsplit("\n\nMemory:", 1)[0]
|
||||
self.excerpts.append(excerpt)
|
||||
kept = [s.strip() for s in re.split(r"(?<=[.!?])\s+", excerpt)
|
||||
if any(f.carried_by(s) for f in self.facts)]
|
||||
detail = re.findall(r"\bat the ([a-z]+ [a-z]+)\b", excerpt.lower())
|
||||
tail = f"The travellers spent time at the {detail[-1]}." if detail else \
|
||||
"The travellers pressed on."
|
||||
return " ".join(dict.fromkeys(kept + [tail]))
|
||||
|
||||
|
||||
#: Words that mean the same thing to `ConceptEmbedder`. The point is only that a
|
||||
#: paraphrase lands near the original; the table is the model of that.
|
||||
CONCEPTS = {
|
||||
"timepiece": ("sundial", "dial", "hour", "hours", "clock", "timepiece"),
|
||||
"vessel": ("teapot", "pot", "kettle", "tea", "jar"),
|
||||
"hid": ("hid", "hide", "hidden", "slipped", "tucked", "put", "stashed"),
|
||||
"weathervane": ("weathervane", "vane"),
|
||||
"waterwheel": ("waterwheel", "wheel", "mill"),
|
||||
}
|
||||
_WORD_TO_CONCEPT = {w: c for c, words in CONCEPTS.items() for w in words}
|
||||
DIMENSIONS = 96
|
||||
|
||||
|
||||
@dataclass
|
||||
class ConceptEmbedder:
|
||||
"""A deterministic embedding: concepts in fixed dimensions, other words hashed."""
|
||||
|
||||
calls: int = 0
|
||||
|
||||
async def embed(self, texts):
|
||||
self.calls += 1
|
||||
return [self.vector(t) for t in texts]
|
||||
|
||||
@staticmethod
|
||||
def vector(text: str) -> list[float]:
|
||||
v = [0.0] * DIMENSIONS
|
||||
v[0] = 0.2 # every text shares a little, as real embeddings do
|
||||
concept_names = list(CONCEPTS)
|
||||
for word in re.findall(r"[a-z]+", text.lower()):
|
||||
concept = _WORD_TO_CONCEPT.get(word)
|
||||
if concept is not None:
|
||||
v[1 + concept_names.index(concept)] += 3.0
|
||||
elif len(word) > 3:
|
||||
bucket = int(hashlib.sha256(word.encode()).hexdigest(), 16)
|
||||
v[1 + len(concept_names) + bucket % (DIMENSIONS - 1 - len(concept_names))] += 1.0
|
||||
norm = math.sqrt(sum(x * x for x in v)) or 1.0
|
||||
return [x / norm for x in v]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- scenarios
|
||||
|
||||
#: Filler places. No word here is in `CONCEPTS`, and none names a planted fact.
|
||||
PLACES = (
|
||||
"north gate", "salt market", "ferry landing", "chapel steps", "rope walk",
|
||||
"fish stalls", "old bridge", "tanner yard", "lamp street", "weir path",
|
||||
"grain store", "boat yard", "watch house", "cloth hall", "eel traps",
|
||||
"sheep fold", "smith forge", "stone quay", "reed beds", "toll booth",
|
||||
)
|
||||
PARAPHRASE_QUERY = "I ask Mara where she tucked the little brass dial that tells the hour."
|
||||
UNRELATED_QUERY = "I ask the ferryman what rope costs at the landing this season."
|
||||
#: Built only from words every fixture memory holds ("travellers", "spent",
|
||||
#: "time"), so its rarity weight is zero everywhere.
|
||||
COMMON_WORDS_QUERY = "The travellers spent time."
|
||||
|
||||
|
||||
def filler_prose(index: int, words: int) -> str:
|
||||
"""Narration that moves on and never touches a planted fact."""
|
||||
place = PLACES[index % len(PLACES)]
|
||||
sentence = (f"At the {place} the travellers stopped, listened to the gulls over the "
|
||||
f"grey water, and talked about the long road north.")
|
||||
reps = max(1, round(words / len(sentence.split())))
|
||||
return " ".join([sentence] * reps)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Scenario:
|
||||
"""One deterministic campaign. Depths: the opening is 0, turn *n*'s player
|
||||
action is 2n-1 and its reply 2n."""
|
||||
|
||||
name: str
|
||||
turns: int = 52
|
||||
capacity: int = 80
|
||||
top_k: int = 5
|
||||
budget: int = 4096
|
||||
prose_words: int = 60
|
||||
plant_turn: int = 1
|
||||
recall_text: str = "I ask Mara where she hid the amber sundial."
|
||||
pin_first_memory: bool = False
|
||||
lineage_control: bool = False
|
||||
diagnose_recall: bool = True
|
||||
#: The narration of the last turn before recall, when a fixture needs the
|
||||
#: scene to say something (WP-B.2's context-dependent question). Must not
|
||||
#: name a planted fact.
|
||||
pre_recall_reply: str = ""
|
||||
|
||||
|
||||
SCENARIOS = {
|
||||
"independent_default": Scenario("independent_default"),
|
||||
"past_capacity": Scenario("past_capacity", capacity=6),
|
||||
"past_capacity_pinned": Scenario("past_capacity_pinned", capacity=6, pin_first_memory=True),
|
||||
# Closer to the shipped ratio (memory_top_k 5 against capacity 80): most of
|
||||
# the bank is not retrieved on a given turn.
|
||||
"past_capacity_low_top_k": Scenario("past_capacity_low_top_k", capacity=8, top_k=2),
|
||||
"long_block_fact_early": Scenario("long_block_fact_early", turns=10, prose_words=850,
|
||||
plant_turn=1, budget=16384),
|
||||
"long_block_fact_late": Scenario("long_block_fact_late", turns=10, prose_words=850,
|
||||
plant_turn=3, budget=16384),
|
||||
"lineage_control": Scenario("lineage_control", lineage_control=True),
|
||||
# v1.1 WP-B.2: the ranking failure B.1 saw on the real model, made
|
||||
# deterministic. Longer narration fills the v1.0.0 query, and `memory_top_k`
|
||||
# is the real run's 4. Below capacity, isolation valid.
|
||||
"ranking_crowded": Scenario("ranking_crowded", prose_words=150, top_k=4),
|
||||
# v1.1 WP-B.2: all three B.1 failures at once. Every block is longer than the
|
||||
# summariser's excerpt, narration crowds the query at `memory_top_k` 4, and
|
||||
# the bank passes a capacity of 8 long before recall at depth 106.
|
||||
"independent_full": Scenario("independent_full", prose_words=850, top_k=4,
|
||||
capacity=8, budget=16384),
|
||||
# v1.1 WP-B.2: a question that names neither the sundial nor the teapot and
|
||||
# cannot be answered without the scene. The last narration puts Mara at the
|
||||
# tavern's top shelf; the player asks "her" what she put "up there".
|
||||
"ranking_context_dependent": Scenario(
|
||||
"ranking_context_dependent", prose_words=150, top_k=4,
|
||||
recall_text="I ask her what she keeps up there.",
|
||||
pre_recall_reply=("Mara stands on a stool at the tavern's top shelf, running a cloth "
|
||||
"around the old kettle up there, and she will not meet your eye.")),
|
||||
}
|
||||
|
||||
|
||||
class ScriptNarrator:
|
||||
"""Stands in for the narrator: returns `next_reply`, with an empty state block."""
|
||||
|
||||
next_reply = ""
|
||||
last_usage = None
|
||||
prompts: list = []
|
||||
|
||||
def __init__(self, *a, **k):
|
||||
pass
|
||||
|
||||
async def generate(self, parts, *, temperature, max_tokens):
|
||||
ScriptNarrator.prompts.append((parts.system, parts.story))
|
||||
yield ("text", ScriptNarrator.next_reply)
|
||||
|
||||
|
||||
def run_scenario(scenario: Scenario) -> dict:
|
||||
"""Plays `scenario` through the real turn route and returns everything measured.
|
||||
|
||||
Uses the database `app.database` is already bound to, creating and dropping
|
||||
its tables, the way the suite's fixtures do. Patches are applied here and
|
||||
removed before returning, so this runs the same under pytest and from the CLI.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures as adventure_routes
|
||||
|
||||
summariser = BestCaseSummariser()
|
||||
embedder = ConceptEmbedder()
|
||||
patches = [
|
||||
(memorybank, "summary_provider", lambda s: summariser),
|
||||
(memorybank, "embedding_provider", lambda s: embedder),
|
||||
# Post-turn work is settled explicitly after each turn, so eviction
|
||||
# happens at a known point rather than whenever a background task runs.
|
||||
(memorybank, "schedule_post_turn", lambda adventure: None),
|
||||
(adventure_routes.turns, "OpenAICompatibleProvider", ScriptNarrator),
|
||||
(limits, "check_row_cap", lambda *a, **k: None),
|
||||
]
|
||||
saved = [(obj, name, getattr(obj, name)) for obj, name, _ in patches]
|
||||
for obj, name, value in patches:
|
||||
setattr(obj, name, value)
|
||||
ScriptNarrator.prompts = []
|
||||
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email=f"b1-{scenario.name}@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="script", endpoint_url="http://127.0.0.1:9/v1",
|
||||
embedding_model="concept-embed", context_token_budget=scenario.budget,
|
||||
max_output_tokens=500, memory_bank_capacity=scenario.capacity,
|
||||
memory_top_k=scenario.top_k,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title=f"B.1 {scenario.name}", memory_bank_enabled=True,
|
||||
auto_summarize=True, persona_name="Aldric",
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="Rain over Westhaven, and the tavern door banging in the wind."))
|
||||
db.commit()
|
||||
adv, user_id = adventure.id, user.id
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
client = TestClient(app)
|
||||
result: dict = {"scenario": scenario.__dict__.copy(), "trace": []}
|
||||
|
||||
def call(method, path, body=None, expect=200):
|
||||
response = client.request(method, f"/api/adventures/{adv}{path}", json=body)
|
||||
assert response.status_code == expect, (path, response.status_code, response.text[:300])
|
||||
return response.json() if response.content else None
|
||||
|
||||
def marks():
|
||||
with SessionLocal() as db:
|
||||
rows = db.execute(select(models.Memory.id, models.Memory.forgotten,
|
||||
models.Memory.embedded).where(
|
||||
models.Memory.adventure_id == adv)).all()
|
||||
summaries_n = db.query(models.Summary).filter_by(adventure_id=adv).count()
|
||||
return tuple(sorted(rows)), summaries_n
|
||||
|
||||
def settle():
|
||||
for _ in range(12):
|
||||
before = marks()
|
||||
asyncio.run(memorybank.run_post_turn(adv))
|
||||
if marks() == before:
|
||||
return
|
||||
|
||||
def memories():
|
||||
with SessionLocal() as db:
|
||||
return [dict(row._mapping) for row in db.execute(select(
|
||||
models.Memory.id, models.Memory.text, models.Memory.source_start,
|
||||
models.Memory.source_end, models.Memory.forgotten, models.Memory.pinned,
|
||||
models.Memory.use_count, models.Memory.last_used_at, models.Memory.branch_id,
|
||||
models.Memory.created_at).where(models.Memory.adventure_id == adv)
|
||||
.order_by(models.Memory.id)).all()]
|
||||
|
||||
def per_turn_isolation(adventure_id):
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adventure_id)
|
||||
active = summaries.current(db, adventure)
|
||||
return {
|
||||
"f_in_state": FACT_F.mentioned_by(json.dumps(adventure.narrative_state or {},
|
||||
default=str)),
|
||||
"f_in_summary": FACT_F.mentioned_by(active.text if active is not None else ""),
|
||||
}
|
||||
|
||||
def turn(kind, text, reply):
|
||||
ScriptNarrator.next_reply = f"{reply}\n```state\n{{\"events\": []}}\n```"
|
||||
response = client.post(f"/api/adventures/{adv}/actions", json={"type": kind, "text": text})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
assert '"type": "error"' not in response.text, response.text[-300:]
|
||||
|
||||
plant_depth = None
|
||||
f_memory_id = None
|
||||
pinned_id = None
|
||||
known: dict[int, dict] = {}
|
||||
g: dict = {}
|
||||
try:
|
||||
for n in range(1, scenario.turns + 1):
|
||||
if n == scenario.plant_turn:
|
||||
turn("story", FACT_F.sentence, filler_prose(n, scenario.prose_words))
|
||||
with SessionLocal() as db:
|
||||
plant_depth = db.query(models.Action.depth).filter_by(
|
||||
adventure_id=adv, text=FACT_F.sentence).scalar()
|
||||
elif scenario.lineage_control and n == 21:
|
||||
call("POST", "/checkpoints", {"name": "before the mill"}, expect=201)
|
||||
turn("story", FACT_G.sentence, filler_prose(n, scenario.prose_words))
|
||||
with SessionLocal() as db:
|
||||
g["plant_depth"] = db.query(models.Action.depth).filter_by(
|
||||
adventure_id=adv, text=FACT_G.sentence).scalar()
|
||||
elif scenario.lineage_control and n == 30:
|
||||
# Line A carries G's memory. Mark it, then abandon it: Undo back
|
||||
# to before G was planted and write something else.
|
||||
g["line_a"] = call("POST", "/checkpoints", {"name": "line A, after the mill"},
|
||||
expect=201)["id"]
|
||||
with SessionLocal() as db:
|
||||
g_rows = [m for m in db.execute(select(models.Memory).where(
|
||||
models.Memory.adventure_id == adv)).scalars() if FACT_G.carried_by(m.text)]
|
||||
g["memory_ids"] = [m.id for m in g_rows]
|
||||
with SessionLocal() as db:
|
||||
g["last_action_id_before_divergence"] = db.query(models.Action.id).filter_by(
|
||||
adventure_id=adv).order_by(models.Action.id.desc()).limit(1).scalar()
|
||||
for _ in range(9):
|
||||
call("POST", "/undo")
|
||||
turn("do", f"I turn away from the mill and walk to the {PLACES[n % len(PLACES)]}.",
|
||||
filler_prose(n + 100, scenario.prose_words))
|
||||
g["diverged_at_turn"] = n
|
||||
elif n == scenario.turns and scenario.pre_recall_reply:
|
||||
turn("do", "I head back to the tavern.", scenario.pre_recall_reply)
|
||||
else:
|
||||
turn("do", f"I walk on to the {PLACES[n % len(PLACES)]}.",
|
||||
filler_prose(n, scenario.prose_words))
|
||||
settle()
|
||||
|
||||
rows = memories()
|
||||
created = [r["id"] for r in rows if r["id"] not in known]
|
||||
newly_forgotten = [r["id"] for r in rows
|
||||
if r["forgotten"] and not known.get(r["id"], {}).get("forgotten")]
|
||||
for r in rows:
|
||||
known[r["id"]] = r
|
||||
if f_memory_id is None and plant_depth is not None:
|
||||
for r in rows:
|
||||
if (r["source_start"] is not None and r["source_start"] <= plant_depth
|
||||
<= r["source_end"] and FACT_F.carried_by(r["text"])):
|
||||
f_memory_id = r["id"]
|
||||
if scenario.pin_first_memory and pinned_id is None:
|
||||
candidate = next((r for r in rows if r["id"] != f_memory_id), None)
|
||||
if candidate is not None:
|
||||
call("PATCH", f"/memories/{candidate['id']}", {"pinned": True})
|
||||
pinned_id = candidate["id"]
|
||||
f_row = known.get(f_memory_id) if f_memory_id else None
|
||||
result["trace"].append({
|
||||
"turn": n,
|
||||
"active": sum(1 for r in rows if not r["forgotten"]),
|
||||
"total": len(rows),
|
||||
"created": created,
|
||||
"evicted": newly_forgotten,
|
||||
"created_and_evicted_same_turn": sorted(set(created) & set(newly_forgotten)),
|
||||
"f_memory_id": f_memory_id,
|
||||
"f_forgotten": bool(f_row and f_row["forgotten"]),
|
||||
"f_use_count": f_row["use_count"] if f_row else None,
|
||||
"coverage": coverage([(r["source_start"], r["source_end"]) for r in rows
|
||||
if not r["forgotten"] and r["source_start"] is not None],
|
||||
None),
|
||||
# Isolation on every turn, not only at recall (WP-B.2 acceptance).
|
||||
**per_turn_isolation(adv),
|
||||
})
|
||||
|
||||
turn("do", scenario.recall_text, filler_prose(999, scenario.prose_words))
|
||||
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv)
|
||||
settings = db.query(models.Settings).filter_by(user_id=user_id).first()
|
||||
recall_action = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
result["plant_depth"] = plant_depth
|
||||
result["recall_depth"] = recall_action.depth
|
||||
result["isolation"] = isolation(
|
||||
db, adventure, FACT_F, plant_depth,
|
||||
recall_snapshot=recall_action.context_snapshot,
|
||||
recall_depth=recall_action.depth)
|
||||
result["diagnosis"] = asyncio.run(diagnose(
|
||||
db, adventure, settings, FACT_F, plant_depth,
|
||||
recall_action=recall_action, embed=embedder.embed))
|
||||
result["summariser_excerpts"] = len(summariser.excerpts)
|
||||
# Provenance as the recall turn recorded it, resolved back to rows.
|
||||
used = (recall_action.context_snapshot.get("memories") or {}).get("used") or []
|
||||
f_entry = next((m for m in used
|
||||
if m.get("id") == result["diagnosis"]["created"]["memory_id"]), None)
|
||||
f_memory = db.get(models.Memory, f_entry["id"]) if f_entry else None
|
||||
block = memorybank.source_block(db, f_memory) if f_memory is not None else []
|
||||
result["provenance"] = {
|
||||
"recorded": f_entry and {k: f_entry.get(k) for k in
|
||||
("id", "source", "semantic_score", "lexical_score",
|
||||
"final_score", "authority")},
|
||||
"range_covers_plant": bool(f_entry and f_entry["source"]["source_start"]
|
||||
<= plant_depth <= f_entry["source"]["source_end"]),
|
||||
"matches_row": bool(f_memory is not None and f_entry["source"] == {
|
||||
"branch_id": f_memory.branch_id, "depth": f_memory.depth,
|
||||
"source_start": f_memory.source_start, "source_end": f_memory.source_end}),
|
||||
"source_block_depths": [a.depth for a in block],
|
||||
"source_block_holds_planting": any(a.text == FACT_F.sentence for a in block),
|
||||
}
|
||||
|
||||
memory_id = result["diagnosis"]["created"]["memory_id"]
|
||||
if memory_id is not None and not result["diagnosis"]["retained"]["forgotten"]:
|
||||
variants = {}
|
||||
base = production_query(adventure, recall_action.id)
|
||||
# WP-B.2's negative control needs a decoy: another memory that
|
||||
# holds a word the question adds, and nothing about F.
|
||||
decoy = next(((m.id, found.group(1)) for m in db.execute(
|
||||
select(models.Memory).where(models.Memory.adventure_id == adventure.id,
|
||||
models.Memory.forgotten.is_(False),
|
||||
models.Memory.id != memory_id)
|
||||
.order_by(models.Memory.id)).scalars()
|
||||
if (found := re.search(r"at the ([a-z]+ [a-z]+)\.", m.text or ""))), None)
|
||||
queries = [
|
||||
("direct", variant_query(base, scenario.recall_text)),
|
||||
("paraphrase", variant_query(base, PARAPHRASE_QUERY)),
|
||||
("unrelated", variant_query(base, UNRELATED_QUERY)),
|
||||
# The player's words with the context taken away.
|
||||
("input_only", {**variant_query(base, scenario.recall_text), "context": ""}),
|
||||
# Words every memory in these fixtures holds, and nothing else.
|
||||
("common_words", variant_query(base, COMMON_WORDS_QUERY)),
|
||||
]
|
||||
if decoy is not None:
|
||||
queries.append(("rare_word_with_paraphrase", variant_query(
|
||||
base, PARAPHRASE_QUERY[:-1] + f", out by the {decoy[1]}.")))
|
||||
for label, query in queries:
|
||||
ranking = asyncio.run(rank_bank(db, adventure, settings, query, embedder.embed))
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
|
||||
decoy_row = next((r for r in ranking["scored"]
|
||||
if decoy is not None and r["memory_id"] == decoy[0]), None)
|
||||
text = query["input"]
|
||||
variants[label] = {"query": text, "rank": row and row["rank"],
|
||||
"selected_count": len(ranking["selected"]),
|
||||
"decoy_memory_id": decoy and decoy[0],
|
||||
"decoy_rank": decoy_row and decoy_row["rank"],
|
||||
"decoy_lexical_score": decoy_row and decoy_row["lexical_score"],
|
||||
"of": len(ranking["scored"]),
|
||||
"similarity": row and row["similarity"],
|
||||
"lexical_score": row and row["lexical_score"],
|
||||
"final_score": row and row["final_score"],
|
||||
"selected": bool(row and row["selected"]),
|
||||
"top_k": ranking["top_k"]}
|
||||
result["ranking_variants"] = variants
|
||||
if memory_id is not None:
|
||||
result["f_first_used_turn"] = next(
|
||||
(t["turn"] for t in result["trace"] if (t["f_use_count"] or 0) > 0), None)
|
||||
result["f_last_use_increase_turn"] = max(
|
||||
(b["turn"] for a, b in zip(result["trace"], result["trace"][1:])
|
||||
if (b["f_use_count"] or 0) > (a["f_use_count"] or 0)), default=None)
|
||||
|
||||
evicted_turn = next((t["turn"] for t in result["trace"] if t["f_forgotten"]), None)
|
||||
first_evictions = next((t["evicted"] for t in result["trace"] if t["evicted"]), [])
|
||||
result["eviction"] = {
|
||||
"capacity": scenario.capacity,
|
||||
"f_evicted_at_turn": evicted_turn,
|
||||
"f_use_count_when_evicted": next(
|
||||
(t["f_use_count"] for t in result["trace"] if t["f_forgotten"]), None),
|
||||
"first_eviction_turn": next(
|
||||
(t["turn"] for t in result["trace"] if t["evicted"]), None),
|
||||
"first_evicted_ids": first_evictions,
|
||||
"f_memory_was_first_evicted": bool(f_memory_id and f_memory_id in first_evictions),
|
||||
"created_and_evicted_same_turn": sorted(
|
||||
{i for t in result["trace"] for i in t["created_and_evicted_same_turn"]}),
|
||||
"pinned_memory_id": pinned_id,
|
||||
"pinned_memory_forgotten": bool(pinned_id and known[pinned_id]["forgotten"]),
|
||||
}
|
||||
|
||||
if scenario.lineage_control:
|
||||
path_clause = lineage.path_of(db, adventure).clause(models.Memory)
|
||||
stored = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id.in_(g.get("memory_ids") or [-1]))).scalars().all()
|
||||
eligible = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id.in_(g.get("memory_ids") or [-1]), path_clause)).scalars().all()
|
||||
used_after = set()
|
||||
injected_text = False
|
||||
# Only turns played after the divergence. Before it, G was on the
|
||||
# active line, and a memory of it being used then is correct.
|
||||
for action in (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv,
|
||||
models.Action.type == "ai",
|
||||
models.Action.id > g["last_action_id_before_divergence"])
|
||||
.options(undefer(models.Action.context_snapshot))):
|
||||
snap = action.context_snapshot or {}
|
||||
for m in (snap.get("memories") or {}).get("used") or []:
|
||||
if m.get("id") in (g.get("memory_ids") or []):
|
||||
used_after.add(action.id)
|
||||
if FACT_G.mentioned_by(_section(snap, MEMORIES_LABEL)):
|
||||
injected_text = True
|
||||
g.update(stored=stored, eligible_on_active_line=eligible,
|
||||
turns_whose_memories_used_named_g=sorted(used_after),
|
||||
g_text_ever_in_used_memories=injected_text)
|
||||
if scenario.lineage_control:
|
||||
call("POST", f"/checkpoints/{g['line_a']}/restore")
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv)
|
||||
eligible = db.execute(select(models.Memory.id).where(
|
||||
models.Memory.id.in_(g.get("memory_ids") or [-1]),
|
||||
lineage.path_of(db, adventure).clause(models.Memory))).scalars().all()
|
||||
g["eligible_after_returning_to_line_a"] = eligible
|
||||
result["lineage_control"] = g
|
||||
return result
|
||||
finally:
|
||||
for obj, name, value in saved:
|
||||
setattr(obj, name, value)
|
||||
app.dependency_overrides.clear()
|
||||
adventure_routes.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
@@ -0,0 +1,551 @@
|
||||
"""v1.1 WP-B.2: how faithfully a real model's memories keep the facts of their block.
|
||||
|
||||
**Diagnostic only. Nothing here is imported by the application, and nothing here
|
||||
is a release gate.**
|
||||
|
||||
Why it exists: the first isolation-valid real-model WP-B run failed at creation.
|
||||
The whole planting block reached the summariser, and the memory it wrote left the
|
||||
planted fact out (`V1.1-WP-B2-REPORT.md` §L.2). B2.4 tried the plan's bounded
|
||||
remedy, a memory prompt instructing the model to keep named facts and objects. It
|
||||
was measured with this module, did not correct the failure, and was **not
|
||||
shipped** (§T). The module stays so the limitation can be measured again, on this
|
||||
model or a different one.
|
||||
|
||||
- **Fixtures.** Short story blocks, each built around a fact a later scene could
|
||||
turn on, with the ordinary texture a real block carries around it. They are
|
||||
genre-neutral (office, contemporary, a science-fiction-neutral station), plus
|
||||
the attempt-2 planting block itself. Each names the facts a memory must keep,
|
||||
whom each belongs to, and what it must not invent.
|
||||
- **The checker** (`evaluate`) is deterministic and reads only the memory text.
|
||||
It is a heuristic, and says so:
|
||||
- a fact counts as kept when one sentence names every part of it;
|
||||
- attribution is the nearest named character before the fact's verb;
|
||||
- it also reports word count, a leading "Memory:" and second-person "you".
|
||||
- **The comparison** (`compare`) sends each fixture to a real model through the
|
||||
application's own provider and `memorybank.memory_user_prompt`. It scores
|
||||
memories under the shipped prompt and under the rejected B2.4 experiment.
|
||||
|
||||
A scripted summariser cannot show what a prompt makes a model do. So the
|
||||
deterministic tests prove only that the checker is right and the fixtures reach
|
||||
the summariser; model quality is measured here, with inference, and reported.
|
||||
|
||||
# the real-model measurement (inference: ask first)
|
||||
.venv/bin/python -m tools.memory_fidelity --endpoint <v1 URL> \\
|
||||
--model qwen2.5:3b-instruct-16k --samples 5 --out "$HOME/v11-evidence/<label>"
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
#: The memory prompt WP-B.2 B2.4 tried and **rejected**. Kept verbatim so the
|
||||
#: experiment in `V1.1-WP-B2-REPORT.md` §T can be repeated; the application
|
||||
#: never uses it.
|
||||
B24_EXPERIMENT_PROMPT = (
|
||||
"You compress interactive-fiction story excerpts into memories. Respond with "
|
||||
"1-2 plain sentences in past tense, in at most 50 words.\n\n"
|
||||
"Keep the concrete facts a later scene could turn on: specific people, "
|
||||
"objects and places; where something is; who has, hid, found, knows, saw or "
|
||||
"promised what; injuries, clues and commitments. A distinctive fact comes "
|
||||
"before mood, scenery, routine movement and small talk. Drop those first, "
|
||||
"however much of the excerpt they fill, and never let a later passage crowd "
|
||||
"out an earlier fact.\n\n"
|
||||
"Keep each fact with the person it belongs to. Never move an action, promise, "
|
||||
"possession, statement or piece of knowledge from one character to another, "
|
||||
"and never add a fact the excerpt does not state.\n\n"
|
||||
'Write in the third person. The narration calls the protagonist "you". The '
|
||||
'protagonist\'s own actions and words are the lines that begin with ">", '
|
||||
'written as "I" or "You", and what they establish is part of the story just '
|
||||
"as the narration is. The protagonist is named in the Cast: refer to them by "
|
||||
'that name, never as "you" or "I". If the Cast gives no name for them, call '
|
||||
'them "the player". Name the other characters too rather than writing "he", '
|
||||
'"she" or "they" on their own — this memory will be read on its own, much '
|
||||
"later, with nothing around it to say who a pronoun meant.\n\n"
|
||||
"No preamble, no commentary."
|
||||
)
|
||||
|
||||
TARGET_WORDS = 50
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Fact:
|
||||
"""One fact a memory must keep.
|
||||
|
||||
`groups`: every group must be matched in one sentence, by any of its terms.
|
||||
`verbs`: the relation. When one is in that sentence, the nearest named
|
||||
character before it is who the memory says the fact belongs to.
|
||||
`actor`: whom it belongs to. `None` for a fact with no owner.
|
||||
"""
|
||||
|
||||
name: str
|
||||
groups: tuple[tuple[str, ...], ...]
|
||||
actor: str | None = None
|
||||
verbs: tuple[str, ...] = ()
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Fixture:
|
||||
fixture_id: str
|
||||
genre: str
|
||||
requirement: str # which measurement this fixture serves
|
||||
protagonist: str
|
||||
others: tuple[str, ...]
|
||||
actions: tuple[tuple[str, str], ...] # (type, text), oldest first
|
||||
facts: tuple[Fact, ...]
|
||||
min_facts: int | None = None # default: all
|
||||
#: Regexes a faithful memory must not match. Each is anchored on the wrong
|
||||
#: character as the subject ("Marcus promised"), because a faithful memory
|
||||
#: may name that character elsewhere in the same sentence ("promised Marcus").
|
||||
forbidden: tuple[str, ...] = ()
|
||||
#: Hand-written memories for the checker's own tests: one that should pass,
|
||||
#: and failures that should not, each with the reason it must report.
|
||||
faithful: str = ""
|
||||
unfaithful: tuple[tuple[str, str], ...] = field(default_factory=tuple)
|
||||
|
||||
@property
|
||||
def cast(self) -> tuple[str, ...]:
|
||||
return (self.protagonist, *self.others)
|
||||
|
||||
@property
|
||||
def raw(self) -> str:
|
||||
return "\n\n".join(text for _, text in self.actions)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ checker
|
||||
|
||||
def _term(term: str) -> re.Pattern:
|
||||
body = r"\s+".join(re.escape(part) for part in term.lower().split())
|
||||
return re.compile(rf"(?<![a-z]){body}(?:s|es|ed|d)?(?![a-z])")
|
||||
|
||||
|
||||
def _sentences(text: str) -> list[str]:
|
||||
return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text or "") if s.strip()]
|
||||
|
||||
|
||||
def _first(low: str, terms) -> int | None:
|
||||
found = [m.start() for t in terms for m in [_term(t).search(low)] if m]
|
||||
return min(found) if found else None
|
||||
|
||||
|
||||
def _attributed_to(sentence: str, fact: Fact, cast: tuple[str, ...]) -> str | None:
|
||||
"""The character this sentence gives the fact to, or None if it names none."""
|
||||
low = sentence.lower()
|
||||
at = _first(low, fact.verbs) if fact.verbs else None
|
||||
names = [(m.start(), name) for name in cast for m in _term(name).finditer(low)]
|
||||
if at is not None:
|
||||
before = [(pos, name) for pos, name in names if pos < at]
|
||||
if before:
|
||||
return max(before)[1]
|
||||
return min(names)[1] if names else None
|
||||
|
||||
|
||||
def evaluate(fixture: Fixture, memory: str) -> dict:
|
||||
"""What a memory kept, whom it gave each fact to, what it invented, and how
|
||||
it is framed."""
|
||||
memory = memory or ""
|
||||
sentences = _sentences(memory)
|
||||
facts = {}
|
||||
for fact in fixture.facts:
|
||||
holding = [s for s in sentences
|
||||
if all(_first(s.lower(), group) is not None for group in fact.groups)]
|
||||
owners = sorted({o for s in holding
|
||||
if (o := _attributed_to(s, fact, fixture.cast)) is not None})
|
||||
facts[fact.name] = {
|
||||
"kept": bool(holding),
|
||||
"attributed_to": owners,
|
||||
"attribution_ok": (fact.actor is None or not holding
|
||||
or (owners != [] and set(owners) == {fact.actor})),
|
||||
}
|
||||
kept = sum(1 for f in facts.values() if f["kept"])
|
||||
needed = len(fixture.facts) if fixture.min_facts is None else fixture.min_facts
|
||||
inventions = [p for p in fixture.forbidden if re.search(p, memory, re.I)]
|
||||
words = len(memory.split())
|
||||
misattributed = [name for name, f in facts.items() if f["kept"] and not f["attribution_ok"]]
|
||||
return {
|
||||
"facts": facts,
|
||||
"kept": kept,
|
||||
"needed": needed,
|
||||
"retained": kept >= needed,
|
||||
"misattributed": misattributed,
|
||||
"inventions": inventions,
|
||||
"words": words,
|
||||
"over_target": words > TARGET_WORDS,
|
||||
# Framing the shipped prompt's rules exist to prevent.
|
||||
"memory_prefix": memory.lstrip().lower().startswith("memory:"),
|
||||
"second_person": re.search(r"\byou(r|rs|rself)?\b", memory, re.I) is not None,
|
||||
"passed": kept >= needed and not misattributed and not inventions,
|
||||
}
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- fixtures
|
||||
|
||||
OFFICE_TEXTURE = (
|
||||
"The open-plan floor hums with keyboards and the air conditioning rattles in "
|
||||
"its vent. Someone has left a birthday card on the printer, and the coffee "
|
||||
"machine gurgles through another pot."
|
||||
)
|
||||
|
||||
FIXTURES: tuple[Fixture, ...] = (
|
||||
Fixture(
|
||||
fixture_id="object_place_office",
|
||||
genre="office",
|
||||
requirement="distinctive object and place",
|
||||
protagonist="Dana", others=("Priya",),
|
||||
actions=(
|
||||
("start", "Monday at the insurance office. Dana is covering the late shift."),
|
||||
("do", "> You check the queue of unanswered claims."),
|
||||
("ai", f"{OFFICE_TEXTURE} Priya walks past your desk carrying a stack of folders, "
|
||||
"and slides the red backup drive into the bottom drawer of the grey filing "
|
||||
"cabinet in the archive room before locking it. The phones ring twice and stop. "
|
||||
"Rain streaks the tall windows while the floor slowly empties."),
|
||||
("do", "> You ask Priya whether the audit is still on for Thursday."),
|
||||
("ai", "Priya shrugs, says nobody tells her anything, and goes back to her own desk. "
|
||||
"The cleaners arrive with their carts and the lights dim on a timer."),
|
||||
("do", "> You log off and pack your bag."),
|
||||
),
|
||||
facts=(Fact("drive in the cabinet", (("backup drive", "drive"), ("drawer", "filing cabinet", "cabinet")),
|
||||
actor="Priya", verbs=("slid", "slide", "put", "placed", "locked", "hid", "stored", "left")),),
|
||||
forbidden=(r"\bdana\s+(had\s+)?(took|takes|has|holds|held|stole|locked|slid|hid)\b[^.]*\bdrive\b",),
|
||||
faithful="Priya locked the red backup drive in the bottom drawer of the grey filing cabinet "
|
||||
"in the archive room while Dana covered the late shift.",
|
||||
unfaithful=(
|
||||
("Dana covered a quiet late shift at the insurance office while rain fell and the "
|
||||
"cleaners arrived.", "not retained"),
|
||||
("Dana locked the red backup drive in the bottom drawer of the filing cabinet.",
|
||||
"misattributed"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
fixture_id="player_fact_station",
|
||||
genre="science-fiction-neutral",
|
||||
requirement="player-established concrete fact",
|
||||
protagonist="Reyes", others=("Okafor",),
|
||||
actions=(
|
||||
("start", "Deck four of the relay station, halfway through the night cycle. Reyes is on maintenance duty."),
|
||||
("do", "> You walk the corridor checking the pressure seals."),
|
||||
("ai", "The corridor lights pulse a dim blue. Condensation beads on the pipes, and "
|
||||
"somewhere below a pump cycles on with a shudder. Chief Okafor passes with a "
|
||||
"tablet under one arm and nods without stopping."),
|
||||
("do", "> I watch Chief Okafor seal the coolant sample in locker nine and log it under a false name."),
|
||||
("ai", "The night cycle drags on. The ventilation hisses, a door chimes somewhere "
|
||||
"down the ring, and the viewport shows the same slow turn of stars it always "
|
||||
"does. You finish the seal checks and sign the maintenance sheet, and the "
|
||||
"corridor settles back into its usual hum."),
|
||||
("do", "> You head back to your bunk."),
|
||||
),
|
||||
facts=(Fact("sample in locker nine", (("coolant sample", "sample"), ("locker",)),
|
||||
actor="Okafor", verbs=("seal", "sealed", "locked", "put", "stored", "hid", "logged", "placed")),),
|
||||
forbidden=(r"\breyes\s+(had\s+)?(sealed|seals|hid|stored|locked|logged)\b[^.]*\bsample\b",),
|
||||
faithful="Reyes saw Chief Okafor seal the coolant sample in locker nine and log it under a false name.",
|
||||
unfaithful=(
|
||||
("Reyes finished the seal checks on deck four during a quiet night cycle.", "not retained"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
fixture_id="promise_contemporary",
|
||||
genre="contemporary",
|
||||
requirement="promise / commitment",
|
||||
protagonist="Dana", others=("Marcus",),
|
||||
actions=(
|
||||
("start", "A Saturday afternoon at the flat Dana is about to rent from Marcus."),
|
||||
("do", "> You look around the empty living room."),
|
||||
("ai", "Sunlight falls across bare floorboards. The radiator ticks, a neighbour's "
|
||||
"radio plays through the wall, and Marcus jingles a ring of keys while he "
|
||||
"talks about the boiler and the bins."),
|
||||
("do", '> You say "Marcus, I will bring you the signed lease by Friday."'),
|
||||
("ai", "Marcus nods and writes something on the back of an envelope. Outside a bus "
|
||||
"pulls away, a dog barks twice, and the afternoon light moves slowly up the wall."),
|
||||
("do", "> You thank him and leave."),
|
||||
),
|
||||
facts=(Fact("lease by Friday", (("lease",), ("friday",)), actor="Dana",
|
||||
verbs=("promise", "promised", "bring", "agreed", "said", "would")),),
|
||||
forbidden=(r"\bmarcus\s+(promised|agreed|will\s+bring|would\s+bring)\b[^.]*\blease\b",),
|
||||
faithful="Dana promised Marcus she would bring him the signed lease for the flat by Friday.",
|
||||
unfaithful=(
|
||||
("Marcus promised to bring Dana the signed lease by Friday.", "misattributed"),
|
||||
("Dana viewed the empty flat on a sunny Saturday while Marcus talked about the boiler.",
|
||||
"not retained"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
fixture_id="attribution_station",
|
||||
genre="science-fiction-neutral",
|
||||
requirement="attribution",
|
||||
protagonist="Reyes", others=("Lena", "Tomas"),
|
||||
actions=(
|
||||
("start", "The survey ship's cargo bay, between jumps."),
|
||||
("do", "> You ask who can open the sealed vault."),
|
||||
("ai", "Lena folds her arms. She is the only one aboard who knows the vault door "
|
||||
"code, and she makes it clear she is keeping it to herself. Tomas taps the "
|
||||
"access badge clipped to his jacket; without it the bay lift will not move."),
|
||||
("do", "> You look from one of them to the other."),
|
||||
("ai", "The bay lights flicker as the drive spools. Crates creak against their "
|
||||
"straps, and the air smells of cold metal and oil."),
|
||||
("do", "> You wait for one of them to speak."),
|
||||
),
|
||||
facts=(
|
||||
Fact("code", (("code",),), actor="Lena", verbs=("knows", "knew", "keeps", "kept", "holds", "held")),
|
||||
Fact("badge", (("badge",),), actor="Tomas",
|
||||
verbs=("carries", "carried", "has", "had", "holds", "held", "wore", "wears", "tapped", "taps")),
|
||||
),
|
||||
forbidden=(r"\breyes\b[^.]*\b(knew|knows)\b[^.]*\bcode\b",),
|
||||
faithful="Lena alone knew the vault door code and kept it to herself; Tomas carried the access "
|
||||
"badge that the bay lift needed.",
|
||||
unfaithful=(
|
||||
("Tomas knew the vault door code, and Lena carried the access badge.", "misattributed"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
fixture_id="clutter_office",
|
||||
genre="office",
|
||||
requirement="clutter pressure",
|
||||
protagonist="Dana", others=("Priya", "Owen"),
|
||||
actions=(
|
||||
("start", "The quarterly offsite at a conference hotel by the motorway."),
|
||||
("do", "> You find a seat near the back."),
|
||||
("ai", "The conference room smells of carpet cleaner and burnt coffee. Chairs scrape, "
|
||||
"a projector fan whines, and someone at the front struggles with the clicker. "
|
||||
"Owen talks about his weekend at length, the traffic on the ring road, a new "
|
||||
"sandwich place, the football, and whether it will rain for the barbecue. The "
|
||||
"slides cycle through charts nobody reads. Outside the window lorries hiss past "
|
||||
"on the wet motorway, and the hotel's muzak drifts in whenever the door opens."),
|
||||
("do", "> You go to the refreshment table."),
|
||||
("ai", "Pastries sweat under cling film. Priya stirs her tea, glances around, and "
|
||||
"quietly tells you that she saw Owen shred the signed supplier contract in the "
|
||||
"copy room last night. Then she talks about the weather, the parking, and the "
|
||||
"long drive home, and laughs at a joke from across the room. The afternoon "
|
||||
"session is announced, people drift back to their seats, and the projector "
|
||||
"fan starts whining again over a long talk about quarterly targets."),
|
||||
("do", "> You take your seat for the afternoon session."),
|
||||
),
|
||||
facts=(Fact("contract shredded", (("contract",), ("shred", "shredded", "destroyed")),
|
||||
actor="Owen", verbs=("shred", "shredded", "destroyed")),),
|
||||
forbidden=(r"\b(priya|dana)\b\s+(had\s+)?(shred|shredded|destroyed)\b",),
|
||||
faithful="At the offsite, Priya told Dana she had seen Owen shred the signed supplier contract "
|
||||
"in the copy room the night before.",
|
||||
unfaithful=(
|
||||
("Dana sat through a dull offsite of charts, pastries and Owen's talk about the weekend.",
|
||||
"not retained"),
|
||||
("Priya shredded the signed supplier contract in the copy room.", "misattributed"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
fixture_id="no_invention_office",
|
||||
genre="office",
|
||||
requirement="no invention",
|
||||
protagonist="Dana", others=("Owen",),
|
||||
actions=(
|
||||
("start", "A short planning meeting in the small room on the third floor."),
|
||||
("do", "> You sit down opposite Owen."),
|
||||
("ai", "A black briefcase sits unclaimed by the door; nobody mentions it. Owen says "
|
||||
"the budget review has moved from Tuesday to Thursday, and asks you to tell "
|
||||
"the team."),
|
||||
("do", "> You agree to pass it on."),
|
||||
("ai", "Owen thanks you, checks his phone, and the meeting ends after ten minutes. "
|
||||
"The briefcase is still by the door when you leave."),
|
||||
("do", "> You walk back to your desk."),
|
||||
),
|
||||
facts=(Fact("review moved", (("budget review", "review"), ("thursday",))),),
|
||||
forbidden=(
|
||||
r"\b(took|taken|stole|hid|hidden|grabbed|pocketed|carried|carries|owns|owned|belong\w*)\b[^.]*\bbriefcase\b",
|
||||
r"\bbriefcase\b[^.]*\b(belong\w*|his|her|owen's|dana's|secret|clue)\b",
|
||||
r"\b(clue|secret|password|code)\b",
|
||||
),
|
||||
faithful="Owen told Dana the budget review had moved from Tuesday to Thursday, and Dana agreed "
|
||||
"to tell the team.",
|
||||
unfaithful=(
|
||||
("Owen told Dana the budget review had moved to Thursday and left his secret briefcase by the door.",
|
||||
"invented"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
fixture_id="multiple_facts_station",
|
||||
genre="science-fiction-neutral",
|
||||
requirement="multiple concrete facts",
|
||||
protagonist="Reyes", others=("Hale", "Varga", "Moreau"),
|
||||
actions=(
|
||||
("start", "The mess hall of the mining outpost after the shift change."),
|
||||
("do", "> You sit with the day crew."),
|
||||
("ai", "Trays clatter and the recycler drones. Engineer Hale admits, half joking, that "
|
||||
"she hid the spare fuse inside the airlock control panel. Doctor Varga says only "
|
||||
"she knows the reactor override phrase, and changes the subject. Pilot Moreau "
|
||||
"grumbles that he owes Hale two shifts of cover."),
|
||||
("do", "> You finish your meal."),
|
||||
("ai", "The lights dim for the rest cycle and the crew drifts off to their bunks."),
|
||||
("do", "> You head to your quarters."),
|
||||
),
|
||||
facts=(
|
||||
Fact("fuse in the panel", (("fuse",), ("panel", "airlock")), actor="Hale",
|
||||
verbs=("hid", "hidden", "put", "stashed", "kept")),
|
||||
Fact("override phrase", (("override",),), actor="Varga", verbs=("knows", "knew", "keeps")),
|
||||
Fact("owes cover", (("owes", "owed", "debt"), ("shift", "cover")), actor="Moreau",
|
||||
verbs=("owes", "owed")),
|
||||
),
|
||||
min_facts=2,
|
||||
faithful="Hale hid the spare fuse inside the airlock control panel, and Varga alone knew the "
|
||||
"reactor override phrase.",
|
||||
unfaithful=(
|
||||
("Varga hid the spare fuse inside the airlock panel, and Hale knew the reactor override phrase.",
|
||||
"misattributed"),
|
||||
),
|
||||
),
|
||||
Fixture(
|
||||
# The planting block of WP-B.2 real-model attempt 2, verbatim: the
|
||||
# harness's own fixture campaign, with no identifier in it. It is the
|
||||
# failure the reference model has not overcome.
|
||||
fixture_id="regression_attempt_2",
|
||||
genre="fantasy (the harness campaign)",
|
||||
requirement="the actual failed-run block",
|
||||
protagonist="Aldric", others=("Mara", "Edrin"),
|
||||
actions=(
|
||||
("start", "Rain over Westhaven. Aldric sits in the Crooked Lantern with a silver key in his "
|
||||
"pocket and no-one to give it to."),
|
||||
("do", "> I tell Mara quietly that the silver key opens the crypt beneath the Old Abbey — "
|
||||
"SILVER-KEY-CRYPT-OLD-ABBEY."),
|
||||
("ai", "The lantern's warmth feels cold against Aldric's hands as he fidgets with the silver "
|
||||
"key. Mara observes him without a word, her eyes unreadable. The rain outside "
|
||||
"persists, a backdrop of nature's indifference to the human drama transpiring. "
|
||||
"Aldric's fingers close over the key, the familiar weight making him feel both more "
|
||||
"alone and more determined. He pulls the key from his pocket, the metal feeling cold "
|
||||
"against his skin. \n\n> \"I have the key,\" Aldric whispers to Mara. \"It leads to "
|
||||
"the crypt beneath the abbey.\"\n\nEdrin, having finished his ale, stands. \"Well, now "
|
||||
"that we know the key leads to somewhere... do you think we should see what it "
|
||||
"opens?\"\n\nThe silver key feels heavy in Aldric's hands. He looks at Mara, then at "
|
||||
"the key, and finally at Edrin. He decides. \"Let's do this.\""),
|
||||
("do", "> I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top "
|
||||
"shelf, and she makes me promise to tell no one."),
|
||||
("ai", "Aldric steps closer to the smoky bar, his gaze shifting from the silver key in his "
|
||||
"pocket to Mara. \"Let's not overthink this,\" he says, his voice steady despite the "
|
||||
"gnawing doubt that tugs at him. \"Let's see what the crypt holds.\"\n\nEdrin raises "
|
||||
"an eyebrow, his curiosity piqued. \"Alright, then. But if this leads somewhere "
|
||||
"dangerous, we're splitting up. If we make it back here, I'll tell no one what we "
|
||||
"found.\"\n\nMara nods her agreement, her eyes still unreadable. \"Agreed. But "
|
||||
"remember, we're not the first to look there. We need to tread carefully.\"\n\nWith a "
|
||||
"slight nod, they step out of the tavern. The rain has only intensified, pelting them "
|
||||
"as they walk toward the Old Abbey's shadowed grounds. The crypt, sealed since its "
|
||||
"founding, stands before them, its door ajar, inviting or perhaps foreboding. As "
|
||||
"Aldric approaches, he feels a strange mixture of excitement and fear, his hands "
|
||||
"curling around the silver key's handle. \n\n> He turns the key, and the lock clicks "
|
||||
"open with a satisfying, if ominous, hiss. The door swings inward, revealing a narrow "
|
||||
"passageway. The air inside is musty and cold, the only light filtering in through "
|
||||
"the cracks. \n\nAldric takes a tentative step inside, his breath visible in the air. "
|
||||
"Mara and Edrin follow, their footsteps echoing in the silence. The walls are lined "
|
||||
"with stone and cobwebs, and the smell of damp and decay is overwhelming. They "
|
||||
"advance cautiously, each of them aware of the potential dangers lurking within."),
|
||||
("do", "> I ask Mara what she has heard about the abbey."),
|
||||
),
|
||||
facts=(Fact("sundial in the teapot", (("sundial",), ("teapot",)), actor="Mara",
|
||||
verbs=("slip", "slipped", "hid", "hidden", "put", "placed", "tucked", "stashed")),),
|
||||
forbidden=(
|
||||
r"\b(aldric|edrin)\b\s+(had\s+)?(slipped|hid|placed|put|tucked|stashed)\b[^.]*\bsundial\b",
|
||||
),
|
||||
faithful="Aldric watched Mara slip the amber sundial inside the cracked teapot on the tavern's "
|
||||
"top shelf and promised her to tell no one; then Aldric, Mara and Edrin entered the "
|
||||
"crypt beneath the Old Abbey.",
|
||||
unfaithful=(
|
||||
# The memory attempt 2 actually stored, verbatim.
|
||||
("Aldric sits in the Crooked Lantern with a silver key in his pocket, no-one to give it to. "
|
||||
"Mara observed him quietly, her eyes unreadable. Edrin finished his ale and stood, asking if "
|
||||
"they should see what the crypt beneath the Old Abbey holds. Aldric decided to go, promising "
|
||||
"not to tell anyone. They walked to the Old Abbey's grounds, the crypt door ajar, inviting "
|
||||
"and foreboding. Inside, the air was musty and cold, with the smell of damp and decay. They "
|
||||
"advanced cautiously, each aware of potential dangers. The silver key, the key to the crypt, "
|
||||
"felt heavy in Aldric's hands.", "not retained"),
|
||||
("Aldric slipped the amber sundial inside the cracked teapot on the tavern's top shelf.",
|
||||
"misattributed"),
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
FIXTURES_BY_ID = {f.fixture_id: f for f in FIXTURES}
|
||||
|
||||
|
||||
def cast_brief_for(fixture: Fixture) -> str:
|
||||
"""The cast brief `memorybank.cast_brief` would build for this fixture."""
|
||||
from app import memorybank
|
||||
|
||||
lines = [memorybank._cast_line(fixture.protagonist, "", protagonist=True)]
|
||||
lines += [memorybank._cast_line(name, "") for name in fixture.others]
|
||||
return "Cast:\n" + "\n".join(lines)
|
||||
|
||||
|
||||
def user_prompt_for(fixture: Fixture) -> str:
|
||||
"""Exactly the user message the application sends for this block."""
|
||||
from app import memorybank
|
||||
|
||||
return memorybank.memory_user_prompt(cast_brief_for(fixture), memorybank.memory_excerpt(fixture.raw))
|
||||
|
||||
|
||||
# --------------------------------------------------------- the measurement
|
||||
|
||||
async def compare(endpoint: str, model: str, samples: int) -> dict:
|
||||
"""Every fixture, `samples` times, under the shipped prompt and the B2.4 experiment."""
|
||||
from app import memorybank, models
|
||||
|
||||
settings = models.Settings(endpoint_url=endpoint, model=model, summary_model="",
|
||||
api_mode=models.Settings.__table__.c.api_mode.default.arg,
|
||||
model_timeout_seconds=300)
|
||||
provider = memorybank.summary_provider(settings)
|
||||
arms = {"shipped": memorybank.MEMORY_SYSTEM_PROMPT, "b2.4-experiment": B24_EXPERIMENT_PROMPT}
|
||||
out: dict = {"endpoint_model": model, "samples": samples, "fixtures": {}}
|
||||
for fixture in FIXTURES:
|
||||
user = user_prompt_for(fixture)
|
||||
row: dict = {"requirement": fixture.requirement, "genre": fixture.genre, "arms": {}}
|
||||
for arm, system in arms.items():
|
||||
runs = []
|
||||
for _ in range(samples):
|
||||
text = (await provider.complete(system, user) or "").strip()
|
||||
runs.append({"memory": text, **evaluate(fixture, text)})
|
||||
summary = {
|
||||
"passed": sum(r["passed"] for r in runs),
|
||||
"retained": sum(r["retained"] for r in runs),
|
||||
"misattributed": sum(bool(r["misattributed"]) for r in runs),
|
||||
"invented": sum(bool(r["inventions"]) for r in runs),
|
||||
"memory_prefix": sum(r["memory_prefix"] for r in runs),
|
||||
"second_person": sum(r["second_person"] for r in runs),
|
||||
"words_median": sorted(r["words"] for r in runs)[len(runs) // 2],
|
||||
"words_max": max(r["words"] for r in runs),
|
||||
"over_target": sum(r["over_target"] for r in runs),
|
||||
}
|
||||
row["arms"][arm] = {"runs": runs, **summary}
|
||||
print(f"{fixture.fixture_id:28} {arm:16} "
|
||||
+ " ".join(f"{k} {v}" for k, v in summary.items()), flush=True)
|
||||
out["fixtures"][fixture.fixture_id] = row
|
||||
return out
|
||||
|
||||
|
||||
def main() -> int:
|
||||
import argparse
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--endpoint", required=True)
|
||||
parser.add_argument("--model", required=True)
|
||||
parser.add_argument("--samples", type=int, default=5)
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
handle = tempfile.NamedTemporaryFile(suffix=".db", delete=False)
|
||||
handle.close()
|
||||
os.environ["AIDND_DB_PATH"] = handle.name # nothing is written; never the real database
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
try:
|
||||
result = asyncio.run(compare(args.endpoint, args.model, args.samples))
|
||||
finally:
|
||||
Path(handle.name).unlink(missing_ok=True)
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
(out / "fidelity.json").write_text(json.dumps(result, indent=2))
|
||||
print(f"written to {out / 'fidelity.json'}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,125 @@
|
||||
"""v1.1 WP-B.1: run the deterministic memory-retention scenarios, or diagnose a real campaign.
|
||||
|
||||
# the deterministic scenarios, against an isolated database in --out
|
||||
.venv/bin/python -m tools.v11_b1_memory scenarios --out "$HOME/v11-evidence/b1/<label>"
|
||||
|
||||
# the four stages for a finished real campaign (reads its database; embeds
|
||||
# the recall query with the campaign's own configured embedding model)
|
||||
AIDND_TEST_ENDPOINT=... AIDND_TEST_EMBED_MODEL=nomic-embed-text:latest \\
|
||||
.venv/bin/python -m tools.v11_b1_memory diagnose --db <campaign.db> \\
|
||||
--plant-depth 3 --out "$HOME/v11-evidence/b1/<label>"
|
||||
|
||||
Run from `backend/`. Nothing here changes memory behaviour; see
|
||||
`tools/memory_diagnostic.py` for what is measured and what the stubs model.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
sub = parser.add_subparsers(dest="command", required=True)
|
||||
scen = sub.add_parser("scenarios")
|
||||
scen.add_argument("--out", required=True)
|
||||
scen.add_argument("--only", action="append", default=[])
|
||||
diag = sub.add_parser("diagnose")
|
||||
diag.add_argument("--db", required=True)
|
||||
diag.add_argument("--plant-depth", type=int, required=True)
|
||||
diag.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
if args.command == "scenarios":
|
||||
db_path = out / "scenarios.db"
|
||||
if db_path.exists():
|
||||
db_path.unlink()
|
||||
os.environ["AIDND_DB_PATH"] = str(db_path)
|
||||
else:
|
||||
# A copy, so diagnosis never writes to the evidence database.
|
||||
copy = out / "diagnosed-copy.db"
|
||||
shutil.copy2(args.db, copy)
|
||||
os.environ["AIDND_DB_PATH"] = str(copy)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from tools import memory_diagnostic as md # after the database is chosen
|
||||
|
||||
if args.command == "scenarios":
|
||||
names = args.only or list(md.SCENARIOS)
|
||||
summary = {}
|
||||
for name in names:
|
||||
result = md.run_scenario(md.SCENARIOS[name])
|
||||
(out / f"{name}.json").write_text(json.dumps(result, indent=2, default=str))
|
||||
d = result.get("diagnosis") or {}
|
||||
summary[name] = {
|
||||
"verdict": d.get("verdict"),
|
||||
"isolation_ok": (result.get("isolation") or {}).get("ok"),
|
||||
"plant_depth": result.get("plant_depth"),
|
||||
"recall_depth": result.get("recall_depth"),
|
||||
"f_evicted_at_turn": (result.get("eviction") or {}).get("f_evicted_at_turn"),
|
||||
}
|
||||
print(f"{name:26} verdict={d.get('verdict')!s:26} "
|
||||
f"isolation_ok={summary[name]['isolation_ok']} "
|
||||
f"plant={result.get('plant_depth')} recall={result.get('recall_depth')}")
|
||||
(out / "summary.json").write_text(json.dumps(summary, indent=2))
|
||||
return 0
|
||||
|
||||
import asyncio
|
||||
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import memorybank, models
|
||||
from app.database import SessionLocal
|
||||
|
||||
endpoint = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
embed_model = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
||||
with SessionLocal() as db:
|
||||
adventure = db.query(models.Adventure).order_by(models.Adventure.id).first()
|
||||
settings = db.query(models.Settings).filter_by(user_id=adventure.user_id).first()
|
||||
if endpoint:
|
||||
settings.endpoint_url = endpoint
|
||||
if embed_model:
|
||||
settings.embedding_model = embed_model
|
||||
recall_action = (db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adventure.id,
|
||||
models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first())
|
||||
embed = memorybank.embedding_provider(settings).embed
|
||||
iso = md.isolation(db, adventure, md.FACT_F, args.plant_depth,
|
||||
recall_snapshot=recall_action.context_snapshot,
|
||||
recall_depth=recall_action.depth)
|
||||
diagnosis = asyncio.run(md.diagnose(db, adventure, settings, md.FACT_F, args.plant_depth,
|
||||
recall_action=recall_action, embed=embed))
|
||||
variants = {}
|
||||
memory_id = diagnosis["created"]["memory_id"]
|
||||
if memory_id is not None and not diagnosis.get("retained", {}).get("forgotten"):
|
||||
base = md.production_query(adventure, recall_action.id)
|
||||
for label, query in (("recall_turn", base),
|
||||
("paraphrase", md.variant_query(base, md.PARAPHRASE_QUERY)),
|
||||
("unrelated", md.variant_query(base, md.UNRELATED_QUERY))):
|
||||
ranking = asyncio.run(md.rank_bank(db, adventure, settings, query, embed))
|
||||
row = next((r for r in ranking["scored"] if r["memory_id"] == memory_id), None)
|
||||
variants[label] = {"rank": row and row["rank"], "of": len(ranking["scored"]),
|
||||
"similarity": row and row["similarity"],
|
||||
"lexical_score": row and row["lexical_score"],
|
||||
"final_score": row and row["final_score"],
|
||||
"selected": bool(row and row["selected"])}
|
||||
db.rollback()
|
||||
report = {"isolation": iso, "diagnosis": diagnosis, "ranking_variants": variants}
|
||||
(out / "diagnosis.json").write_text(json.dumps(report, indent=2, default=str))
|
||||
print(json.dumps({"isolation_ok": iso["ok"], "verdict": diagnosis["verdict"]}, indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,187 @@
|
||||
"""v1.1: does a real v1.0.0 database open unchanged?
|
||||
|
||||
# the same database, opened by each tree, snapshotted read-only
|
||||
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
|
||||
--tree <v1.0.0 worktree>/backend --label v100 --out "$HOME/v11-evidence/compat"
|
||||
.venv/bin/python -m tools.v11_compat_check --db <v1 campaign.db> \\
|
||||
--label v11 --exercise --out "$HOME/v11-evidence/compat"
|
||||
|
||||
Run from `backend/`. The source database is never opened. It is copied into
|
||||
`--out` first, and the copy is what the application opens.
|
||||
|
||||
A read-only snapshot is taken through the API, the same way a reader sees the
|
||||
campaign:
|
||||
|
||||
- the export bundle, which carries the whole tree, the head, the Save Points,
|
||||
state, events, summaries, memories and knowledge, and has no timestamp of its
|
||||
own;
|
||||
- the narrative state and its events;
|
||||
- the Save Points, the imported knowledge, the memories, the derived status and
|
||||
the settings;
|
||||
- the database schema and `PRAGMA user_version`, before and after the
|
||||
application opened it.
|
||||
|
||||
Two snapshots of the same database from two trees are then compared. Identical
|
||||
means v1.1 read it exactly as v1.0.0 did, and a matching schema and version mean
|
||||
nothing migrated.
|
||||
|
||||
`--exercise` then uses the v1.1 copy: undo, redo, a Save Point restore, a
|
||||
context dry run (knowledge retrieval), an export, and an import of that export.
|
||||
It first points the copy's endpoint at a loopback port that refuses, so nothing
|
||||
here reaches an inference server.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def _schema(path: Path) -> dict:
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
version = connection.execute("PRAGMA user_version").fetchone()[0]
|
||||
rows = connection.execute(
|
||||
"SELECT type, name, sql FROM sqlite_master WHERE name NOT LIKE 'sqlite_%' "
|
||||
"ORDER BY type, name").fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
return {"user_version": version, "objects": [list(r) for r in rows]}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--db", required=True)
|
||||
parser.add_argument("--tree", default="", help="a backend/ directory to import the app from")
|
||||
parser.add_argument("--label", required=True)
|
||||
parser.add_argument("--exercise", action="store_true")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
copy = out / f"{args.label}.db"
|
||||
if copy.exists():
|
||||
print(f"{copy} exists; choose a new --label or --out")
|
||||
return 2
|
||||
shutil.copy2(args.db, copy)
|
||||
schema_before = _schema(copy)
|
||||
|
||||
if args.tree:
|
||||
sys.path.insert(0, str(Path(args.tree).resolve()))
|
||||
os.environ["AIDND_DB_PATH"] = str(copy)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, models
|
||||
from app.database import SessionLocal, get_db
|
||||
from app.main import app
|
||||
|
||||
print(f"app imported from {Path(sys.modules['app'].__file__).parent}")
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
with SessionLocal() as db:
|
||||
owner = db.query(models.Adventure.user_id).order_by(models.Adventure.id).first()
|
||||
user_id = owner[0] if owner else db.query(models.User.id).first()[0]
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
|
||||
def call(client, method, url, body=None, expect=200):
|
||||
response = client.request(method, f"/api{url}", json=body)
|
||||
if response.status_code != expect:
|
||||
raise SystemExit(f"{method} {url}: HTTP {response.status_code} {response.text[:300]}")
|
||||
return response.json() if response.content else None
|
||||
|
||||
report: dict = {"label": args.label, "schema_before": schema_before}
|
||||
with TestClient(app) as client:
|
||||
with SessionLocal() as db:
|
||||
adventure_ids = [a for (a,) in db.query(models.Adventure.id)
|
||||
.filter(models.Adventure.user_id == user_id)
|
||||
.order_by(models.Adventure.id)]
|
||||
snapshot = {"settings": call(client, "GET", "/settings"), "adventures": {}}
|
||||
for adv in adventure_ids:
|
||||
snapshot["adventures"][str(adv)] = {
|
||||
"export": call(client, "GET", f"/adventures/{adv}/export"),
|
||||
"state": call(client, "GET", f"/adventures/{adv}/state"),
|
||||
"state_events": call(client, "GET", f"/adventures/{adv}/state/events"),
|
||||
"checkpoints": call(client, "GET", f"/adventures/{adv}/checkpoints"),
|
||||
"knowledge": call(client, "GET", f"/adventures/{adv}/knowledge"),
|
||||
"memories": call(client, "GET", f"/adventures/{adv}/memories"),
|
||||
"derived": call(client, "GET", f"/adventures/{adv}/derived"),
|
||||
"newest_actions": call(client, "GET", f"/adventures/{adv}/actions?limit=5"),
|
||||
}
|
||||
report["snapshot"] = snapshot
|
||||
|
||||
if args.exercise and adventure_ids:
|
||||
adv = adventure_ids[0]
|
||||
ex: dict = {}
|
||||
call(client, "PUT", "/settings", {"endpoint_url": "http://127.0.0.1:9/v1",
|
||||
"embedding_model": ""})
|
||||
before = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
|
||||
ex["before"] = {k: before[k] for k in ("total", "can_undo", "can_redo")}
|
||||
undone = call(client, "POST", f"/adventures/{adv}/undo")
|
||||
ex["after_undo"] = {k: undone[k] for k in ("total", "can_undo", "can_redo")}
|
||||
redone = call(client, "POST", f"/adventures/{adv}/redo")
|
||||
ex["after_redo"] = {k: redone[k] for k in ("total", "can_undo", "can_redo")}
|
||||
points = call(client, "GET", f"/adventures/{adv}/checkpoints")
|
||||
if points:
|
||||
point = points[0]
|
||||
restored = call(client, "POST",
|
||||
f"/adventures/{adv}/checkpoints/{point['id']}/restore")
|
||||
page = call(client, "GET", f"/adventures/{adv}/actions?limit=1")
|
||||
ex["restore"] = {"save_point": point["name"], "total": page["total"],
|
||||
"can_redo": page["can_redo"],
|
||||
"response_keys": sorted(restored or {})}
|
||||
context = client.get(f"/api/adventures/{adv}/context")
|
||||
body = context.json()
|
||||
ex["context"] = {
|
||||
"status": context.status_code,
|
||||
"knowledge_used": len(((body.get("knowledge") or {}).get("used")) or []),
|
||||
"canon_section": any(s["label"] == "campaign_canon"
|
||||
for s in body.get("sections") or []),
|
||||
"state_section": any(s["label"] == "narrative_state"
|
||||
for s in body.get("sections") or []),
|
||||
"summary": body.get("summary"),
|
||||
"tokens": body.get("tokens"),
|
||||
"window": body.get("window"),
|
||||
}
|
||||
bundle = call(client, "GET", f"/adventures/{adv}/export")
|
||||
imported = call(client, "POST", "/adventures/import", bundle, expect=201)
|
||||
new_id = imported["id"]
|
||||
reimport = call(client, "GET", f"/adventures/{new_id}/export")
|
||||
ex["import"] = {
|
||||
"new_id": new_id,
|
||||
"actions_in_bundle": len(bundle.get("actions") or []),
|
||||
"actions_after_import": len(reimport.get("actions") or []),
|
||||
"head_same": (bundle.get("headBranch") is not None
|
||||
and bundle.get("headDepth") == reimport.get("headDepth")),
|
||||
"checkpoints": [len(bundle.get("checkpoints") or []),
|
||||
len(reimport.get("checkpoints") or [])],
|
||||
"memories": [len(bundle.get("memories") or []),
|
||||
len(reimport.get("memories") or [])],
|
||||
"narrative_state_same": bundle.get("narrativeState") == reimport.get("narrativeState"),
|
||||
}
|
||||
report["exercise"] = ex
|
||||
|
||||
app.dependency_overrides.clear()
|
||||
report["schema_after"] = _schema(copy)
|
||||
(out / f"{args.label}.json").write_text(json.dumps(report, indent=2, sort_keys=True, default=str))
|
||||
same_schema = report["schema_before"] == report["schema_after"]
|
||||
print(f"schema unchanged by opening: {same_schema} "
|
||||
f"(user_version {report['schema_before']['user_version']} -> "
|
||||
f"{report['schema_after']['user_version']})")
|
||||
if "exercise" in report:
|
||||
print(json.dumps(report["exercise"], indent=2, default=str)[:3000])
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,307 @@
|
||||
"""v1.1 release smoke test: the shipped image, as a reader would meet it.
|
||||
|
||||
python -m tools.v11_release_smoke --image <tag> --out <dir under $HOME>
|
||||
|
||||
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT` (an **HTTPS** Ollama on the
|
||||
trusted LAN) and `AIDND_TEST_MODEL`. `--ca` names the private CA to install
|
||||
inside the container, defaulting to this machine's own.
|
||||
|
||||
Supplemental release evidence, not a replacement for the gates: it asks whether
|
||||
the artefact that ships actually runs, reaches its approved narrator, refuses an
|
||||
unapproved one, and keeps a campaign across a container restart.
|
||||
|
||||
## The two things this is careful about
|
||||
|
||||
**The CA is installed, not bypassed.** `app/tlstrust.ssl_context()` is
|
||||
`ssl.create_default_context()` — the platform's own store — unioned with
|
||||
certifi's. So the private CA is mounted into
|
||||
`/usr/local/share/ca-certificates/` and registered with
|
||||
`update-ca-certificates`, and verification is then ordinary. Nothing sets
|
||||
`verify=False`, and a check inside the container proves the handshake succeeds
|
||||
through that store.
|
||||
|
||||
**Loopback means the published port.** The process inside the container listens
|
||||
on `0.0.0.0` because that is the only address a published port can reach
|
||||
(`docker-compose.yml` says so). What must be loopback-only is the *publish*, so
|
||||
the container is started with `-p 127.0.0.1:<port>:8000` and the check is that
|
||||
the host's LAN address refuses the same port.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import socket
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from tools.m11_webdriver import Browser, free_port, require_under_home # noqa: E402
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
NAME = "v11-release-smoke"
|
||||
VOLUME = "v11-release-smoke-data"
|
||||
#: An endpoint the policy must refuse whatever else is true: a public host.
|
||||
PUBLIC_ENDPOINT = "https://api.openai.com/v1"
|
||||
|
||||
|
||||
class Checks:
|
||||
def __init__(self) -> None:
|
||||
self.rows: list[dict] = []
|
||||
|
||||
def record(self, name: str, ok: bool, detail: str = "") -> bool:
|
||||
self.rows.append({"check": name, "result": "PASS" if ok else "FAIL",
|
||||
"detail": detail})
|
||||
print(f" {'ok ' if ok else 'FAIL'} {name}" + (f" — {detail}" if detail else ""),
|
||||
flush=True)
|
||||
return ok
|
||||
|
||||
@property
|
||||
def failed(self) -> list[dict]:
|
||||
return [r for r in self.rows if r["result"] == "FAIL"]
|
||||
|
||||
|
||||
def run(*args: str, **kwargs) -> subprocess.CompletedProcess:
|
||||
return subprocess.run(args, capture_output=True, text=True, **kwargs)
|
||||
|
||||
|
||||
def api(base: str, method: str, path: str, payload=None, timeout=900):
|
||||
data = json.dumps(payload).encode() if payload is not None else None
|
||||
request = urllib.request.Request(
|
||||
f"{base}/api{path}", data=data, method=method,
|
||||
headers={"Content-Type": "application/json"} if data else {})
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
body = response.read().decode()
|
||||
return json.loads(body) if body else None
|
||||
|
||||
|
||||
def stream_turn(base: str, adv: int, text: str) -> list[dict]:
|
||||
request = urllib.request.Request(
|
||||
f"{base}/api/adventures/{adv}/actions",
|
||||
data=json.dumps({"type": "do", "text": text}).encode(),
|
||||
method="POST", headers={"Content-Type": "application/json"})
|
||||
events: list[dict] = []
|
||||
with urllib.request.urlopen(request, timeout=900) as response:
|
||||
for raw in response:
|
||||
line = raw.decode(errors="replace").strip()
|
||||
if line.startswith("data:"):
|
||||
try:
|
||||
events.append(json.loads(line[5:].strip()))
|
||||
except json.JSONDecodeError:
|
||||
pass
|
||||
return events
|
||||
|
||||
|
||||
def lan_address() -> str | None:
|
||||
"""This machine's own LAN address, for the loopback-only check."""
|
||||
probe = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
|
||||
try:
|
||||
probe.connect(("192.0.2.1", 9)) # TEST-NET-1: routed nowhere, sends nothing
|
||||
return probe.getsockname()[0]
|
||||
except OSError:
|
||||
return None
|
||||
finally:
|
||||
probe.close()
|
||||
|
||||
|
||||
def wait_ready(base: str, *, timeout: float = 180) -> bool:
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
try:
|
||||
urllib.request.urlopen(f"{base}/api/settings", timeout=3)
|
||||
return True
|
||||
except Exception:
|
||||
time.sleep(1)
|
||||
return False
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--image", required=True)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument("--ca", default="/usr/local/share/ca-certificates/draco.crt")
|
||||
parser.add_argument(
|
||||
"--add-host", default="", metavar="NAME:ADDRESS",
|
||||
help=("resolve the narrator's hostname inside the container. A `.local` "
|
||||
"name is mDNS, and a container has no mDNS resolver, so the "
|
||||
"endpoint policy refuses an address it cannot classify and "
|
||||
"`PUT /api/settings` answers 400. Mapping the name — rather than "
|
||||
"using the address — keeps the hostname the certificate is issued "
|
||||
"for, which is the thing this test verifies."))
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL):
|
||||
print("set AIDND_TEST_ENDPOINT (https://…) and AIDND_TEST_MODEL")
|
||||
return 2
|
||||
if not ENDPOINT.startswith("https://"):
|
||||
print("the smoke test needs an HTTPS endpoint: that is what it verifies")
|
||||
return 2
|
||||
ca = Path(args.ca)
|
||||
if not ca.exists():
|
||||
print(f"no CA at {ca}")
|
||||
return 2
|
||||
|
||||
out = require_under_home(Path(args.out).expanduser())
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
checks = Checks()
|
||||
port = free_port()
|
||||
base = f"http://127.0.0.1:{port}"
|
||||
started = datetime.now()
|
||||
|
||||
run("docker", "rm", "-f", NAME)
|
||||
run("docker", "volume", "rm", VOLUME)
|
||||
run("docker", "volume", "create", VOLUME)
|
||||
|
||||
print(f"starting {args.image} on 127.0.0.1:{port} with a fresh volume …")
|
||||
start = run(
|
||||
"docker", "run", "-d", "--name", NAME,
|
||||
"-p", f"127.0.0.1:{port}:8000",
|
||||
"-v", f"{VOLUME}:/data",
|
||||
"-v", f"{ca}:/usr/local/share/ca-certificates/{ca.name}:ro",
|
||||
*(("--add-host", args.add_host) if args.add_host else ()),
|
||||
args.image,
|
||||
"sh", "-c",
|
||||
"update-ca-certificates >/dev/null 2>&1; "
|
||||
"exec uvicorn app.main:app --host 0.0.0.0 --port 8000",
|
||||
)
|
||||
if start.returncode != 0:
|
||||
print(start.stderr[:400])
|
||||
return 1
|
||||
container = start.stdout.strip()[:12]
|
||||
|
||||
try:
|
||||
checks.record("the container starts", True, container)
|
||||
ready = wait_ready(base)
|
||||
if not checks.record("the application answers on loopback", ready, base):
|
||||
logs = run("docker", "logs", NAME)
|
||||
(out / "container.log").write_text(logs.stdout + logs.stderr)
|
||||
return 1
|
||||
|
||||
published = run("docker", "port", NAME).stdout.strip()
|
||||
checks.record("the port is published on loopback only",
|
||||
"127.0.0.1" in published and "0.0.0.0" not in published, published)
|
||||
|
||||
lan = lan_address()
|
||||
if lan:
|
||||
try:
|
||||
urllib.request.urlopen(f"http://{lan}:{port}/api/settings", timeout=4)
|
||||
reachable = True
|
||||
except Exception:
|
||||
reachable = False
|
||||
checks.record("the LAN address does not serve the application", not reachable,
|
||||
f"port {port} on this machine's LAN address")
|
||||
|
||||
page = urllib.request.urlopen(base + "/", timeout=30)
|
||||
html = page.read().decode(errors="replace")
|
||||
checks.record("the first page loads", page.status == 200 and "<div id=\"root\"" in html,
|
||||
f"HTTP {page.status}, {len(html)} bytes")
|
||||
remote = [chunk for chunk in html.split('"')
|
||||
if chunk.startswith("http://") or chunk.startswith("https://")]
|
||||
checks.record("the shell references no remote origin", not remote, str(remote[:3]))
|
||||
csp = page.headers.get("content-security-policy") or ""
|
||||
checks.record("a CSP is served", bool(csp), csp[:80])
|
||||
|
||||
# The approved endpoint, verified through the private CA *inside* the
|
||||
# container, with the application's own trust context and no bypass.
|
||||
probe = run("docker", "exec", NAME, "python", "-c",
|
||||
"import json,urllib.request,ssl,sys;"
|
||||
"sys.path.insert(0,'/app/backend');"
|
||||
"from app.tlstrust import ssl_context;"
|
||||
f"r=urllib.request.urlopen('{ENDPOINT}/models',"
|
||||
" timeout=20, context=ssl_context());"
|
||||
"print(r.status)")
|
||||
checks.record("the approved HTTPS narrator verifies through the private CA",
|
||||
probe.returncode == 0 and "200" in probe.stdout,
|
||||
(probe.stdout + probe.stderr).strip()[:160])
|
||||
|
||||
api(base, "PUT", "/settings", {"endpoint_url": ENDPOINT, "model": MODEL,
|
||||
"max_output_tokens": 300,
|
||||
"model_timeout_seconds": 600})
|
||||
try:
|
||||
api(base, "PUT", "/settings", {"endpoint_url": PUBLIC_ENDPOINT, "model": MODEL})
|
||||
refused = False
|
||||
detail = "accepted"
|
||||
except urllib.error.HTTPError as exc:
|
||||
refused = 400 <= exc.code < 500
|
||||
detail = f"HTTP {exc.code}"
|
||||
checks.record("a public endpoint is refused", refused, detail)
|
||||
# Put the approved one back, whatever happened above.
|
||||
api(base, "PUT", "/settings", {"endpoint_url": ENDPOINT, "model": MODEL,
|
||||
"max_output_tokens": 300,
|
||||
"model_timeout_seconds": 600})
|
||||
|
||||
created = api(base, "POST", "/adventures", {
|
||||
"title": "Release Smoke",
|
||||
"opening": "Rain over Westhaven, and the abbey bell tolling.",
|
||||
"persona_name": "Aldric"})
|
||||
adv = created["id"]
|
||||
checks.record("a campaign is created", bool(adv), f"id {adv}")
|
||||
|
||||
events = stream_turn(base, adv, "I ask Mara what the bell means.")
|
||||
errors = [e for e in events if e.get("type") == "error"]
|
||||
page_after = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {}
|
||||
checks.record("one real narrator turn is accepted",
|
||||
not errors and (page_after.get("total") or 0) >= 2,
|
||||
errors[0].get("detail", "")[:160] if errors else
|
||||
f"{page_after.get('total')} actions")
|
||||
before = [(a.get("type"), (a.get("text") or "")[:120])
|
||||
for a in (page_after.get("actions") or [])]
|
||||
state_before = api(base, "GET", f"/adventures/{adv}/state") or {}
|
||||
|
||||
print("restarting the container …")
|
||||
run("docker", "restart", NAME)
|
||||
ready = wait_ready(base)
|
||||
checks.record("the container restarts and serves again", ready)
|
||||
|
||||
page_reopened = api(base, "GET", f"/adventures/{adv}/actions?limit=50") or {}
|
||||
after = [(a.get("type"), (a.get("text") or "")[:120])
|
||||
for a in (page_reopened.get("actions") or [])]
|
||||
checks.record("the transcript survived the restart", after == before,
|
||||
f"{len(before)} -> {len(after)} actions")
|
||||
state_after = api(base, "GET", f"/adventures/{adv}/state") or {}
|
||||
checks.record("the narrative state survived the restart",
|
||||
state_after == state_before)
|
||||
|
||||
browser = Browser(headless=True, log=out / "geckodriver.log")
|
||||
try:
|
||||
browser.go(f"{base}/play/{adv}")
|
||||
browser.wait_for(".story-controls", timeout=60)
|
||||
story = browser.js(
|
||||
"const el = document.querySelector('.story');"
|
||||
" return el ? el.textContent.trim().length : 0;")
|
||||
checks.record("Firefox renders the reopened campaign",
|
||||
isinstance(story, int) and story > 0, f"{story} characters of story")
|
||||
browser.screenshot(out / "reopened-campaign.png")
|
||||
finally:
|
||||
browser.quit()
|
||||
finally:
|
||||
logs = run("docker", "logs", NAME)
|
||||
(out / "container.log").write_text(logs.stdout + logs.stderr)
|
||||
run("docker", "rm", "-f", NAME)
|
||||
run("docker", "volume", "rm", VOLUME)
|
||||
|
||||
report = {
|
||||
"image": args.image,
|
||||
"started": started.isoformat(timespec="seconds"),
|
||||
"seconds": round((datetime.now() - started).total_seconds()),
|
||||
"endpoint_class": "trusted-LAN HTTPS with a private CA",
|
||||
"checks": checks.rows,
|
||||
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
|
||||
"failed": len(checks.failed),
|
||||
}
|
||||
(out / "smoke-report.json").write_text(json.dumps(report, indent=2))
|
||||
print(f"\n{report['passed']} passed, {report['failed']} failed "
|
||||
f"-> {out / 'smoke-report.json'}")
|
||||
return 1 if checks.failed else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,245 @@
|
||||
"""v1.1 WP-A2: replay real stored narration through the v1.0.0 and current extractors.
|
||||
|
||||
The A2 extractor changes remove more text from a narrator's reply than v1.0.0
|
||||
did. Removing story is worse than leaving protocol (`TECHNICAL-DESIGN.md`
|
||||
§15.4), so every change is shown to a person rather than summarised. This tool
|
||||
takes every real reply the evidence kept, runs it through both extractors, and
|
||||
writes each turn whose prose differs with:
|
||||
|
||||
- the v1.0.0 prose and the current prose;
|
||||
- every line removed, and the rule that explains it;
|
||||
- any removal no rule explains, which fails the replay.
|
||||
|
||||
**Input.** A reply is read from the turn's stored `raw_output` where the
|
||||
evidence database kept one: that is exactly what the narrator sent. A bundle
|
||||
carries no raw reply, so a bundle's turns are replayed from their stored text,
|
||||
which is v1.0.0's output already. For those the old prose is the input itself,
|
||||
and the comparison is still exact.
|
||||
|
||||
**The v1.0.0 extractor** is read from the release tag with `git show`, not
|
||||
copied, so this tool compares against what shipped. It shares `events` and
|
||||
`render` with the current tree. A2 does not change `events.SPECS` or
|
||||
`render.SECTION_HEADINGS`, and the report verifies that with `git diff`.
|
||||
|
||||
Evidence stays outside the repository. Replayed text is fiction from the
|
||||
acceptance fixtures, but it is still somebody's run.
|
||||
|
||||
.venv/bin/python -m tools.v11_replay_extractor \\
|
||||
--db "$HOME/m11-evidence/**/*.db" \\
|
||||
--bundle "$HOME/m11-evidence/closeout-3652dc6/identity/turn-99/bundle.json" \\
|
||||
--out "$HOME/v11-evidence/a2-replay"
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import difflib
|
||||
import glob
|
||||
import hashlib
|
||||
import importlib.util
|
||||
import json
|
||||
import sqlite3
|
||||
import subprocess
|
||||
import sys
|
||||
import types
|
||||
import zlib
|
||||
from pathlib import Path
|
||||
|
||||
RELEASE = "v1.0.0"
|
||||
EXTRACT_PATH = "backend/app/narrative/extract.py"
|
||||
#: A removal larger than this share of the v1.0.0 prose is flagged for review
|
||||
#: even when every line is explained, because a rule that eats most of a reply
|
||||
#: is the shape a false positive takes.
|
||||
LARGE_REMOVAL_SHARE = 0.25
|
||||
|
||||
|
||||
def load_release_extractor(repo: Path) -> types.ModuleType:
|
||||
"""`app.narrative.extract` as it was at the release tag."""
|
||||
source = subprocess.run(
|
||||
["git", "-C", str(repo), "show", f"{RELEASE}:{EXTRACT_PATH}"],
|
||||
check=True, capture_output=True, text=True,
|
||||
).stdout
|
||||
import app.narrative # noqa: F401 the package the relative import needs
|
||||
|
||||
spec = importlib.util.spec_from_loader("app.narrative._extract_release", loader=None)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
module.__package__ = "app.narrative"
|
||||
exec(compile(source, f"{RELEASE}:{EXTRACT_PATH}", "exec"), module.__dict__)
|
||||
return module
|
||||
|
||||
|
||||
def _unpack(blob):
|
||||
if blob is None:
|
||||
return None
|
||||
try:
|
||||
return json.loads(zlib.decompress(bytes(blob)).decode("utf-8"))
|
||||
except (zlib.error, ValueError, UnicodeDecodeError):
|
||||
return None
|
||||
|
||||
|
||||
def turns_from_db(path: str):
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute(
|
||||
"SELECT id, depth, text, context_snapshot FROM actions WHERE type = 'ai'"
|
||||
).fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
for action_id, depth, text, blob in rows:
|
||||
snapshot = _unpack(blob) or {}
|
||||
raw = snapshot.get("raw_output")
|
||||
if isinstance(raw, str) and raw.strip():
|
||||
yield {"source": path, "id": action_id, "depth": depth,
|
||||
"input": raw, "input_kind": "raw_output"}
|
||||
elif text:
|
||||
yield {"source": path, "id": action_id, "depth": depth,
|
||||
"input": text, "input_kind": "stored_text"}
|
||||
|
||||
|
||||
def turns_from_bundle(path: str):
|
||||
bundle = json.loads(Path(path).read_text())
|
||||
for action in bundle.get("actions") or []:
|
||||
if action.get("type") == "ai" and action.get("text"):
|
||||
yield {"source": path, "id": action.get("id"), "depth": action.get("depth"),
|
||||
"input": action["text"], "input_kind": "stored_text"}
|
||||
|
||||
|
||||
def removed_lines(before: str, after: str) -> list[str]:
|
||||
"""Lines present in `before` and gone from `after`, in order.
|
||||
|
||||
Compared with trailing whitespace ignored. The extractor strips the end of
|
||||
every reply it cuts, so a story line that becomes the last line loses a
|
||||
trailing space. A first version of this tool reported that space as a
|
||||
rewritten line of story, which it is not.
|
||||
"""
|
||||
old = [line.rstrip() for line in before.split("\n")]
|
||||
new = [line.rstrip() for line in after.split("\n")]
|
||||
matcher = difflib.SequenceMatcher(a=old, b=new, autojunk=False)
|
||||
gone: list[str] = []
|
||||
for tag, a0, a1, _b0, _b1 in matcher.get_opcodes():
|
||||
if tag in ("delete", "replace"):
|
||||
gone.extend(old[a0:a1])
|
||||
return gone
|
||||
|
||||
|
||||
def replay(inputs, old, new) -> dict:
|
||||
seen: set[str] = set()
|
||||
unchanged = 0
|
||||
changed: list[dict] = []
|
||||
duplicates = 0
|
||||
for turn in inputs:
|
||||
digest = hashlib.sha256(turn["input"].encode()).hexdigest()
|
||||
if digest in seen:
|
||||
duplicates += 1
|
||||
continue
|
||||
seen.add(digest)
|
||||
old_prose, _old_parsed, _old_raw = old.split(turn["input"])
|
||||
new_prose, _new_parsed, _new_raw = new.split(turn["input"])
|
||||
if old_prose == new_prose:
|
||||
unchanged += 1
|
||||
continue
|
||||
lines = []
|
||||
unexplained = 0
|
||||
for line in removed_lines(old_prose, new_prose):
|
||||
if not line.strip():
|
||||
continue
|
||||
rule = new.explain_removed_line(line)
|
||||
if rule is None:
|
||||
unexplained += 1
|
||||
lines.append({"line": line, "rule": rule})
|
||||
added = [line for line in removed_lines(new_prose, old_prose) if line.strip()]
|
||||
share = 1 - len(new_prose) / max(1, len(old_prose))
|
||||
flags = []
|
||||
if unexplained:
|
||||
flags.append("unexplained_removal")
|
||||
if added:
|
||||
flags.append("text_added_or_rewritten")
|
||||
if share > LARGE_REMOVAL_SHARE:
|
||||
flags.append("large_removal")
|
||||
changed.append({
|
||||
**{k: turn[k] for k in ("source", "id", "depth", "input_kind")},
|
||||
"sha256": digest,
|
||||
"old_prose": old_prose,
|
||||
"new_prose": new_prose,
|
||||
"removed": lines,
|
||||
"added_or_rewritten": added,
|
||||
"removed_chars": len(old_prose) - len(new_prose),
|
||||
"removed_share": round(share, 4),
|
||||
"flags": flags,
|
||||
})
|
||||
return {
|
||||
"replayed": unchanged + len(changed),
|
||||
"duplicates_skipped": duplicates,
|
||||
"unchanged": unchanged,
|
||||
"changed": len(changed),
|
||||
"flagged": sum(1 for c in changed if c["flags"]),
|
||||
"turns": changed,
|
||||
}
|
||||
|
||||
|
||||
def write_markdown(result: dict, path: Path) -> None:
|
||||
out = [
|
||||
"# A2 extractor replay",
|
||||
"",
|
||||
f"- replayed (unique replies): **{result['replayed']}**",
|
||||
f"- duplicates skipped: {result['duplicates_skipped']}",
|
||||
f"- unchanged: {result['unchanged']}",
|
||||
f"- changed: **{result['changed']}**",
|
||||
f"- flagged: **{result['flagged']}**",
|
||||
"",
|
||||
]
|
||||
for index, turn in enumerate(result["turns"], 1):
|
||||
out += [
|
||||
f"## {index}. {Path(turn['source']).parent.name}/{Path(turn['source']).name}"
|
||||
f" action {turn['id']} depth {turn['depth']} ({turn['input_kind']})",
|
||||
"",
|
||||
f"- removed chars: {turn['removed_chars']} ({turn['removed_share']:.1%})",
|
||||
f"- flags: {', '.join(turn['flags']) or 'none'}",
|
||||
"",
|
||||
"Removed lines:",
|
||||
"",
|
||||
]
|
||||
for item in turn["removed"]:
|
||||
out.append(f"- `{item['rule'] or 'UNEXPLAINED'}` — {item['line']!r}")
|
||||
out += ["", "<details><summary>v1.0.0 prose</summary>", "", "```text",
|
||||
turn["old_prose"], "```", "</details>", "",
|
||||
"<details><summary>current prose</summary>", "", "```text",
|
||||
turn["new_prose"], "```", "</details>", ""]
|
||||
path.write_text("\n".join(out))
|
||||
|
||||
|
||||
def main(argv=None) -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--db", action="append", default=[],
|
||||
help="an evidence database, or a glob of them")
|
||||
parser.add_argument("--bundle", action="append", default=[],
|
||||
help="an exported bundle whose campaign has no database here")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args(argv)
|
||||
|
||||
repo = Path(__file__).resolve().parents[2]
|
||||
from app.narrative import extract as current
|
||||
|
||||
old = load_release_extractor(repo)
|
||||
dbs = sorted({p for pattern in args.db for p in glob.glob(pattern, recursive=True)})
|
||||
|
||||
def inputs():
|
||||
for path in dbs:
|
||||
yield from turns_from_db(path)
|
||||
for path in args.bundle:
|
||||
yield from turns_from_bundle(path)
|
||||
|
||||
result = replay(inputs(), old, current)
|
||||
result["databases"] = dbs
|
||||
result["bundles"] = args.bundle
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
(out / "replay.json").write_text(json.dumps(result, indent=2, ensure_ascii=False))
|
||||
write_markdown(result, out / "replay.md")
|
||||
print(json.dumps({k: result[k] for k in
|
||||
("replayed", "duplicates_skipped", "unchanged", "changed", "flagged")}))
|
||||
return 1 if result["flagged"] else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,381 @@
|
||||
"""v1.1 release Gate 9: a real v1.0.0 campaign, opened by the candidate.
|
||||
|
||||
python -m tools.v11_upgrade_check --v100 <worktree> --out <dir under $HOME>
|
||||
|
||||
Run from `backend/`. Reads `AIDND_TEST_ENDPOINT`, `AIDND_TEST_MODEL` and
|
||||
`AIDND_TEST_EMBED_MODEL`: the campaign has to be *played*, because memories,
|
||||
summaries and narrative state are things a narrator produces. A schema-only
|
||||
fixture would prove nothing about an upgrade, which is why §11 item 9 asks for a
|
||||
database the v1.0.0 application built.
|
||||
|
||||
Three phases, each its own server process, so everything that survives crosses
|
||||
as bytes on disk:
|
||||
|
||||
1. **v1.0.0 builds and plays.** The `432f041` tree serves the application: a
|
||||
campaign is created, canon knowledge imported, turns played, a Save Point
|
||||
taken, a turn undone so the head is not at the tip and Redo is available, and
|
||||
a narration length chosen. Then that server stops, and a census is taken.
|
||||
2. **The candidate opens the same file.** Nothing is copied; the candidate's own
|
||||
migrations run against it. The census is taken again and compared field by
|
||||
field.
|
||||
3. **Bundles cross both ways.** The v1.0.0 export is imported by the candidate.
|
||||
The candidate's export is offered back to v1.0.0, and whatever happens is
|
||||
reported — the format string is unchanged, which is a reason to test backward
|
||||
import, not a reason to assume it.
|
||||
|
||||
Settings hold only what the caller's environment names, and the evidence
|
||||
directory lives under `$HOME`.
|
||||
|
||||
## Shapes this had to be written against, not guessed
|
||||
|
||||
- A turn is **SSE**: `POST /adventures/{id}/actions`, and a failed turn is an
|
||||
`error` *event* inside an HTTP 200. Reading the status code would call every
|
||||
failure a success.
|
||||
- History is `GET /{id}/actions?limit=N` -> `{actions, total, has_more,
|
||||
can_undo, can_redo}`. There is **no head-id field**, so the active head is
|
||||
compared as the newest action plus the two flags.
|
||||
- Save Points are **checkpoints**. Creating one after an Undo names the undone
|
||||
position, deliberately.
|
||||
- There is **no summaries route**; summaries and post-turn health both come from
|
||||
`GET /{id}/derived`.
|
||||
- Campaign switches are `PATCH /adventures/{id}` with `memory_bank_enabled` and
|
||||
`auto_summarize` — not the names a reader would guess.
|
||||
- Knowledge import is **multipart**, as `m11_long_run` and `m11_offline` build it.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from tools.m11_webdriver import free_port, require_under_home # noqa: E402
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
EMBED_MODEL = os.environ.get("AIDND_TEST_EMBED_MODEL", "")
|
||||
TURN_TIMEOUT = 900
|
||||
#: What the evidence database is left holding. The campaign has to be played
|
||||
#: against a real narrator, but nothing about the *upgrade* depends on which
|
||||
#: host that was, and §11 item 9 asks for no real hostname in this database.
|
||||
PLACEHOLDER_ENDPOINT = "http://127.0.0.1:11434/v1"
|
||||
|
||||
#: What the upgrade must preserve, compared exactly on both sides. `settings` is
|
||||
#: included because a migration that silently rewrote an endpoint would be a real
|
||||
#: defect; `schema_version` is read from the file rather than the API.
|
||||
CENSUS = ("transcript", "newest_action", "total", "can_undo", "can_redo",
|
||||
"checkpoints", "state", "memories", "summaries", "knowledge",
|
||||
"narration_length", "memory_bank_enabled", "auto_summarize",
|
||||
"settings", "schema_version")
|
||||
|
||||
CANON_MD = """# Westhaven
|
||||
|
||||
The abbey bell is rung only for a death. Mara keeps the harbour ledger.
|
||||
Aldric carries a silver key he will not explain.
|
||||
"""
|
||||
|
||||
|
||||
class App:
|
||||
"""One application process, from whichever tree it is given."""
|
||||
|
||||
def __init__(self, tree: Path, db: Path, log: Path, label: str):
|
||||
self.tree, self.db, self.label = tree, db, label
|
||||
self.port = free_port()
|
||||
self.log = log
|
||||
handle = open(log, "ab")
|
||||
self.proc = subprocess.Popen(
|
||||
[str(Path(__file__).resolve().parent.parent / ".venv/bin/uvicorn"),
|
||||
"app.main:app", "--host", "127.0.0.1", "--port", str(self.port)],
|
||||
cwd=str(tree / "backend"), stdout=handle, stderr=subprocess.STDOUT,
|
||||
env={**os.environ, "AIDND_DB_PATH": str(db),
|
||||
"AIDND_DATABASE_URL": "", "DATABASE_URL": ""},
|
||||
)
|
||||
self.url = f"http://127.0.0.1:{self.port}"
|
||||
deadline = time.monotonic() + 120
|
||||
while time.monotonic() < deadline:
|
||||
if self.proc.poll() is not None:
|
||||
raise RuntimeError(f"{label} exited early; see {log}")
|
||||
try:
|
||||
urllib.request.urlopen(self.url + "/api/settings", timeout=2)
|
||||
print(f" {label} serving {db.name} on {self.url}", flush=True)
|
||||
return
|
||||
except Exception:
|
||||
time.sleep(0.2)
|
||||
raise RuntimeError(f"{label} never became ready; see {log}")
|
||||
|
||||
def call(self, method: str, path: str, payload=None, timeout=120):
|
||||
data = json.dumps(payload).encode() if payload is not None else None
|
||||
request = urllib.request.Request(
|
||||
f"{self.url}/api{path}", data=data, method=method,
|
||||
headers={"Content-Type": "application/json"} if data else {})
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
body = response.read().decode()
|
||||
return json.loads(body) if body else None
|
||||
|
||||
def stream(self, path: str, payload) -> list[dict]:
|
||||
"""A turn. A failed turn is an event in the stream, not a status code."""
|
||||
request = urllib.request.Request(
|
||||
f"{self.url}/api{path}", data=json.dumps(payload).encode(),
|
||||
method="POST", headers={"Content-Type": "application/json"})
|
||||
events: list[dict] = []
|
||||
with urllib.request.urlopen(request, timeout=TURN_TIMEOUT) as response:
|
||||
for raw in response:
|
||||
line = raw.decode(errors="replace").strip()
|
||||
if line.startswith("data:"):
|
||||
try:
|
||||
events.append(json.loads(line[5:].strip()))
|
||||
except json.JSONDecodeError:
|
||||
pass
|
||||
return events
|
||||
|
||||
def upload(self, adv: int, name: str, body: str, classification: str) -> dict:
|
||||
boundary = "----v11upgrade"
|
||||
parts = (
|
||||
f"--{boundary}\r\nContent-Disposition: form-data; name=\"classification\""
|
||||
f"\r\n\r\n{classification}\r\n"
|
||||
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
|
||||
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
|
||||
f"--{boundary}--\r\n"
|
||||
).encode()
|
||||
request = urllib.request.Request(
|
||||
f"{self.url}/api/adventures/{adv}/knowledge", data=parts, method="POST",
|
||||
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"})
|
||||
with urllib.request.urlopen(request, timeout=120) as response:
|
||||
return json.loads(response.read().decode())
|
||||
|
||||
def stop(self) -> None:
|
||||
if self.proc.poll() is None:
|
||||
self.proc.terminate()
|
||||
try:
|
||||
self.proc.wait(timeout=30)
|
||||
except subprocess.TimeoutExpired:
|
||||
self.proc.kill()
|
||||
|
||||
|
||||
def schema_version(db: Path) -> int:
|
||||
connection = sqlite3.connect(f"file:{db}?mode=ro", uri=True)
|
||||
try:
|
||||
return connection.execute("PRAGMA user_version").fetchone()[0]
|
||||
finally:
|
||||
connection.close()
|
||||
|
||||
|
||||
def census(app: App, adv: int, db: Path) -> dict:
|
||||
page = app.call("GET", f"/adventures/{adv}/actions?limit=500") or {}
|
||||
actions = page.get("actions") or []
|
||||
derived = app.call("GET", f"/adventures/{adv}/derived") or {}
|
||||
adventure = app.call("GET", f"/adventures/{adv}") or {}
|
||||
settings = app.call("GET", "/settings") or {}
|
||||
checkpoints = app.call("GET", f"/adventures/{adv}/checkpoints") or []
|
||||
memories = app.call("GET", f"/adventures/{adv}/memories") or []
|
||||
state = app.call("GET", f"/adventures/{adv}/state") or {}
|
||||
knowledge = app.call("GET", f"/adventures/{adv}/knowledge") or []
|
||||
newest = actions[-1] if actions else {}
|
||||
return {
|
||||
"transcript": [(a.get("type"), (a.get("text") or "")[:300]) for a in actions],
|
||||
# No head id is exposed; the head is the newest action on the read line
|
||||
# plus the two flags the page carries.
|
||||
"newest_action": ((newest.get("type"), (newest.get("text") or "")[:300])
|
||||
if newest else None),
|
||||
"total": page.get("total"),
|
||||
"can_undo": page.get("can_undo"),
|
||||
"can_redo": page.get("can_redo"),
|
||||
"checkpoints": sorted((c.get("name"), c.get("depth"), c.get("branch_id"),
|
||||
c.get("on_path"), c.get("resolved"))
|
||||
for c in checkpoints),
|
||||
"state": state.get("document") if isinstance(state, dict) else state,
|
||||
"memories": sorted((m.get("text") or "")[:200] for m in memories),
|
||||
"summaries": sorted((s.get("text") or "")[:200]
|
||||
for s in (derived.get("summaries") or [])),
|
||||
"knowledge": sorted((k.get("filename") or k.get("title"),
|
||||
k.get("classification")) for k in knowledge),
|
||||
"narration_length": adventure.get("narration_length"),
|
||||
"memory_bank_enabled": adventure.get("memory_bank_enabled"),
|
||||
"auto_summarize": adventure.get("auto_summarize"),
|
||||
"settings": {k: settings.get(k)
|
||||
for k in ("endpoint_url", "model", "max_output_tokens")},
|
||||
"schema_version": schema_version(db),
|
||||
}
|
||||
|
||||
|
||||
def play(app: App, adv: int, text: str) -> bool:
|
||||
events = app.stream(f"/adventures/{adv}/actions", {"type": "do", "text": text})
|
||||
errors = [e for e in events if e.get("type") == "error"]
|
||||
if errors:
|
||||
print(f" turn refused: {errors[0].get('detail', '')[:150]}", flush=True)
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
def build_v100_campaign(app: App) -> int:
|
||||
created = app.call("POST", "/adventures", {
|
||||
"title": "Upgrade Evidence",
|
||||
"opening": "Rain over Westhaven, and the abbey bell tolling.",
|
||||
"canon_rules": ["The dead do not return."],
|
||||
"persona_name": "Aldric",
|
||||
})
|
||||
adv = created["id"]
|
||||
app.call("PUT", "/settings", {
|
||||
"endpoint_url": ENDPOINT, "model": MODEL,
|
||||
"embedding_model": EMBED_MODEL,
|
||||
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
|
||||
# Memory and summaries are per-campaign switches defaulting to off, so the
|
||||
# census would otherwise have nothing to compare.
|
||||
app.call("PATCH", f"/adventures/{adv}", {
|
||||
"narration_length": "brief", "memory_bank_enabled": True,
|
||||
"auto_summarize": True})
|
||||
app.upload(adv, "westhaven-canon.md", CANON_MD, "canon")
|
||||
|
||||
# Enough turns that memories and summaries actually exist. A memory needs
|
||||
# MEMORY_INTERVAL (6) actions plus SETTLE_SLACK (1) settled past the anchor,
|
||||
# and a summary needs SUMMARY_INTERVAL (15) uncovered actions — both counted
|
||||
# in *actions*, and a turn writes two. A first pass at this gate played five
|
||||
# turns, wrote neither, and compared 0 against 0, which proves nothing about
|
||||
# whether the upgrade preserves them.
|
||||
beats = [
|
||||
"I ask Mara what the bell means.",
|
||||
"I show her the silver key.",
|
||||
"I follow her to the harbour ledger.",
|
||||
"I ask who else knows about the key.",
|
||||
"I read the ledger's last page aloud.",
|
||||
"I ask the ferryman about the fen road.",
|
||||
"I wait out the rain and watch the harbour.",
|
||||
"I ask Mara about the abbey's sealed crypt.",
|
||||
"I count the entries against the tide table.",
|
||||
"I ask who signed for the last shipment.",
|
||||
"I walk the quay to the chandler's door.",
|
||||
"I ask the chandler what he remembers of that night.",
|
||||
"I show the chandler the key.",
|
||||
"I return to Mara with what he said.",
|
||||
"I ask Mara what she means to do now.",
|
||||
"I agree to meet her at first light.",
|
||||
"I take the long way back along the ridge.",
|
||||
"I check whether anyone followed me.",
|
||||
"I write down what I have learned so far.",
|
||||
"I sleep, and wake before the bell.",
|
||||
]
|
||||
played = 0
|
||||
for text in beats:
|
||||
if play(app, adv, text):
|
||||
played += 1
|
||||
print(f" {played} turns accepted by v1.0.0", flush=True)
|
||||
|
||||
app.call("POST", f"/adventures/{adv}/checkpoints", {"name": "before the ledger"})
|
||||
# One Undo, so the head is not at the retained tip and Redo is available.
|
||||
app.call("POST", f"/adventures/{adv}/undo", {})
|
||||
|
||||
# §11 item 9: the database must carry **loopback or placeholder settings with
|
||||
# no real hostnames**. The turns above needed a real narrator, so the
|
||||
# endpoint is reset to loopback once the story exists — before the census is
|
||||
# taken and before either bundle is exported.
|
||||
app.call("PUT", "/settings", {
|
||||
"endpoint_url": PLACEHOLDER_ENDPOINT, "model": MODEL,
|
||||
"embedding_model": EMBED_MODEL,
|
||||
"max_output_tokens": 300, "model_timeout_seconds": TURN_TIMEOUT})
|
||||
return adv
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--v100", required=True, help="the v1.0.0 worktree")
|
||||
parser.add_argument("--out", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL and EMBED_MODEL):
|
||||
print("set AIDND_TEST_ENDPOINT, AIDND_TEST_MODEL and AIDND_TEST_EMBED_MODEL")
|
||||
return 2
|
||||
|
||||
out = require_under_home(Path(args.out).expanduser())
|
||||
shutil.rmtree(out, ignore_errors=True)
|
||||
out.mkdir(parents=True)
|
||||
v100_tree = Path(args.v100).expanduser().resolve()
|
||||
candidate_tree = Path(__file__).resolve().parent.parent.parent
|
||||
db = out / "campaign.db"
|
||||
results: dict = {"started": datetime.now().isoformat(timespec="seconds"),
|
||||
"v100_tree": str(v100_tree), "candidate": str(candidate_tree)}
|
||||
failures: list[str] = []
|
||||
|
||||
print("phase 1 — v1.0.0 builds and plays the campaign")
|
||||
app = App(v100_tree, db, out / "v100-server.log", "v1.0.0")
|
||||
try:
|
||||
adv = build_v100_campaign(app)
|
||||
before = census(app, adv, db)
|
||||
v100_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
|
||||
(out / "v100-export.json").write_text(json.dumps(v100_bundle))
|
||||
finally:
|
||||
app.stop()
|
||||
results["adventure"], results["before"] = adv, before
|
||||
print(f" {before['total']} actions, redo={before['can_redo']}, "
|
||||
f"checkpoints={len(before['checkpoints'])}, memories={len(before['memories'])}, "
|
||||
f"summaries={len(before['summaries'])}, knowledge={len(before['knowledge'])}, "
|
||||
f"schema={before['schema_version']}")
|
||||
|
||||
print("\nphase 2 — the candidate opens that same database file")
|
||||
app = App(candidate_tree, db, out / "candidate-server.log", "candidate")
|
||||
try:
|
||||
after = census(app, adv, db)
|
||||
v11_bundle = app.call("GET", f"/adventures/{adv}/export", timeout=600)
|
||||
(out / "v11-export.json").write_text(json.dumps(v11_bundle))
|
||||
try:
|
||||
imported = app.call("POST", "/adventures/import", v100_bundle, timeout=600)
|
||||
results["v100_bundle_into_v11"] = {"status": "imported",
|
||||
"id": (imported or {}).get("id")}
|
||||
print(" the v1.0.0 bundle imported into the candidate")
|
||||
except urllib.error.HTTPError as exc:
|
||||
results["v100_bundle_into_v11"] = {
|
||||
"status": "refused", "code": exc.code,
|
||||
"detail": exc.read().decode()[:300]}
|
||||
failures.append("v100_bundle_into_v11")
|
||||
print(f" the candidate REFUSED the v1.0.0 bundle: {exc.code}")
|
||||
finally:
|
||||
app.stop()
|
||||
results["after"] = after
|
||||
|
||||
print(" comparing the census, field by field:")
|
||||
for field in CENSUS:
|
||||
same = before.get(field) == after.get(field)
|
||||
print(f" {'ok ' if same else 'DIFF'} {field}")
|
||||
if not same:
|
||||
failures.append(field)
|
||||
results.setdefault("differences", {})[field] = {
|
||||
"before": before.get(field), "after": after.get(field)}
|
||||
|
||||
print("\nphase 3 — the candidate's bundle offered back to v1.0.0")
|
||||
app = App(v100_tree, out / "backward.db", out / "v100-backward.log", "v1.0.0")
|
||||
try:
|
||||
try:
|
||||
back = app.call("POST", "/adventures/import", v11_bundle, timeout=600)
|
||||
results["v11_bundle_into_v100"] = {"status": "imported",
|
||||
"id": (back or {}).get("id")}
|
||||
print(" v1.0.0 ACCEPTED the v1.1 bundle")
|
||||
except urllib.error.HTTPError as exc:
|
||||
detail = exc.read().decode()[:400]
|
||||
results["v11_bundle_into_v100"] = {"status": "refused", "code": exc.code,
|
||||
"detail": detail}
|
||||
# Reported, not failed: §11 item 9 asks for the result, and the
|
||||
# owner's brief asks whether a refusal breaks the compatibility
|
||||
# promise — a judgement, not an assertion this script may make.
|
||||
print(f" v1.0.0 REFUSED the v1.1 bundle: {exc.code} {detail[:160]}")
|
||||
finally:
|
||||
app.stop()
|
||||
|
||||
results["failures"] = failures
|
||||
(out / "upgrade-report.json").write_text(json.dumps(results, indent=2, default=str))
|
||||
print(f"\n{'PASS' if not failures else 'FAIL'}: {len(failures)} field(s) differ "
|
||||
f"-> {out / 'upgrade-report.json'}")
|
||||
return 1 if failures else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,211 @@
|
||||
"""v1.1 WP-A1: real turns, and what the server said it read.
|
||||
|
||||
AIDND_TEST_ENDPOINT=https://<host>:<port>/v1 AIDND_TEST_MODEL=<model> \\
|
||||
.venv/bin/python -m tools.v11_window_accounting \\
|
||||
--bundle "$HOME/m11-evidence/m04-final/bundle.json" --turns 4 \\
|
||||
--out "$HOME/v11-evidence/a1-accounting/<label>"
|
||||
|
||||
Run from `backend/`. The evidence that matters is at the edge of the window, and
|
||||
a new campaign takes dozens of turns to reach it. So this imports a long
|
||||
campaign, by default the v1 evidence run's 207-action bundle, and every turn is
|
||||
assembled against a full window from the first. A real narrator is then asked
|
||||
for `--turns` turns, and each one is written out with:
|
||||
|
||||
configured budget, verified window, and where the window came from
|
||||
the application's estimate of what it sent (its count, plus the text the
|
||||
provider adds)
|
||||
the server's own prompt-token count, from the usage it reported
|
||||
the reply allocation and the safety reserve
|
||||
the observed margin: window - reply allocation - the server's count
|
||||
the accounting status: fits, exceeded, truncation_suspected or unknown
|
||||
|
||||
The first turn on a cold model cannot verify the window: `/api/ps` knows nothing
|
||||
until the model is loaded. That is M11's behaviour, and the row says so rather
|
||||
than hiding the turn.
|
||||
|
||||
The database lives in `--out`, not in `/tmp`, so the stored snapshots behind
|
||||
every row can be read again. The endpoint is read from the environment and is
|
||||
never written into a committed file.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
ENDPOINT = os.environ.get("AIDND_TEST_ENDPOINT", "")
|
||||
MODEL = os.environ.get("AIDND_TEST_MODEL", "")
|
||||
|
||||
TURNS = [
|
||||
"I look around carefully and take stock of where I am.",
|
||||
"I ask the nearest person what has happened since I was last here.",
|
||||
"I check what I am carrying.",
|
||||
"I move on towards the place I meant to reach.",
|
||||
"I wait and listen.",
|
||||
"I say, \"Tell me the part you left out.\"",
|
||||
]
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
|
||||
parser.add_argument("--bundle", default="",
|
||||
help="a campaign bundle to import, so turns start at a full window")
|
||||
parser.add_argument("--turns", type=int, default=4)
|
||||
parser.add_argument("--budget", type=int, default=16384)
|
||||
parser.add_argument("--max-output", type=int, default=500)
|
||||
parser.add_argument("--timeout", type=int, default=900)
|
||||
parser.add_argument("--out", required=True)
|
||||
parser.add_argument("--unload-first", action="store_true",
|
||||
help="ask the configured server to unload the model before turn 1, "
|
||||
"so the first turn starts cold (the A1 corrective test)")
|
||||
args = parser.parse_args()
|
||||
|
||||
if not (ENDPOINT and MODEL):
|
||||
print("set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL")
|
||||
return 2
|
||||
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
db_path = out / "accounting.db"
|
||||
if db_path.exists():
|
||||
print(f"{db_path} exists; choose a new --out")
|
||||
return 2
|
||||
os.environ["AIDND_DB_PATH"] = str(db_path)
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
from app import auth, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
Base.metadata.create_all(bind=engine)
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email="accounting@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model=MODEL, endpoint_url=ENDPOINT, embedding_model="",
|
||||
context_token_budget=args.budget, max_output_tokens=args.max_output,
|
||||
model_timeout_seconds=args.timeout,
|
||||
))
|
||||
db.commit()
|
||||
user_id = user.id
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
client = TestClient(app)
|
||||
|
||||
if args.bundle:
|
||||
bundle = json.loads(Path(args.bundle).read_text())
|
||||
# The evidence campaign had its memory bank and auto-summarise on. Here
|
||||
# they would only add post-turn model calls between the measured turns,
|
||||
# on the same host, and fail noisily as the in-process client closes.
|
||||
# This tool measures the turn's own prompt, which neither changes.
|
||||
bundle["memoryBankEnabled"] = False
|
||||
bundle["autoSummarize"] = False
|
||||
imported = client.post("/api/adventures/import", json=bundle)
|
||||
imported.raise_for_status()
|
||||
adv = imported.json()["id"]
|
||||
else:
|
||||
created = client.post("/api/adventures", json={
|
||||
"title": "Window accounting", "opening": "A quiet road at dusk."})
|
||||
created.raise_for_status()
|
||||
adv = created.json()["id"]
|
||||
|
||||
if args.unload_first:
|
||||
# The configured endpoint only, under the same policy and TLS trust as a turn.
|
||||
import asyncio
|
||||
import httpx
|
||||
from app import contextwindow, endpoints, tlstrust
|
||||
reason = endpoints.rejection_reason(ENDPOINT)
|
||||
if reason:
|
||||
print(f"endpoint refused: {reason}")
|
||||
return 2
|
||||
base = contextwindow.native_base(ENDPOINT)
|
||||
with httpx.Client(verify=tlstrust.ssl_context(), timeout=120) as http:
|
||||
unloaded = http.post(f"{base}/api/generate", json={"model": MODEL, "keep_alive": 0})
|
||||
resident = [m.get("name") for m in http.get(f"{base}/api/ps").json().get("models", [])]
|
||||
contextwindow.cache_clear()
|
||||
print(f"unload: HTTP {unloaded.status_code} {unloaded.text[:120]} | resident now: {resident}")
|
||||
|
||||
rows: list[dict] = []
|
||||
timeline = (out / "turns.jsonl").open("a")
|
||||
print(f"model {MODEL}, budget {args.budget}, reply {args.max_output}")
|
||||
print(f"{'#':>2} {'status':22} {'window':>13} {'estimate':>8} {'server':>7} "
|
||||
f"{'reserve':>7} {'margin':>7} {'sec':>5}")
|
||||
for index in range(args.turns):
|
||||
text = TURNS[index % len(TURNS)]
|
||||
started = time.monotonic()
|
||||
response = client.post(f"/api/adventures/{adv}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
seconds = round(time.monotonic() - started, 1)
|
||||
error = None
|
||||
if response.status_code != 200 or '"type": "error"' in response.text:
|
||||
error = response.text[-400:]
|
||||
with SessionLocal() as db:
|
||||
action = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv, models.Action.type == "ai")
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
.order_by(models.Action.id.desc()).first()
|
||||
)
|
||||
snapshot = (action.context_snapshot or {}) if action else {}
|
||||
tokens = snapshot.get("tokens") or {}
|
||||
window = snapshot.get("window") or {}
|
||||
accounting = snapshot.get("accounting") or {}
|
||||
row = {
|
||||
"turn": index + 1,
|
||||
"action_id": action.id if action else None,
|
||||
"seconds": seconds,
|
||||
"error": error,
|
||||
"configured_budget": tokens.get("configured_budget"),
|
||||
"effective_budget": tokens.get("budget"),
|
||||
"window_verified": window.get("verified"),
|
||||
"window_tokens": window.get("tokens"),
|
||||
"window_source": window.get("source"),
|
||||
"preflight_attempted": (window.get("preflight") or {}).get("attempted"),
|
||||
"preflight_loaded": (window.get("preflight") or {}).get("loaded"),
|
||||
"preflight_verified_before": (window.get("preflight") or {}).get("verified_before"),
|
||||
"preflight_verified_after": (window.get("preflight") or {}).get("verified_after"),
|
||||
"preflight_detail": (window.get("preflight") or {}).get("detail"),
|
||||
"app_prompt_tokens": tokens.get("total"),
|
||||
"transport_tokens": tokens.get("transport"),
|
||||
"app_estimate": tokens.get("estimate"),
|
||||
"output_reserve": tokens.get("output_reserve"),
|
||||
"safety_reserve": tokens.get("safety_reserve"),
|
||||
"server_prompt_tokens": accounting.get("server_prompt_tokens"),
|
||||
"difference": accounting.get("difference"),
|
||||
"observed_margin": accounting.get("observed_margin"),
|
||||
"status": accounting.get("status"),
|
||||
"history_included": (snapshot.get("history") or {}).get("included"),
|
||||
"history_total": (snapshot.get("history") or {}).get("total"),
|
||||
}
|
||||
rows.append(row)
|
||||
timeline.write(json.dumps(row) + "\n")
|
||||
timeline.flush()
|
||||
print(f"{row['turn']:>2} {str(row['status'] if not error else 'ERROR'):22} "
|
||||
f"{str(row['window_tokens'])+('v' if row['window_verified'] else '?'):>13} "
|
||||
f"{str(row['app_estimate']):>8} {str(row['server_prompt_tokens']):>7} "
|
||||
f"{str(row['safety_reserve']):>7} {str(row['observed_margin']):>7} {seconds:>5}")
|
||||
if error:
|
||||
print(f" error: {error[:200]}")
|
||||
timeline.close()
|
||||
(out / "summary.json").write_text(json.dumps({
|
||||
"model": MODEL, "budget": args.budget, "max_output_tokens": args.max_output,
|
||||
"bundle": args.bundle, "rows": rows,
|
||||
}, indent=2))
|
||||
app.dependency_overrides.clear()
|
||||
return 0 if all(r["error"] is None for r in rows) else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,353 @@
|
||||
"""v1.1 WP-D criterion 3: the backup completes **through the UI**.
|
||||
|
||||
python -m tools.wpd_backup_ui --out <dir under $HOME> [--case real|large|both] [--show]
|
||||
|
||||
Run from `backend/`, with `frontend/dist` already built.
|
||||
|
||||
WP-D's first pass drove `POST /api/backups` — the endpoint the *Back up now*
|
||||
button calls — on a 2.3 MB campaign database and a 117 MB one. That is evidence
|
||||
about the implementation, and the plan's criterion 3 asks for something else:
|
||||
that the backup *completes through the UI* on both. A reader does not call an
|
||||
endpoint. This drives the reader-facing control in a real Firefox, against the
|
||||
production build served by FastAPI, exactly as WP-C's harness does.
|
||||
|
||||
**No narrator and no inference.** A backup needs neither, so nothing here
|
||||
touches a model host.
|
||||
|
||||
What it refuses to call a pass, per the brief:
|
||||
|
||||
- the click does nothing (no toast, no file);
|
||||
- the request fails (an error toast);
|
||||
- no backup file appears on disk;
|
||||
- the finished copy does not pass a full `PRAGMA integrity_check`;
|
||||
- the UI reports an error.
|
||||
|
||||
Every wait is on a condition the page or the filesystem can show. Nothing here
|
||||
sleeps and then asserts.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import shutil
|
||||
import sqlite3
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from tools.m11_webdriver import ( # noqa: E402
|
||||
Browser, Site, geckodriver_version, require_under_home,
|
||||
)
|
||||
|
||||
BACKEND = Path(__file__).resolve().parent.parent
|
||||
HOME = Path.home()
|
||||
|
||||
#: The two databases criterion 3 names. Both are copied before use: the first is
|
||||
#: the M11 evidence campaign and must not be written to, and the second is the
|
||||
#: 117 MB application database built for WP-D's timing.
|
||||
DEFAULT_REAL = HOME / "m11-evidence/m04-final/campaign.db"
|
||||
DEFAULT_LARGE = HOME / "v11-evidence/wp-d/endpoint/campaign-100mb/campaign.db"
|
||||
|
||||
BLOCK = '[data-testid="database-backup"]'
|
||||
BUTTON = f'{BLOCK} button.primary'
|
||||
|
||||
|
||||
class Checks:
|
||||
"""Results, with the discipline that an unrun check is not a passing one."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.rows: list[dict] = []
|
||||
|
||||
def record(self, case: str, name: str, ok: bool, detail: str = "") -> bool:
|
||||
self.rows.append({"case": case, "check": name,
|
||||
"result": "PASS" if ok else "FAIL", "detail": detail})
|
||||
print(f" {'ok ' if ok else 'FAIL'} {case:6} {name}"
|
||||
+ (f" — {detail}" if detail else ""), flush=True)
|
||||
return ok
|
||||
|
||||
@property
|
||||
def failed(self) -> list[dict]:
|
||||
return [r for r in self.rows if r["result"] == "FAIL"]
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- helpers
|
||||
|
||||
def integrity_of(path: Path) -> tuple[str, float]:
|
||||
"""The full check on a finished copy, and what it cost, measured here.
|
||||
|
||||
The application runs its own `integrity_check` before keeping the file; this
|
||||
is an independent second opinion on the artefact the UI produced, and it is
|
||||
where criterion 3's 'time for integrity_check' comes from.
|
||||
"""
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
try:
|
||||
started = time.perf_counter()
|
||||
rows = connection.execute("PRAGMA integrity_check").fetchall()
|
||||
elapsed = time.perf_counter() - started
|
||||
finally:
|
||||
connection.close()
|
||||
return ", ".join(str(r[0]) for r in rows), elapsed
|
||||
|
||||
|
||||
def digest(path: Path) -> str:
|
||||
return hashlib.sha256(path.read_bytes()).hexdigest()[:16]
|
||||
|
||||
|
||||
def backups_in(db_path: Path) -> dict[str, int]:
|
||||
directory = db_path.parent / "backups"
|
||||
if not directory.exists():
|
||||
return {}
|
||||
return {p.name: p.stat().st_size for p in directory.iterdir() if p.is_file()}
|
||||
|
||||
|
||||
def wait_for_new_backup(db_path: Path, before: set[str], *, timeout: float = 300,
|
||||
poll: float = 0.2) -> Path | None:
|
||||
"""The backup file the application wrote, once it is really there.
|
||||
|
||||
A file condition, not a sleep: a name that was not there before, not a
|
||||
`.partial`, and a size that has stopped growing.
|
||||
"""
|
||||
directory = db_path.parent / "backups"
|
||||
deadline = time.monotonic() + timeout
|
||||
sizes: dict[str, int] = {}
|
||||
while time.monotonic() < deadline:
|
||||
if directory.exists():
|
||||
for path in directory.iterdir():
|
||||
if not path.is_file() or path.name in before:
|
||||
continue
|
||||
if path.name.endswith(".partial"):
|
||||
continue
|
||||
size = path.stat().st_size
|
||||
if size > 0 and sizes.get(path.name) == size:
|
||||
return path
|
||||
sizes[path.name] = size
|
||||
time.sleep(poll)
|
||||
return None
|
||||
|
||||
|
||||
def toasts(browser: Browser) -> list[dict]:
|
||||
"""What the page is telling the reader — the message apart from its mark.
|
||||
|
||||
A toast renders a decorative mark before the message:
|
||||
`<span class="toast-mark" aria-hidden="true">❖</span><span>…</span>`. So the
|
||||
button's `textContent` begins with that character, and the first run of this
|
||||
tool matched `textContent.startswith('Backup written:')` and reported six
|
||||
*passing* behaviours as failures — the UI had said exactly what it should,
|
||||
and the assertion was reading the mark. `message` is the message span alone;
|
||||
`text` is kept whole for the evidence record.
|
||||
"""
|
||||
return browser.js("""
|
||||
const host = document.querySelector('.toast-host');
|
||||
if (!host) return [];
|
||||
return [...host.querySelectorAll('button.toast')].map(b => {
|
||||
const span = b.querySelector('span:not(.toast-mark)');
|
||||
return {
|
||||
error: b.classList.contains('toast-error'),
|
||||
message: (span ? span.textContent : b.textContent).trim(),
|
||||
text: b.textContent.trim(),
|
||||
};
|
||||
});
|
||||
""") or []
|
||||
|
||||
|
||||
def open_backup_panel(browser: Browser, site: Site, checks: Checks, case: str) -> bool:
|
||||
"""Reach the control the way a reader does: the nav link, then the panel."""
|
||||
browser.go(site.url)
|
||||
browser.wait_for(".topnav", timeout=60)
|
||||
link = browser.find('.nav-links a[href="/settings"]', required=False)
|
||||
if link is None:
|
||||
return checks.record(case, "the Settings link is in the navigation", False)
|
||||
browser.click(link)
|
||||
opened = browser.wait_js(f"!!document.querySelector('{BLOCK}')", timeout=60)
|
||||
if not checks.record(case, "Settings opens from the navigation link", opened,
|
||||
browser.url):
|
||||
return False
|
||||
|
||||
summary = browser.find(f"{BLOCK} summary", required=False)
|
||||
if summary is None:
|
||||
return checks.record(case, "the backup panel has a disclosure", False)
|
||||
browser.click(summary)
|
||||
# The panel is a <details>: it loads what is already on disk when it opens,
|
||||
# so waiting for the directory line proves the application answered.
|
||||
shown = browser.wait_js(
|
||||
f"document.querySelector('{BLOCK}').open === true"
|
||||
f" && !!document.querySelector('{BUTTON}')", timeout=30)
|
||||
checks.record(case, "the backup panel opens and shows its control", shown)
|
||||
listed = browser.wait_js(f"!!document.querySelector('{BLOCK} code')", timeout=30)
|
||||
checks.record(case, "the panel reports where backups are written", listed)
|
||||
return shown
|
||||
|
||||
|
||||
def click_back_up_now(browser: Browser, site: Site, db: Path, checks: Checks,
|
||||
case: str, label: str) -> dict:
|
||||
"""One press of the button, judged by what the page and the disk then show."""
|
||||
before = set(backups_in(db))
|
||||
# Clear anything still on screen, so the toast this press produces is the
|
||||
# one that is read back rather than a leftover from the previous press.
|
||||
browser.js("""
|
||||
for (const b of document.querySelectorAll('.toast-host button.toast')) b.click();
|
||||
return true;
|
||||
""")
|
||||
button = browser.find(BUTTON, required=False)
|
||||
if button is None:
|
||||
checks.record(case, f"{label}: the control is on the page", False)
|
||||
return {}
|
||||
|
||||
started = time.perf_counter()
|
||||
browser.click(button)
|
||||
|
||||
# Either outcome ends the wait, so a failure is reported as a failure rather
|
||||
# than as a timeout.
|
||||
settled = browser.wait_js(
|
||||
"(() => { const t = [...document.querySelectorAll('.toast-host button.toast')];"
|
||||
" return t.length > 0; })()", timeout=300)
|
||||
elapsed = time.perf_counter() - started
|
||||
|
||||
shown = toasts(browser)
|
||||
errors = [t for t in shown if t["error"]]
|
||||
written = [t for t in shown if t["message"].startswith("Backup written:")]
|
||||
|
||||
checks.record(case, f"{label}: the click produced a visible result", settled,
|
||||
json.dumps(shown)[:200])
|
||||
checks.record(case, f"{label}: the UI reports no error",
|
||||
not errors, json.dumps(errors)[:300])
|
||||
checks.record(case, f"{label}: the UI reports the backup was written",
|
||||
bool(written), json.dumps(written)[:200])
|
||||
|
||||
idle = browser.wait_js(
|
||||
"(() => { const b = document.querySelector(%s);"
|
||||
" return !!b && b.textContent.trim() === 'Back up now'; })()"
|
||||
% json.dumps(BUTTON), timeout=120)
|
||||
checks.record(case, f"{label}: the control returns from 'Backing up…'", idle)
|
||||
|
||||
produced = wait_for_new_backup(db, before)
|
||||
checks.record(case, f"{label}: a backup file was physically produced",
|
||||
produced is not None, str(produced))
|
||||
if produced is None:
|
||||
return {"seconds": round(elapsed, 3), "toasts": shown}
|
||||
|
||||
named = any(produced.name in t["message"] for t in written)
|
||||
checks.record(case, f"{label}: the UI names the file that appeared", named,
|
||||
f"{produced.name} — {written[0]['message'] if written else ''}")
|
||||
|
||||
relisted = browser.wait_js(
|
||||
"[...document.querySelectorAll('%s .backup-list code')]"
|
||||
".some(c => c.textContent.trim() === %s)" % (BLOCK, json.dumps(produced.name)),
|
||||
timeout=60)
|
||||
checks.record(case, f"{label}: the new backup appears in the panel's list", relisted)
|
||||
|
||||
verdict, integrity_seconds = integrity_of(produced)
|
||||
checks.record(case, f"{label}: the finished copy passes full integrity_check",
|
||||
verdict == "ok", f"{verdict} in {integrity_seconds * 1000:.1f} ms")
|
||||
|
||||
return {
|
||||
"file": str(produced),
|
||||
"bytes": produced.stat().st_size,
|
||||
"sha256_16": digest(produced),
|
||||
"seconds": round(elapsed, 3),
|
||||
"integrity": verdict,
|
||||
"integrity_seconds": round(integrity_seconds, 4),
|
||||
"toasts": shown,
|
||||
}
|
||||
|
||||
|
||||
def run_case(case: str, source: Path, out: Path, checks: Checks, *, show: bool,
|
||||
twice: bool) -> dict:
|
||||
print(f"\n=== {case}: {source} ===", flush=True)
|
||||
work = out / case
|
||||
shutil.rmtree(work, ignore_errors=True)
|
||||
(work / "data").mkdir(parents=True)
|
||||
db = work / "data" / "campaign.db"
|
||||
copy_started = time.perf_counter()
|
||||
shutil.copy2(source, db)
|
||||
copy_seconds = time.perf_counter() - copy_started
|
||||
size = db.stat().st_size
|
||||
print(f" copied {size:,} bytes in {copy_seconds:.2f}s -> {db}", flush=True)
|
||||
|
||||
site = Site(BACKEND, db, work / "server.log")
|
||||
browser = Browser(headless=not show, log=work / "geckodriver.log")
|
||||
result: dict = {"database": str(source), "bytes": size,
|
||||
"served_at": site.url, "firefox": browser.version}
|
||||
try:
|
||||
if not open_backup_panel(browser, site, checks, case):
|
||||
return result
|
||||
result["first"] = click_back_up_now(browser, site, db, checks, case, "backup")
|
||||
browser.screenshot(work / "back-up-now.png")
|
||||
|
||||
if twice and result["first"].get("file"):
|
||||
kept = Path(result["first"]["file"])
|
||||
before_bytes, before_digest = kept.stat().st_size, digest(kept)
|
||||
result["second"] = click_back_up_now(browser, site, db, checks, case,
|
||||
"second backup")
|
||||
still_there = kept.exists()
|
||||
checks.record(case, "the earlier backup still exists", still_there)
|
||||
if still_there:
|
||||
checks.record(
|
||||
case, "and is byte-for-byte what it was",
|
||||
kept.stat().st_size == before_bytes and digest(kept) == before_digest,
|
||||
f"{before_bytes:,} bytes, sha256:{before_digest}")
|
||||
verdict, _ = integrity_of(kept)
|
||||
checks.record(case, "and still passes integrity_check", verdict == "ok",
|
||||
verdict)
|
||||
if result["second"].get("file"):
|
||||
checks.record(case, "the second backup is a different file",
|
||||
result["second"]["file"] != result["first"]["file"],
|
||||
Path(result["second"]["file"]).name)
|
||||
finally:
|
||||
browser.quit()
|
||||
site.stop()
|
||||
return result
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--out", required=True,
|
||||
help="evidence directory, which must be under $HOME")
|
||||
parser.add_argument("--case", default="both", choices=["real", "large", "both"])
|
||||
parser.add_argument("--show", action="store_true", help="run Firefox visibly")
|
||||
parser.add_argument("--real-db", default=str(DEFAULT_REAL))
|
||||
parser.add_argument("--large-db", default=str(DEFAULT_LARGE))
|
||||
args = parser.parse_args()
|
||||
|
||||
out = require_under_home(Path(args.out).expanduser())
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
if not (BACKEND.parent / "frontend/dist/index.html").exists():
|
||||
print("frontend/dist is not built", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
checks = Checks()
|
||||
started = datetime.now()
|
||||
results: dict[str, dict] = {}
|
||||
wanted = [("real", Path(args.real_db).expanduser(), True)] if args.case != "large" else []
|
||||
if args.case != "real":
|
||||
wanted.append(("large", Path(args.large_db).expanduser(), False))
|
||||
|
||||
print(f"WP-D criterion 3 — the backup through the UI. geckodriver "
|
||||
f"{geckodriver_version()}", flush=True)
|
||||
for case, source, twice in wanted:
|
||||
if not source.exists():
|
||||
checks.record(case, "the database is present", False, str(source))
|
||||
continue
|
||||
results[case] = run_case(case, source, out, checks, show=args.show, twice=twice)
|
||||
|
||||
report = {
|
||||
"started": started.isoformat(timespec="seconds"),
|
||||
"seconds": round((datetime.now() - started).total_seconds()),
|
||||
"cases": results,
|
||||
"checks": checks.rows,
|
||||
"passed": len([r for r in checks.rows if r["result"] == "PASS"]),
|
||||
"failed": len(checks.failed),
|
||||
}
|
||||
(out / "backup-ui-report.json").write_text(json.dumps(report, indent=2))
|
||||
print(f"\n{report['passed']} passed, {report['failed']} failed "
|
||||
f"-> {out / 'backup-ui-report.json'}")
|
||||
return 1 if checks.failed else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
+26
-1
@@ -151,7 +151,32 @@ export const api = {
|
||||
sendAction: (advId, payload, handlers, signal) =>
|
||||
streamSSE(`/adventures/${advId}/actions`, payload, handlers, signal),
|
||||
retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal),
|
||||
exportAdventure: (id) => request(`/adventures/${id}/export`),
|
||||
// v1.1 WP-D: the bundle, plus what the server says about importing it back.
|
||||
// The body is the bundle and nothing else — the browser saves exactly those
|
||||
// bytes — so the size and the warning come back in headers.
|
||||
exportAdventure: async (id) => {
|
||||
const resp = await fetch(`/api/adventures/${id}/export`, {
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
})
|
||||
if (!resp.ok) {
|
||||
let detail = resp.statusText
|
||||
try { detail = (await resp.json()).detail || detail } catch { /* non-JSON */ }
|
||||
throw new Error(detail)
|
||||
}
|
||||
const bundle = await resp.json()
|
||||
const number = (name) => {
|
||||
const raw = Number(resp.headers.get(name))
|
||||
return Number.isFinite(raw) && raw > 0 ? raw : null
|
||||
}
|
||||
return {
|
||||
bundle,
|
||||
exportBytes: number('X-Export-Bytes'),
|
||||
importLimitBytes: number('X-Import-Limit-Bytes'),
|
||||
// Absent header (an older server) means nothing is claimed either way.
|
||||
importable: resp.headers.get('X-Importable-By-This-Version') !== 'false',
|
||||
warning: resp.headers.get('X-Export-Warning') || null,
|
||||
}
|
||||
},
|
||||
importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }),
|
||||
|
||||
// M9. A verified copy of the whole database, which is a different tool from
|
||||
|
||||
@@ -45,6 +45,25 @@ export function classifyError(message) {
|
||||
const detail = String(message || '').trim() || 'No detail was reported.'
|
||||
const low = detail.toLowerCase()
|
||||
|
||||
// ---- A correction the story refused ----
|
||||
//
|
||||
// v1.1 WP-C, found in the browser. The State panel's refusals ("That
|
||||
// correction can't be applied — no fact 'f1' to invalidate.") matched none of
|
||||
// the rules below and fell through to "Generation failed", with a Retry button
|
||||
// and a line about what you typed. No turn was attempted and nothing was
|
||||
// typed: the reader corrected the story's state and the rules refused it. So
|
||||
// it is said as that, and the reason stays under the technical details.
|
||||
if (low.includes("correction can't be applied")) {
|
||||
return {
|
||||
kind: KIND.STATE,
|
||||
title: 'That correction was not applied',
|
||||
detail,
|
||||
hint: 'Nothing in the story or its state was changed. The reason is in the technical details.',
|
||||
retryable: false,
|
||||
keptInput: false,
|
||||
}
|
||||
}
|
||||
|
||||
// ---- Model / endpoint: the story cannot be told at all ----
|
||||
|
||||
// "No model configured — set one in Settings."
|
||||
|
||||
@@ -66,10 +66,12 @@ export default function Campaigns() {
|
||||
|
||||
const exportOne = async (campaign) => {
|
||||
try {
|
||||
const bundle = await api.exportAdventure(campaign.id)
|
||||
const { bundle, warning } = await api.exportAdventure(campaign.id)
|
||||
const safe = (campaign.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
|
||||
downloadJSON(bundle, `${safe}.json`)
|
||||
toast('Campaign exported.')
|
||||
// v1.1 WP-D: the file is written either way. A campaign too large for this
|
||||
// version to import back says so now, not when it is needed.
|
||||
toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
|
||||
} catch (err) {
|
||||
toast(classifyError(err.message).detail, 'error')
|
||||
}
|
||||
|
||||
@@ -55,9 +55,13 @@ export function FailureNotice({ failure, partial, onRetry, onDismiss }) {
|
||||
|
||||
{failure.hint && <p className="failure-hint">{failure.hint}</p>}
|
||||
|
||||
<p className="failure-kept">
|
||||
Your story is unchanged, and what you typed is still in the box below.
|
||||
</p>
|
||||
{/* Every failed turn keeps what was typed (A05). A refused correction
|
||||
had nothing typed in the box, so it does not claim to (v1.1 WP-C). */}
|
||||
{failure.keptInput !== false && (
|
||||
<p className="failure-kept">
|
||||
Your story is unchanged, and what you typed is still in the box below.
|
||||
</p>
|
||||
)}
|
||||
|
||||
{partial ? (
|
||||
<details className="failure-partial">
|
||||
|
||||
@@ -29,8 +29,8 @@ beforeEach(() => { vi.restoreAllMocks() })
|
||||
|
||||
describe('the failure notice states only what is true', () => {
|
||||
it('claims the typed words were kept — so the caller must actually keep them', async () => {
|
||||
// This assertion is the contract the defect broke. The notice is
|
||||
// unconditional, so every path that shows it owes the reader their text.
|
||||
// This assertion is the contract the defect broke. Every failed turn shows
|
||||
// it, so every path that shows it owes the reader their text.
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError('Could not connect to http://127.0.0.1:9/v1')}
|
||||
partial={null} onRetry={vi.fn()} onDismiss={vi.fn()} />,
|
||||
@@ -64,6 +64,39 @@ describe('the failure notice states only what is true', () => {
|
||||
expect(onRetry).toHaveBeenCalled()
|
||||
})
|
||||
|
||||
// v1.1 WP-C, found in the browser: a correction the story refused fell through
|
||||
// to "Generation failed", offered to try the turn again, and claimed the
|
||||
// typed words were kept — none of which a State-panel correction involves.
|
||||
const REFUSAL = "That correction can't be applied — no fact 'f1' to invalidate."
|
||||
|
||||
it('classifies a refused correction as a state refusal, not a failed turn', () => {
|
||||
const refusal = classifyError(REFUSAL)
|
||||
expect(refusal.kind).toBe(KIND.STATE)
|
||||
expect(refusal.title).toBe('That correction was not applied')
|
||||
expect(refusal.retryable).toBe(false)
|
||||
expect(refusal.keptInput).toBe(false)
|
||||
expect(refusal.detail).toBe(REFUSAL)
|
||||
})
|
||||
|
||||
it('shows a refused correction with its reason, and no turn to retry', async () => {
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError(REFUSAL)} partial={null}
|
||||
onRetry={vi.fn()} onDismiss={vi.fn()} />,
|
||||
)
|
||||
expect(screen.getByText('That correction was not applied')).toBeInTheDocument()
|
||||
expect(screen.getByText(/no fact 'f1' to invalidate/)).toBeInTheDocument()
|
||||
expect(screen.queryByRole('button', { name: 'Try that turn again' })).toBeNull()
|
||||
expect(screen.queryByText(/what you typed is still in the box/)).toBeNull()
|
||||
})
|
||||
|
||||
it('still claims the typed words were kept for a failed turn', async () => {
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError('boom')} partial={null}
|
||||
onRetry={vi.fn()} onDismiss={vi.fn()} />,
|
||||
)
|
||||
expect(screen.getByText(/what you typed is still in the box/)).toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('labels partial prose as not kept rather than showing it as story', async () => {
|
||||
await renderWith(
|
||||
<FailureNotice failure={classifyError('boom')} partial="The tavern door swung"
|
||||
|
||||
@@ -79,10 +79,12 @@ export function CampaignSettingsPanel({ adventure, setAdventure, onError, moment
|
||||
|
||||
const exportCampaign = async () => {
|
||||
try {
|
||||
const bundle = await api.exportAdventure(adventure.id)
|
||||
const { bundle, warning } = await api.exportAdventure(adventure.id)
|
||||
const safe = (adventure.title || 'campaign').replace(/[^\w-]+/g, '_').slice(0, 60)
|
||||
downloadJSON(bundle, `${safe}.json`)
|
||||
toast('Campaign exported.')
|
||||
// v1.1 WP-D: see Campaigns.jsx. The export is delivered; the warning says
|
||||
// this version could not import the file back.
|
||||
toast(warning || 'Campaign exported.', warning ? 'error' : undefined)
|
||||
} catch (err) {
|
||||
onError(classifyError(err.message).detail)
|
||||
}
|
||||
|
||||
@@ -47,6 +47,63 @@ function Section({ title, count, children, open = false, testId }) {
|
||||
)
|
||||
}
|
||||
|
||||
/* v1.1 WP-A1: what the server said it read, set against what was sent.
|
||||
*
|
||||
* Only a turn that was actually sent has this — the next-turn view has not
|
||||
* been sent yet. The two conditions that mean something went wrong are shown
|
||||
* as alerts, because the failure they describe is otherwise silent: Ollama
|
||||
* answers 200 whether or not it cut the front of the prompt off. */
|
||||
const ACCOUNTING = {
|
||||
fits: {
|
||||
title: 'The server read the whole prompt',
|
||||
body: 'Its count stayed inside the room kept for the reply and the safety margin.',
|
||||
},
|
||||
exceeded: {
|
||||
title: 'The prompt was larger than the server allowed for',
|
||||
body: 'The server counted more tokens than the safety margin covers, so the '
|
||||
+ 'reply may have been cut short. The turn is kept.',
|
||||
alert: true,
|
||||
},
|
||||
truncation_suspected: {
|
||||
title: 'The server may have cut the start of the prompt',
|
||||
body: 'It read far fewer tokens than were sent, which is what happens when a '
|
||||
+ 'prompt is larger than the window the model was loaded with. The '
|
||||
+ 'narrator’s rules and the campaign canon are at the start. The turn is kept.',
|
||||
alert: true,
|
||||
},
|
||||
unknown: {
|
||||
title: 'The server did not say how much it read',
|
||||
body: 'Nothing here can confirm whether the whole prompt was used.',
|
||||
},
|
||||
}
|
||||
|
||||
function AccountingReport({ accounting }) {
|
||||
if (!accounting) return null
|
||||
const copy = ACCOUNTING[accounting.status] || ACCOUNTING.unknown
|
||||
return (
|
||||
<div
|
||||
className={copy.alert ? 'notice error' : 'ctx-accounting'}
|
||||
role={copy.alert ? 'alert' : undefined}
|
||||
data-testid="ctx-accounting"
|
||||
data-status={accounting.status}
|
||||
>
|
||||
<strong>{copy.title}</strong>
|
||||
<p>{copy.body}</p>
|
||||
{accounting.server_prompt_tokens != null && (
|
||||
<p className="notice-detail">
|
||||
Sent {accounting.estimate?.toLocaleString()} by this app’s count; the
|
||||
server read {accounting.server_prompt_tokens.toLocaleString()}.
|
||||
{accounting.observed_margin != null
|
||||
&& ` ${accounting.observed_margin.toLocaleString()} tokens were left beside the reply.`}
|
||||
</p>
|
||||
)}
|
||||
{accounting.window_verified === false && (
|
||||
<p className="notice-detail">The model’s window was not verified for this turn.</p>
|
||||
)}
|
||||
</div>
|
||||
)
|
||||
}
|
||||
|
||||
function TokenBar({ sections, total }) {
|
||||
if (!sections.length || total <= 0) return null
|
||||
return (
|
||||
@@ -122,6 +179,12 @@ export function ContextPanel({
|
||||
<span>{tokens.output_reserve.toLocaleString()}</span>
|
||||
</div>
|
||||
)}
|
||||
{tokens.safety_reserve > 0 && (
|
||||
<div className="ctx-token-line dim">
|
||||
<span>Kept free as a safety margin</span>
|
||||
<span>{tokens.safety_reserve.toLocaleString()}</span>
|
||||
</div>
|
||||
)}
|
||||
<div className="ctx-token-line dim">
|
||||
<span>What the model can hold</span>
|
||||
<span>{tokens.budget.toLocaleString()}</span>
|
||||
@@ -135,6 +198,8 @@ export function ContextPanel({
|
||||
)}
|
||||
</div>
|
||||
|
||||
<AccountingReport accounting={report.accounting} />
|
||||
|
||||
{failing.length > 0 && (
|
||||
<div className="notice error" role="alert" data-testid="ctx-derived-failing">
|
||||
<strong>Background work is failing</strong>
|
||||
|
||||
@@ -220,6 +220,50 @@ describe('context inspector (§23, §55)', () => {
|
||||
expect(tokens).toHaveTextContent('8,000')
|
||||
})
|
||||
|
||||
it('shows the safety margin kept free beside the reply (v1.1 A1)', async () => {
|
||||
api.getAdventureContext.mockResolvedValue({
|
||||
...REPORT, tokens: { ...REPORT.tokens, safety_reserve: 410 },
|
||||
})
|
||||
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
|
||||
expect(screen.getByTestId('ctx-tokens')).toHaveTextContent(/safety margin\s*410/)
|
||||
})
|
||||
|
||||
it('says nothing about accounting for a turn that has not been sent', async () => {
|
||||
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
|
||||
expect(screen.queryByTestId('ctx-accounting')).toBeNull()
|
||||
})
|
||||
|
||||
it('alerts when the server may have cut the start of a sent prompt (v1.1 A1)', async () => {
|
||||
vi.spyOn(api, 'getActionContext').mockResolvedValue({
|
||||
...REPORT,
|
||||
accounting: {
|
||||
status: 'truncation_suspected', estimate: 6316, server_prompt_tokens: 2050,
|
||||
observed_margin: 1546, window_verified: true,
|
||||
},
|
||||
})
|
||||
await renderWith(
|
||||
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
|
||||
const el = screen.getByTestId('ctx-accounting')
|
||||
expect(el).toHaveAttribute('data-status', 'truncation_suspected')
|
||||
expect(el).toHaveAttribute('role', 'alert')
|
||||
expect(el).toHaveTextContent('6,316')
|
||||
expect(el).toHaveTextContent('2,050')
|
||||
expect(el).toHaveTextContent(/turn is kept/)
|
||||
})
|
||||
|
||||
it('reports an unknown count plainly, without claiming the prompt fit', async () => {
|
||||
vi.spyOn(api, 'getActionContext').mockResolvedValue({
|
||||
...REPORT, accounting: { status: 'unknown', server_prompt_tokens: null },
|
||||
})
|
||||
await renderWith(
|
||||
<ContextPanel advId="1" inspectActionId="9" refreshKey="x" onClearInspect={() => {}} />)
|
||||
const el = screen.getByTestId('ctx-accounting')
|
||||
expect(el).toHaveAttribute('data-status', 'unknown')
|
||||
expect(el).not.toHaveAttribute('role')
|
||||
expect(el).toHaveTextContent(/did not say/)
|
||||
expect(el.textContent).not.toMatch(/read the whole prompt/)
|
||||
})
|
||||
|
||||
it('shows a retrieved passage with its source, class, heading and score', async () => {
|
||||
await renderWith(<ContextPanel advId="1" refreshKey="x" />)
|
||||
const row = document.querySelector('[data-chunk-id="11"]')
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
/* v1.1 WP-D: an export that says whether this version could import it back.
|
||||
*
|
||||
* The file is delivered either way — a campaign too large to re-import is not a
|
||||
* damaged one, and refusing to write it would destroy the copy the reader was
|
||||
* making. What changes is what they are told, and both reader-facing Export
|
||||
* controls have to tell them: the one on the campaign card and the one in the
|
||||
* campaign's own settings.
|
||||
*
|
||||
* The server decides. These assert that the page shows what it was given and
|
||||
* keeps delivering the file, not that it re-derives the size policy.
|
||||
*/
|
||||
|
||||
import { screen, waitFor } from '@testing-library/react'
|
||||
import userEvent from '@testing-library/user-event'
|
||||
import { beforeEach, describe, expect, it, vi } from 'vitest'
|
||||
import { api } from '../api'
|
||||
import * as components from '../components'
|
||||
import Campaigns from './Campaigns'
|
||||
import { CampaignSettingsPanel } from './Play/panels/CampaignSettingsPanel'
|
||||
import { mockModelStatus, renderWith } from '../test/helpers'
|
||||
|
||||
const BUNDLE = { format: 'ai-dnd-adventure-v3', title: 'Long Campaign', actions: [] }
|
||||
|
||||
const WARNING =
|
||||
"This export is larger than this version's 20 MB import limit (21,230,000 bytes). "
|
||||
+ 'The file was exported successfully, but this version cannot import it.'
|
||||
|
||||
const CAMPAIGN = {
|
||||
id: 4, title: 'Long Campaign', action_count: 900,
|
||||
updated_at: '2026-09-15T10:00:00', snippet: 'Rain over the harbour.',
|
||||
}
|
||||
|
||||
const ADVENTURE = {
|
||||
id: 4, title: 'Long Campaign', ai_instructions: '', narration_length: 'brief',
|
||||
canon_rules: [], persona_name: 'Aldric', persona_desc: '',
|
||||
}
|
||||
|
||||
// restoreAllMocks does not undo stubGlobal, and the fetch stub below would
|
||||
// otherwise outlive its own describe block.
|
||||
beforeEach(() => { vi.restoreAllMocks(); vi.unstubAllGlobals() })
|
||||
|
||||
function exportReturns({ warning = null } = {}) {
|
||||
return vi.spyOn(api, 'exportAdventure').mockResolvedValue({
|
||||
bundle: BUNDLE,
|
||||
exportBytes: warning ? 21_230_000 : 12_000,
|
||||
importLimitBytes: 20 * 1024 * 1024,
|
||||
importable: !warning,
|
||||
warning,
|
||||
})
|
||||
}
|
||||
|
||||
async function library() {
|
||||
mockModelStatus(api)
|
||||
vi.spyOn(api, 'listAdventures').mockResolvedValue([CAMPAIGN])
|
||||
await renderWith(<Campaigns />)
|
||||
await screen.findByText('Long Campaign')
|
||||
}
|
||||
|
||||
async function settingsPanel() {
|
||||
mockModelStatus(api)
|
||||
await renderWith(
|
||||
<CampaignSettingsPanel adventure={ADVENTURE} setAdventure={vi.fn()}
|
||||
onError={vi.fn()} moments={900} />,
|
||||
)
|
||||
}
|
||||
|
||||
describe('reading what the server said', () => {
|
||||
// The tests below mock api.exportAdventure, so nothing there exercises the
|
||||
// header names. These do: a typo in one of them would otherwise leave the
|
||||
// whole suite green and the reader silently uninformed.
|
||||
function serverSends(headers) {
|
||||
const body = JSON.stringify(BUNDLE)
|
||||
vi.stubGlobal('fetch', vi.fn().mockResolvedValue(new Response(body, {
|
||||
status: 200,
|
||||
headers: { 'Content-Type': 'application/json', ...headers },
|
||||
})))
|
||||
}
|
||||
|
||||
it('reports a warning the server sent, with the sizes it named', async () => {
|
||||
serverSends({
|
||||
'X-Export-Bytes': '21230000',
|
||||
'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
|
||||
'X-Importable-By-This-Version': 'false',
|
||||
'X-Export-Warning': WARNING,
|
||||
})
|
||||
const result = await api.exportAdventure(4)
|
||||
expect(result.bundle).toEqual(BUNDLE)
|
||||
expect(result.warning).toBe(WARNING)
|
||||
expect(result.importable).toBe(false)
|
||||
expect(result.exportBytes).toBe(21_230_000)
|
||||
expect(result.importLimitBytes).toBe(20 * 1024 * 1024)
|
||||
})
|
||||
|
||||
it('claims nothing when an older server sends no headers', async () => {
|
||||
serverSends({})
|
||||
const result = await api.exportAdventure(4)
|
||||
expect(result.bundle).toEqual(BUNDLE)
|
||||
expect(result.warning).toBeNull()
|
||||
expect(result.importable).toBe(true)
|
||||
expect(result.exportBytes).toBeNull()
|
||||
expect(result.importLimitBytes).toBeNull()
|
||||
})
|
||||
|
||||
it('treats an ordinary export as importable', async () => {
|
||||
serverSends({
|
||||
'X-Export-Bytes': '12000',
|
||||
'X-Import-Limit-Bytes': String(20 * 1024 * 1024),
|
||||
'X-Importable-By-This-Version': 'true',
|
||||
})
|
||||
const result = await api.exportAdventure(4)
|
||||
expect(result.importable).toBe(true)
|
||||
expect(result.warning).toBeNull()
|
||||
expect(result.exportBytes).toBe(12_000)
|
||||
})
|
||||
})
|
||||
|
||||
describe('the campaign library export', () => {
|
||||
it('delivers the file and says nothing more when it can be imported back', async () => {
|
||||
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
|
||||
exportReturns()
|
||||
await library()
|
||||
await userEvent.click(screen.getByRole('button', { name: 'Export' }))
|
||||
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
|
||||
expect(download.mock.calls[0][0]).toEqual(BUNDLE)
|
||||
expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
|
||||
expect(screen.queryByText(/cannot import/)).toBeNull()
|
||||
})
|
||||
|
||||
it('still delivers the file when it is too large, and says so', async () => {
|
||||
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
|
||||
exportReturns({ warning: WARNING })
|
||||
await library()
|
||||
await userEvent.click(screen.getByRole('button', { name: 'Export' }))
|
||||
// The file is written first: the warning is about importing it back, not
|
||||
// about the export having failed.
|
||||
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
|
||||
expect(download.mock.calls[0][0]).toEqual(BUNDLE)
|
||||
const notice = await screen.findByText(/cannot import it/)
|
||||
expect(notice).toBeInTheDocument()
|
||||
expect(notice.textContent).toContain('20 MB')
|
||||
expect(notice.textContent).toContain('exported successfully')
|
||||
expect(screen.queryByText('Campaign exported.')).toBeNull()
|
||||
})
|
||||
})
|
||||
|
||||
describe('the campaign settings export', () => {
|
||||
it('delivers the file and confirms it when it can be imported back', async () => {
|
||||
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
|
||||
exportReturns()
|
||||
await settingsPanel()
|
||||
await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
|
||||
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
|
||||
expect(await screen.findByText('Campaign exported.')).toBeInTheDocument()
|
||||
expect(screen.queryByText(/cannot import/)).toBeNull()
|
||||
})
|
||||
|
||||
it('still delivers the file when it is too large, and says so', async () => {
|
||||
const download = vi.spyOn(components, 'downloadJSON').mockImplementation(() => {})
|
||||
exportReturns({ warning: WARNING })
|
||||
await settingsPanel()
|
||||
await userEvent.click(screen.getByRole('button', { name: 'Export campaign' }))
|
||||
await waitFor(() => expect(download).toHaveBeenCalledTimes(1))
|
||||
const notice = await screen.findByText(/cannot import it/)
|
||||
expect(notice.textContent).toContain('20 MB')
|
||||
expect(screen.queryByText('Campaign exported.')).toBeNull()
|
||||
})
|
||||
})
|
||||
@@ -32,6 +32,16 @@
|
||||
font-variant-numeric: tabular-nums;
|
||||
}
|
||||
.ctx-warn { margin: 8px 0 0; color: var(--danger); font-size: 0.76rem; }
|
||||
/* v1.1 A1: what the server said it read. The two failure states render as a
|
||||
`.notice.error` alert instead; this is the quiet form for `fits` and
|
||||
`unknown`. */
|
||||
.ctx-accounting {
|
||||
padding: 9px 13px;
|
||||
border: 1px solid var(--border);
|
||||
border-radius: 7px;
|
||||
color: var(--text-dim);
|
||||
}
|
||||
.ctx-accounting p { margin: 4px 0 0; }
|
||||
|
||||
.token-bar {
|
||||
display: flex;
|
||||
@@ -49,7 +59,22 @@
|
||||
.slice-4 { background: #6f9e8c; }
|
||||
.slice-5 { background: #c48a6a; }
|
||||
.slice-6 { background: #7c86b8; }
|
||||
.slice-7 { background: var(--border-bright); }
|
||||
/* v1.1 WP-E: pinned to the literal value this slice already rendered, instead
|
||||
of borrowing --border-bright. A chart fill and a control edge have different
|
||||
jobs: WP-E raised --border-bright to clear WCAG 1.4.11 (3:1) for control
|
||||
boundaries, and that dragged this slice to #7a7aaa, an OKLab dE of 0.035
|
||||
from .slice-6 (#7c86b8) — two neighbouring slices the same colour. The
|
||||
other slices sit 0.100-0.119 from their nearest neighbour; at #3d3d55 this
|
||||
one sits 0.251, the most separated in the set.
|
||||
|
||||
1.4.11's 3:1 does not govern this: it is a proportional fill in a labelled
|
||||
breakdown, not the boundary of a control, and what it needs is to be
|
||||
distinguishable from the seven slices beside it. Reassigning it to a freer
|
||||
hue was considered and rejected — inside the palette's own chroma and
|
||||
lightness bands the only hues that beat 0.100 are pinks near 14 degrees,
|
||||
which is --danger's territory and would paint an ordinary prompt section in
|
||||
the colour this application reserves for failure. */
|
||||
.slice-7 { background: #3d3d55; }
|
||||
|
||||
.ctx-block {
|
||||
border: 1px solid var(--border);
|
||||
|
||||
@@ -3,8 +3,15 @@
|
||||
--bg-panel: #131320;
|
||||
--bg-panel-glass: rgba(19, 19, 32, 0.82);
|
||||
--bg-input: #1a1a2a;
|
||||
--border: #2b2b3d;
|
||||
--border-bright: #3d3d55;
|
||||
/* v1.1 WP-E: control boundaries carry WCAG 1.4.11 (3:1 non-text contrast) on
|
||||
their own, rather than leaning on the control's text label. The floor is
|
||||
measured against --bg-input (#1a1a2a), not --bg-panel: inputs and buttons
|
||||
are drawn on --bg-input (styles/forms.css), and it is the lightest of the
|
||||
three backgrounds a border sits on, so it is the worst case.
|
||||
--border 3.21:1 and --border-bright 4.24:1 there; higher on the others
|
||||
(backend/tools/contrast_audit.py, which now fails below 3:1). */
|
||||
--border: #676792;
|
||||
--border-bright: #7a7aaa;
|
||||
--text: #e2ddd0;
|
||||
--text-dim: #918c7d;
|
||||
--accent: #d4a94e;
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Adventure Storyteller — Production Build Milestones
|
||||
|
||||
**Status:** In implementation. M1-M8 complete and accepted (M1 and M2: 2026-09-02; M3 and M4: 2026-09-03; M5: 2026-09-04; M6, M7 and M8: 2026-09-06, each of the last four after an independent review and a corrective pass). **M9 — Export, Backup, Recovery, and Migration Hardening — is next, and has not been started.**
|
||||
**Status:** **Complete. M1-M11 are closed, and v1.0.0 was released on 2026-09-14** (signed tag `v1.0.0` on signed commit `432f041`, which `main` also points at). This document is the v1 milestone history and is not extended. Post-v1 work is planned as v1.1 work packages in `V1.1-PLAN.md`, not as further milestones.
|
||||
**Base:** AI-DnD `d72f7c1bda0f34fccd84afb7a25c34eb01c901de`
|
||||
|
||||
## 1. Purpose
|
||||
@@ -1466,6 +1466,10 @@ applicable (report §T). The acceptance takes effect with the owner's signed
|
||||
closeout commit. **No release tag exists**: tagging `v1.0.0` is a separate
|
||||
decision, and the tag must point at that signed commit.
|
||||
|
||||
*Post-release note (2026-09-14):* both events have since happened. The closeout
|
||||
commit was signed as `432f041`, `main` was fast-forwarded to it, and the signed
|
||||
tag `v1.0.0` points at it. The paragraph above is kept as it stood at closeout.
|
||||
|
||||
**M1-M11 are all complete. There is no M12.** Post-v1 work is backlog, listed
|
||||
below under *Post-v1 backlog*, and none of it is an unfinished v1 milestone.
|
||||
|
||||
@@ -1550,6 +1554,10 @@ Work recorded for after v1. None of it is a v1 requirement or an unfinished v1
|
||||
milestone, and none of it has a brief. Each item needs one before work begins.
|
||||
Sources are the M11 report's §P and §S.6.
|
||||
|
||||
**Triaged on 2026-09-14 in `V1.1-PLAN.md`**, which orders it into v1.1 work
|
||||
packages, a v1.2 list and future work. The list below is kept as recorded at the
|
||||
M11 closeout; `V1.1-PLAN.md` is where its disposition now lives.
|
||||
|
||||
- **Context-window safety margin.** The largest prompts leave 23-42 real tokens,
|
||||
and Ollama cuts an over-window prompt with no error. Consider a deliberate
|
||||
reserve, or counting with the narrator's own tokenizer.
|
||||
|
||||
@@ -451,6 +451,50 @@ The memory system should favor:
|
||||
- uniqueness,
|
||||
- continuity relevance.
|
||||
|
||||
### As implemented (v1.1 WP-B.2): what the summariser is shown
|
||||
|
||||
A memory is written from one block of `MEMORY_INTERVAL` (6) story actions. The
|
||||
summariser is given the cast brief, then the block, inside a budget of
|
||||
`MEMORY_EXCERPT_TOKENS` (2,000).
|
||||
|
||||
- **A block that fits** is sent whole, exactly as v1.0.0 sent it.
|
||||
- **A longer block** was cut to its last 2,000 tokens in v1.0.0, so a fact near
|
||||
its start never reached the summariser (WP-B.1). It is now sent as its
|
||||
opening and its end, in order, with a visible marker between them
|
||||
(`EXCERPT_OMISSION_MARKER`, "[… the middle of this stretch of story is left
|
||||
out here …]"). The marker and its blank lines are paid for first, and the rest
|
||||
is halved, the odd token going to the end: 992 + 993 + 15 = 2,000 tokens. The
|
||||
rejoined text is measured, and the opening gives up tokens until the whole is
|
||||
within budget.
|
||||
- **The marker is never stored.** `summarize_block` removes it from anything the
|
||||
model repeats back.
|
||||
- **Limit.** A fact in the middle of a very long block is still left out. The
|
||||
input stays bounded; it is not a summary of everything.
|
||||
|
||||
Existing memories are not rewritten. `tools/rewrite_memories.py`, which is
|
||||
opt-in, uses the same function.
|
||||
|
||||
### As implemented (v1.1): what a memory can be relied on to keep
|
||||
|
||||
The memory prompt is v1.0.0's, unchanged. v1.1 changed what the summariser is
|
||||
shown (above), not what it is told.
|
||||
|
||||
**Known limitation, accepted for v1.1.** The application now delivers the whole
|
||||
relevant block to the summariser, and keeps, ranks and injects the memory it
|
||||
writes (§18, §20, §21). But the reference summariser, `qwen2.5:3b-instruct-16k`,
|
||||
can still be given a block that states a distinctive fact and write a memory
|
||||
that:
|
||||
- omits the fact, or the specific objects in it;
|
||||
- attributes it to the wrong character;
|
||||
- prefers the generic narration that follows it.
|
||||
|
||||
So **independent recovery from memory is proven for the application's
|
||||
mechanisms, and is not guaranteed with the reference model.** A bounded prompt
|
||||
change aimed at this was tried and rejected
|
||||
(`reports/v1.1/V1.1-WP-B2-REPORT.md` §T). A stronger dedicated summariser,
|
||||
structured fact extraction, or separate factual and narrative memory are
|
||||
future options, and none is implemented.
|
||||
|
||||
## 16. Memory Retrieval
|
||||
|
||||
Retrieval should be local.
|
||||
@@ -513,6 +557,27 @@ abbey crypt
|
||||
prior discoveries
|
||||
```
|
||||
|
||||
### As implemented (v1.1 WP-B.2)
|
||||
|
||||
v1.0.0 embedded the newest four actions cut to 600 tokens, so the player's
|
||||
one-line question arrived after three turns of narration and barely moved the
|
||||
embedding (WP-B.1: a direct question's cosine fell from 0.708 alone to 0.241 in
|
||||
that query). The query is now two short texts, embedded in one call:
|
||||
|
||||
| Component | What it is | Bound |
|
||||
| --- | --- | --- |
|
||||
| **input** | the player's own action this turn (`do`, `say` or `story`) | last 200 tokens |
|
||||
| **context** | the scene from the authoritative state (its summary, the location's name, the names of who is present), then the end of the newest narration | 60 + 120 tokens |
|
||||
|
||||
- A continue turn, and the Insights dry run, have no input; the context alone is
|
||||
searched.
|
||||
- A retry searches with the input being retried; the discarded attempt is not in
|
||||
the context.
|
||||
- The full entity list, threads and older narration are deliberately left out,
|
||||
so a long scene or a large cast cannot outweigh the question by length.
|
||||
- What was searched for is recorded per turn in the context snapshot
|
||||
(`memories.query`: input, context, input terms, weights).
|
||||
|
||||
## 19. Memory Retrieval Filtering
|
||||
|
||||
Before ranking memories, filter by:
|
||||
@@ -553,6 +618,41 @@ future work rather than something M6 delivered.
|
||||
What M6 does implement, because similarity alone proved insufficient, is
|
||||
redundancy suppression before the final selection: see §22.
|
||||
|
||||
### As implemented (v1.1 WP-B.2)
|
||||
|
||||
One transparent lexical term is added to similarity. For each eligible memory:
|
||||
|
||||
```text
|
||||
semantic_score = 0.6 * cos(input, memory) + 0.4 * cos(context, memory)
|
||||
(either cosine alone when the other text is empty)
|
||||
lexical_score = sum of w(t) over the input's terms the memory holds
|
||||
/ sum of w(t) over all the input's terms in [0, 1]
|
||||
w(t) = ln((N + 1) / (df(t) + 1)) N eligible memories, df holding t
|
||||
final_score = semantic_score + 0.15 * lexical_score
|
||||
```
|
||||
|
||||
- **Terms** are the knowledge path's tokenizer and stop list (`knowledge.fts`),
|
||||
with possessives dropped and a plural `s` folded. No stemmer, no dependency.
|
||||
- **Rarity** is computed per turn over the eligible candidates only. There is no
|
||||
index and no stored field. A word every candidate holds (a protagonist's
|
||||
name) weighs 0; a word the question shares with one memory weighs most.
|
||||
- **Only the player's input** is matched lexically, never the context.
|
||||
- **The weight** was chosen by sweep over 0, 0.05, 0.1, 0.15, 0.2, 0.3 and 0.5:
|
||||
0.15 is the smallest at which the lexical term alone lifts the planting-era
|
||||
memory into `memory_top_k` against the v1.0.0 query, while no rare-word
|
||||
negative control lets an unrelated memory pass a semantically relevant one.
|
||||
A memory can gain at most 0.15 from wording, so it cannot pass one more than
|
||||
0.15 ahead of it in meaning.
|
||||
- **Pins** are unchanged: always used, counted toward `memory_top_k`.
|
||||
- **Ties** on the final score are broken by memory id.
|
||||
- **Provenance.** Each used memory records `semantic_score`, `lexical_score` and
|
||||
`final_score`; `similarity` keeps its v1.0.0 meaning, the semantic score, so
|
||||
the inspector's "closeness" is unchanged.
|
||||
|
||||
Importance, recency, entity overlap and story-thread overlap remain
|
||||
unimplemented. Redundancy suppression (§22) is unchanged and runs over the final
|
||||
order.
|
||||
|
||||
## 21. Memory Budget
|
||||
|
||||
Retrieved memories should have a bounded token budget.
|
||||
@@ -564,6 +664,44 @@ Recommended behavior:
|
||||
- include only the highest-value items that fit,
|
||||
- preserve source IDs for inspection.
|
||||
|
||||
### As implemented (v1.1 WP-B.2): selection and the bank's capacity
|
||||
|
||||
**Selection** is unchanged in shape: every eligible, embedded memory on the
|
||||
active lineage is scored (§20), pinned memories are taken first, the rest fill
|
||||
`memory_top_k` (default 5) best first, skipping repeats (§22). The Memories
|
||||
section is priced into the protected context like any other live section, so
|
||||
its budget is unchanged (F03). There is still no relevance floor: a full
|
||||
`memory_top_k` is used whenever the bank holds that many.
|
||||
|
||||
**Capacity** (`memory_bank_capacity`, default 80) is unchanged. Eviction still
|
||||
runs after each post-turn pass, over the whole adventure rather than one
|
||||
lineage, and marks rows `forgotten` rather than deleting them. What changed is
|
||||
the order (`memorybank.eviction_order`), because least-recently-used alone
|
||||
discarded the only memory of an early stretch first (WP-B.1):
|
||||
|
||||
- **Coverage signal.** Memories with a source range say which stretch they
|
||||
describe. Each is judged by the hole its removal would leave between the end
|
||||
of the memory before it and the start of the memory after it. The smallest
|
||||
hole goes first, so the bank thins where it is densest. A memory whose start
|
||||
another memory shares leaves no hole.
|
||||
- **Boundaries.** The earliest and the latest memory by position are not
|
||||
coverage candidates: they are the only records of the opening and of the most
|
||||
recent stretch. This also keeps a memory written this turn from being evicted
|
||||
by the pass that wrote it (the frozen bank).
|
||||
- **Recency signal.** Among equal holes, the least recently used goes first
|
||||
(`coalesce(last_used_at, created_at)`), then the less used, then the lower id.
|
||||
- **Pinned rows** are never taken, and count as coverage.
|
||||
- **Fallback.** Memories with no range (typed by the player, or migrated) and
|
||||
boundaries are taken least recently used first, as in v1.0.0, once no coverage
|
||||
candidate remains. The bank stays bounded either way; only an all-pinned bank
|
||||
may exceed capacity.
|
||||
|
||||
Measured on banks where nothing is ever retrieved, the kept bank starts at the
|
||||
opening and its largest uncovered stretch stays within about 1.3 times the
|
||||
average spacing (story length / capacity). The v1.0.0 order kept only the newest
|
||||
stretch. The rule reads no text and no vectors, and lineage eligibility is
|
||||
unaffected: it decides only `forgotten`.
|
||||
|
||||
## 22. Duplicate Suppression
|
||||
|
||||
Do not include the same fact repeatedly through:
|
||||
|
||||
@@ -69,6 +69,18 @@ out of tokens partway through it. All of that is removed before the prose is
|
||||
stored, because stored prose is replayed as history. `TECHNICAL-DESIGN.md` §15.4
|
||||
has the rules.
|
||||
|
||||
**Implementation note (v1.1 WP-A2).** A small model also copies the protocol's
|
||||
*instructions*: the vocabulary written as calls, the length hint, and the scene
|
||||
line. The fix is on both sides:
|
||||
|
||||
- **Prompt:** the vocabulary is shown in the wire format, and the fixed example
|
||||
uses genre-neutral placeholders.
|
||||
- **Extractor:** it recognises those echoes only by strings and names the
|
||||
application owns.
|
||||
|
||||
No event type, field, validation rule or proposal record changed.
|
||||
`TECHNICAL-DESIGN.md` §15.4 lists the four rules.
|
||||
|
||||
### Where it lives
|
||||
|
||||
- `adventures.narrative_state` — the current authoritative document. This is
|
||||
|
||||
+50
-27
@@ -2,11 +2,22 @@
|
||||
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
**milestones M1 through M11 complete**. **M11 was accepted at its closeout
|
||||
(2026-09-14), and the v1 release gate passed** on the release-candidate tree.
|
||||
There is no further planned milestone. The signed closeout commit and any
|
||||
`v1.0.0` tag are separate events and the repository owner's to perform. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
**Current state:** **v1.0.0 remains the released version. Every v1.1 work
|
||||
package is implemented, accepted and signed** — WP-A1/A2 (`d63804f`), WP-B.1
|
||||
(`beb17ad`), WP-B.2 (`0c1ba83`), WP-C (`59b5ebc`), WP-D and WP-E (`87a4032`).
|
||||
**Integrated release validation passed on candidate `87a4032`**
|
||||
(`reports/v1.1/V1.1-RELEASE-REPORT.md`): the v1 contract holds (81 PASS, H09 NOT
|
||||
APPLICABLE), every suite and build passes, and the browser, offline, long-run,
|
||||
identity, recovery, upgrade and smoke gates are clean. WP-B ships with a
|
||||
documented reference-model memory limitation. **No release commit, no `main`
|
||||
update and no `v1.1.0` tag exist yet** — those are the owner's separate events.
|
||||
Phase 0 complete; AI-DnD forked as the production base; **milestones M1
|
||||
through M11 complete and closed**. M11 was accepted at its closeout
|
||||
(2026-09-14), the v1 release gate passed on the release-candidate tree, and the
|
||||
owner then signed the closeout commit (`432f041`), fast-forwarded `main` to it,
|
||||
and created and pushed the signed tag `v1.0.0` pointing at it. v1.1 development
|
||||
is on the `v1.1-development` branch, from that commit, and is planned in
|
||||
`V1.1-PLAN.md` as work packages, not milestones. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
|
||||
@@ -46,13 +57,18 @@ The closeout repeated the browser, offline and identity runs on the exact
|
||||
release-candidate tree (`3652dc6`), and recorded the acceptance (report §S and
|
||||
§T). **The v1 release gate passed.**
|
||||
|
||||
**Not yet happened**, and each a separate event:
|
||||
- the owner signing the closeout commit;
|
||||
- merging to `main`, which still points at M10;
|
||||
- any `v1.0.0` tag.
|
||||
**The release, 2026-09-14.** The three events the closeout left to the owner
|
||||
have all happened:
|
||||
- the closeout commit is signed, as `432f041`;
|
||||
- `main` was fast-forwarded to `432f041`;
|
||||
- the signed tag `v1.0.0` points at `432f041`, and is pushed.
|
||||
|
||||
**There is no further planned milestone and no M12.** Post-v1 work is backlog
|
||||
(`BUILD-MILESTONES.md`, *Post-v1 backlog*).
|
||||
The M11 report's §T still lists the last two as not done. That table records the
|
||||
state at closeout and is deliberately left as written.
|
||||
|
||||
**There is no M12.** Post-v1 work is organised as **v1.1 work packages** in
|
||||
`V1.1-PLAN.md`, which triages the *Post-v1 backlog* recorded in
|
||||
`BUILD-MILESTONES.md`.
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
@@ -122,7 +138,8 @@ Two standing qualifications:
|
||||
| `SECURITY-THREAT-MODEL.md` | The trust boundary, and the inference endpoint policy as implemented. |
|
||||
| `MEDIA-EXTENSION-CONTRACT.md` | The contract future image/video/audio/TTS/STT work must fit. |
|
||||
| `BROWSER-UX-SPEC.md` | The browser surface, and what is deliberately not in it. |
|
||||
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. |
|
||||
| `BUILD-MILESTONES.md` | M1-M11, what each delivers, what is done, and the notes each milestone leaves its successors. Closed history; not extended. |
|
||||
| `V1.1-PLAN.md` | **Post-v1 work.** The triaged backlog, the ordered v1.1 work packages with acceptance criteria, the v1.1 release criteria, and what is deferred. |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | The pass/fail contract v1 is measured against. |
|
||||
| `TEST-CAMPAIGN-FIXTURE.md` | The standard campaign the acceptance tests are run on. |
|
||||
| `VERSION.md` | Package revision history: what each closeout changed. |
|
||||
@@ -133,7 +150,8 @@ Two standing qualifications:
|
||||
1. This file.
|
||||
2. `SPECIFICATION.md`
|
||||
3. `TECHNICAL-DESIGN.md`
|
||||
4. `BUILD-MILESTONES.md` — find the milestone you are being asked to do.
|
||||
4. `V1.1-PLAN.md` — find the work package you are being asked to do.
|
||||
`BUILD-MILESTONES.md` is the v1 history behind it.
|
||||
5. `STORY-BRANCH-SEMANTICS.md`
|
||||
6. `DATA-MODEL.md`
|
||||
7. `CONTEXT-AND-MEMORY.md`
|
||||
@@ -185,7 +203,9 @@ that is the one the next milestone's planning has to consult:
|
||||
It stays in `reports/` after acceptance. The rotation moves a report to the
|
||||
archive when the next milestone's report is written, and after M11 there is no
|
||||
next milestone. Nothing here calls for moving it, so no new convention was
|
||||
invented to do so.
|
||||
invented to do so. `V1.1-PLAN.md` asks each v1.1 work package for a report of
|
||||
its own; when the first is written, it joins this one in `reports/` and the
|
||||
rotation is decided then.
|
||||
|
||||
M9's and M10's reports both moved to `archive/milestone-reports/` when this one
|
||||
was written. M9's had been kept here past its turn because M9 was unaccepted;
|
||||
@@ -337,27 +357,30 @@ Milestone M11 COMPLETE / ACCEPTED (2026-09-14)
|
||||
|
|
||||
v
|
||||
v1 release gate PASSED (2026-09-14)
|
||||
signed release commit and v1.0.0 tag:
|
||||
the repository owner's, not yet done
|
||||
|
|
||||
v
|
||||
Post-v1 backlog only; no milestone planned
|
||||
v1.0.0 RELEASED (2026-09-14)
|
||||
signed tag on signed commit 432f041;
|
||||
main points at the same commit
|
||||
|
|
||||
v
|
||||
v1.1 PLANNING (from 2026-09-14)
|
||||
branch v1.1-development; V1.1-PLAN.md
|
||||
work packages, not milestones
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
**One work package at a time. Do not begin a work package before its brief
|
||||
exists.**
|
||||
|
||||
**Every planned milestone is complete, and M11 is accepted.** The v1 release gate
|
||||
passed on 2026-09-14 (M11 report §T).
|
||||
**Every v1 milestone is complete, M11 is accepted, and v1.0.0 is released**
|
||||
(2026-09-14, tag `v1.0.0` on `432f041`).
|
||||
|
||||
The next actions are the owner's:
|
||||
1. sign the closeout commit;
|
||||
2. decide on the `v1.0.0` tag, pointed at that signed commit.
|
||||
|
||||
Neither is a milestone. **Do not begin post-v1 work as though it were a v1
|
||||
milestone.** Anything after v1 starts from the *Post-v1 backlog* in
|
||||
`BUILD-MILESTONES.md`, with a brief of its own.
|
||||
**Do not begin post-v1 work as though it were a v1 milestone, and do not create
|
||||
M12.** v1.1 work is the ordered work packages in `V1.1-PLAN.md`. Each starts
|
||||
only from a coding brief of its own, and stops at its own boundary for review.
|
||||
The plan names the brief to write first.
|
||||
|
||||
All three questions the M8 debt raised against M9 are settled and recorded:
|
||||
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
|
||||
|
||||
@@ -1230,6 +1230,75 @@ whether there is a number to cap to at all, and that is what the builder and the
|
||||
declaration everywhere it appears, and the connection test says plainly that
|
||||
nothing has checked it against the server.
|
||||
|
||||
**As implemented (v1.1 WP-A1): a safety reserve, and the server's own count.**
|
||||
A ceiling in the application's tokens is not a ceiling in the narrator's. The
|
||||
builder counts with `cl100k_base`, and the v1 evidence left the largest prompts
|
||||
23-42 real tokens from the edge of a 16,384 window. Past the edge Ollama does not
|
||||
refuse. Measured on Ollama 0.33 at a 4,096 window, a 6,316-token prompt returned
|
||||
200 with `prompt_tokens` 2,050.
|
||||
|
||||
- **The reserve.** `contextwindow.safety_reserve(budget)` is
|
||||
`max(256, ceil(5% of the effective budget))`, computed with integer rounding
|
||||
up: 256 at 4,096, 410 at 8,192, 820 at 16,384. It is taken before any history
|
||||
is chosen. The effective budget is the verified or declared window when there
|
||||
is one, and the configured budget otherwise. It is fixed and documented, is not
|
||||
a setting, and is not calibrated per model.
|
||||
- **What replaced the 64-token margin.** M6's `OUTPUT_SAFETY_MARGIN` absorbed two
|
||||
unrelated things.
|
||||
- The application's own text added after pricing: separators between
|
||||
sections, and the chat hint the provider appends to every request. This is
|
||||
now priced exactly as `transport`.
|
||||
- Tokenizer drift. This is now the reserve.
|
||||
|
||||
The reply allocation is exactly `max_output_tokens`. Protected context is
|
||||
`sections + transport + reply + reserve`, and `ContextOverflow` is raised
|
||||
before the model call when that does not fit.
|
||||
- **The server's count.** Streaming requests set `stream_options.include_usage`.
|
||||
Without it Ollama sends no usage, and none of the 514 AI turns in the v1
|
||||
evidence has one. After the reply, `contextwindow.classify_usage` compares the
|
||||
server's `prompt_tokens` with `tokens.estimate`: the assembled text plus what
|
||||
the provider adds.
|
||||
- **Accounting states.** The turn's snapshot records `accounting`, whose status
|
||||
is one of the following, checked in this order:
|
||||
|
||||
| Status | Meaning |
|
||||
| --- | --- |
|
||||
| `unknown` | No positive integer count was reported. It is never read as `fits`. |
|
||||
| `truncation_suspected` | The server read fewer tokens than the estimate by more than the reserve. |
|
||||
| `exceeded` | The server's count plus the reply allocation is over the budget. |
|
||||
| `fits` | Otherwise. |
|
||||
|
||||
The record also carries the server's count, the difference, the reserve and
|
||||
the observed margin (`budget - reply - server count`).
|
||||
- **Surfacing.** The record is returned on the turn's `done` event, logged as a
|
||||
warning when it is `exceeded` or `truncation_suspected`, and shown in the
|
||||
context inspector, where those two statuses are an alert.
|
||||
- **The turn is kept.** A discrepancy found after the reply is recorded, never
|
||||
enforced. The narration has already streamed to the reader, and the accepted
|
||||
turn is not discarded.
|
||||
- **Accounting belongs to one attempt.** It sits in `attempts.ATTEMPT_KEYS`
|
||||
beside `usage`. When a retry or a take selection moves the shared prompt,
|
||||
each take keeps the accounting for its own call.
|
||||
- **A cold model is loaded before its turn is built (v1.1 A1 corrective).** A
|
||||
model that is not resident cannot report its window. The A1 evidence caught
|
||||
exactly that: 13,875 tokens were sent to a server that read 2,050.
|
||||
`contextwindow.ensure_window` works like this:
|
||||
- it probes;
|
||||
- if the window is unverified and the server answered, it makes one bounded
|
||||
`POST /api/generate` naming only the model, with no prompt. Ollama loads
|
||||
the model and generates nothing ("done_reason": "load");
|
||||
- it probes again, bypassing the cache;
|
||||
- the turn is built to whatever that second probe says.
|
||||
|
||||
A failed load, or a window still unverified afterwards, changes nothing: the
|
||||
configured budget stands and the accounting still catches a cut. The request
|
||||
goes to the configured endpoint only, under the same policy and TLS path. It
|
||||
writes nothing, and it is recorded as `window.preflight` in the turn's
|
||||
snapshot. The context dry run never loads a model.
|
||||
|
||||
No schema change: the accounting lives in the snapshot JSON, and a turn from
|
||||
v1.0.0 simply has none. The bundle format is unchanged, for the same reason.
|
||||
|
||||
### 15.3 The history window moves in blocks (post-M11)
|
||||
|
||||
§15.2 makes the window a ceiling. This is about what happens at that ceiling.
|
||||
@@ -1300,6 +1369,53 @@ headings. A lone heading followed by prose stays, and so do JSON a character
|
||||
typed and a fact restated inside a sentence. That last case is how a narrator can
|
||||
still carry authoritative state into its prose (M11 report §G.4 and §P).
|
||||
|
||||
**As implemented (v1.1 WP-A2): the source and the sink together.** The M11
|
||||
closeout's identity run stored four shapes the extractor left. All of them were
|
||||
application text. v1.1 changes both the prompt that taught them and the
|
||||
extractor that missed them. Every new removal is anchored to something the
|
||||
application owns, never to what prose looks like.
|
||||
|
||||
- **The prompt.**
|
||||
- `events.vocabulary_for_prompt` shows each event as the object the model
|
||||
must send (`{"type": "set_possession", "item": "<key>", "owner": "<key>"}`),
|
||||
not as `set_possession(item, owner)`. The call notation was never the wire
|
||||
format, and the narrator copied it.
|
||||
- `EMIT_RULE`'s example uses the placeholders `character-1`, `item-1` and
|
||||
`location-1`, not the fantasy fixture's `mara`, `silver-key`, `old-abbey` and
|
||||
`aldric`. The narrator had proposed `silver-key` in an office meeting.
|
||||
- The length hint's opening and closing words are named constants shared by
|
||||
the builder and the extractor.
|
||||
- **The extractor**, rules R1-R4 (`RULE_*` in `narrative/extract.py`):
|
||||
- **R1:** a whole line that begins with a call to an event in `events.SPECS`,
|
||||
optionally `>`-quoted. Not inside a fenced code block, not mid-sentence, and
|
||||
not for a call-shaped name the vocabulary lacks.
|
||||
- **R2:** a trailing bracket that opens `Hard limit:` and carries the hint's
|
||||
own wording ("append the state block", or "turn must not exceed *N* words").
|
||||
- **R3:** the renderer's scene line left as the reply's last line. It is
|
||||
removed when it ends in the renderer's `(at <location>)`, or when protocol
|
||||
was already cut from the same reply.
|
||||
- **R4:** a ```` ```json ```` or bare ```` ``` ```` opener left as the last
|
||||
line with nothing after it, counted as an opener rather than a closer.
|
||||
- **Proven on real narration.** Every stored real reply in the v1 evidence was
|
||||
replayed through the v1.0.0 and v1.1 extractors (`tools/v11_replay_extractor.py`).
|
||||
Every changed line is attributed to one of the four rules, and a person reviewed
|
||||
every change. The results are in the WP-A1/A2 report.
|
||||
- **Deliberately still left:** a fact restated inside a sentence; model-invented
|
||||
headings; and a bracket that starts `Hard limit:` but carries none of the
|
||||
application's wording.
|
||||
- **R5, the echoed instruction tail (v1.1 A2 corrective).** A v1.1 identity turn
|
||||
ended in a reworded continue hint, "[… Continue the story here, directly.
|
||||
Output only story text.]". Because nothing recognised it, nothing above it was
|
||||
ever trailing, and the reminder, a reworded length hint and a scene line all
|
||||
stayed. R5 makes two changes:
|
||||
- **The continue hint is recognised by its own sentence.** A trailing bracket
|
||||
containing "Output only story text" is an echoed instruction.
|
||||
`CONTINUE_HINT_PHRASE` is pinned by a test to `CHAT_CONTINUE_HINT`.
|
||||
- **A bracket opening with the length hint's own `Hard limit:` is removed only
|
||||
directly above an echoed instruction already cut from the same reply's end.**
|
||||
An in-world "[Hard limit: forty days]" stays when it is the last line, and
|
||||
when a state block follows it. Any other bracket above an echo stays.
|
||||
|
||||
## 16. Database Direction
|
||||
|
||||
SQLite remains the selected v1 authoritative store.
|
||||
|
||||
@@ -0,0 +1,940 @@
|
||||
# Adventure Storyteller — v1.1 Plan
|
||||
|
||||
**Status:** IN PROGRESS. Written 2026-09-14 on `v1.1-development`, from the
|
||||
signed v1.0.0 release commit `432f041`.
|
||||
|
||||
**WP-A1 and WP-A2** are committed and signed as `d63804f`
|
||||
(`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`). **WP-B.1**, the memory-retention
|
||||
diagnostic, is committed and signed as `beb17ad`
|
||||
(`reports/v1.1/V1.1-WP-B1-REPORT.md`). **WP-B.2**, the retention fix, is complete and
|
||||
staged for the owner's signed commit. It is **accepted with a documented
|
||||
real-model limitation** (owner decision, 2026-09-15):
|
||||
- B2.1 ranking, B2.2 eviction and B2.3 excerpt creation ship;
|
||||
- deterministic independent-memory recovery passes;
|
||||
- the one isolation-valid reference-model run failed at memory creation;
|
||||
- the B2.4 prompt experiment did not fix that and was reverted.
|
||||
|
||||
The evidence is in `reports/v1.1/V1.1-WP-B2-REPORT.md` §S and §T.
|
||||
**WP-B.2** is committed and signed as `0c1ba83`. **WP-C**, browser release
|
||||
coverage, is committed and signed as `59b5ebc`. Its final run passed 91 checks
|
||||
(the 38 existing and 53 new) with 0 failed and 0 skipped, over trusted-LAN HTTPS,
|
||||
including real export downloads (`reports/v1.1/V1.1-WP-C-REPORT.md`).
|
||||
**WP-D** (recovery honesty) and **WP-E**
|
||||
(control-boundary contrast) are committed and signed as `87a4032`, each with its
|
||||
own report: `V1.1-WP-D-REPORT.md` and `V1.1-WP-E-REPORT.md`.
|
||||
|
||||
**Integrated release validation has since run on candidate `87a4032` and
|
||||
passed** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): the 82 REQUIRED v1 tests hold
|
||||
(81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a
|
||||
`--no-cache` image whose SPA is file-for-file identical to the local build,
|
||||
offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window
|
||||
run passing M01-M04 with every turn `fits` and an A2 leak count of 0, identity
|
||||
0 signals / 0 protocol shapes, recovery 16/16, a real v1.0.0 upgrade comparing
|
||||
identical on all 15 fields with both bundle directions importing, and a release
|
||||
smoke of 15/15.
|
||||
|
||||
WP-B's reference-model memory limitation is carried as an accepted residual, as
|
||||
are the mid-reply instruction echo and the doubled full stop; K1 is classified as
|
||||
v1.2 backlog. **No `v1.1.0` tag exists, `main` is unchanged, and no release
|
||||
commit has been made** — those three remain the owner's separate events.
|
||||
|
||||
This document replaces nothing. `BUILD-MILESTONES.md` stays as the closed v1
|
||||
history (M1-M11), and its *Post-v1 backlog* is this plan's input. The v1
|
||||
specification and acceptance contract are unchanged: every package below
|
||||
improves how an existing requirement is met. No package adds a requirement.
|
||||
|
||||
---
|
||||
|
||||
## 1. Terminology
|
||||
|
||||
- **Work package (WP).** One independently reviewable unit of v1.1 work, with
|
||||
its own coding brief, its own report, and its own review. Work packages are
|
||||
lettered. **They are not milestones, and there is no M12.**
|
||||
- **Brief.** The coding prompt for one work package. It holds only the
|
||||
objective, the source documents, the scope and non-scope, the acceptance
|
||||
criteria, and the stop condition. No work package begins before its brief
|
||||
exists.
|
||||
- **Reference narrator.** `qwen2.5:3b-instruct`, with and without `num_ctx`
|
||||
16,384 baked in, and `nomic-embed-text` for embeddings. These are the models
|
||||
the v1 evidence used (M11 report §E.1). v1.1 real-model evidence uses the
|
||||
same models so that its numbers compare with v1's.
|
||||
- **Reference hosts.** The CPU reference host serves Ollama over HTTPS with a
|
||||
private CA, which is the A06 evidence path. The GPU inference host serves
|
||||
plain HTTP on the LAN and is used for long runs (M11 report §E.1).
|
||||
- **Planning package version** (`VERSION.md`, v4.x) and **product version**
|
||||
(v1.0.0, v1.1.0) are separate numbers. `SECURITY-THREAT-MODEL.md`'s own
|
||||
"Status: v1.1" line is that document's revision label from Phase 0B. It does
|
||||
not refer to this release.
|
||||
|
||||
## 2. Baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Release | **v1.0.0**, 2026-09-14 |
|
||||
| Release commit | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, signed by the owner |
|
||||
| Tag | `v1.0.0`, a signed annotated tag on that commit, pushed |
|
||||
| `main` | `432f041`, the same commit |
|
||||
| v1.1 branch | `v1.1-development`, created at `432f041` |
|
||||
| v1 contract | 82 REQUIRED FOR V1 tests: 81 PASS and H09 NOT APPLICABLE (M11 report §T) |
|
||||
| Schema | migrations up to 94 (`LATEST_VERSION` 94) |
|
||||
| Bundle format | `ai-dnd-adventure-v3` |
|
||||
| Provenance | the AI-DnD fork point `d72f7c1b` is an ancestor. `LICENSE` and `PROVENANCE.md` are unchanged since `1013c94` |
|
||||
|
||||
## 3. v1.1 goals
|
||||
|
||||
1. **No silent loss of what the narrator is given.** A prompt must not reach the
|
||||
server's window edge by an accident of tokenizer arithmetic. A turn the
|
||||
server truncated must be reported.
|
||||
2. **No protocol in the story.** The shapes of the application's own protocol
|
||||
that a narrator copies must not be stored as narration. The fixed
|
||||
instructions must not hand every campaign one genre's nouns to copy. Story
|
||||
prose must not be removed in the process.
|
||||
3. **Memory that remembers on its own.** An old fact must be recoverable from
|
||||
the memory bank itself, with provenance, when nothing else carries it. This
|
||||
must not weaken lineage safety.
|
||||
4. **Release evidence without API-only gaps.** Every reader-facing history,
|
||||
state, length, failure and export behaviour must be driven in a real browser.
|
||||
5. **Recovery that says when it cannot recover.** Backups get a full integrity
|
||||
check, and an export that import would refuse must say so.
|
||||
6. **Control boundaries that meet WCAG 1.4.11.**
|
||||
|
||||
And throughout: **every v1.0.0 campaign opens unchanged in v1.1.**
|
||||
|
||||
## 4. Non-goals
|
||||
|
||||
- New product features: media generation, TTS, STT, story search, whole-transcript
|
||||
copy, a discarded-history recovery screen, a restore button, tablet redesign.
|
||||
- Changes to the history, branch, head or Save Point architecture (ADR 005,
|
||||
ADR 012).
|
||||
- Changes to the authoritative state model, its event vocabulary or its
|
||||
validator (ADR 010, ADR 013). Changes to the prompt *wording* that describes
|
||||
the vocabulary are in scope, in WP-A2.
|
||||
- Changes to the knowledge authority classes or their retrieval.
|
||||
- A new bundle format version.
|
||||
- Any new runtime dependency, any network access beyond the configured
|
||||
inference endpoint, and any first-use download, including per-model
|
||||
tokenizers.
|
||||
- Supporting inference servers other than Ollama, beyond what
|
||||
`context_window_override` already allows.
|
||||
- Re-writing stored narration or memories of existing campaigns automatically.
|
||||
- Redesigning identity or state handling on the strength of post-M8 finding D,
|
||||
which has not been reproduced.
|
||||
|
||||
## 5. Rules for every work package
|
||||
|
||||
The v1 milestone rules (`BUILD-MILESTONES.md` §2) apply to every work package,
|
||||
and so do the following:
|
||||
|
||||
1. **Compatibility default.** Existing v1.0.0 campaign databases open
|
||||
unchanged. Any migration is forward-only and additive, and is tested by
|
||||
opening a real v1.0.0 database. Fresh-install and upgraded schemas must still
|
||||
compare identical (`test_m11_migration.py`). v1.0.0 bundles still import.
|
||||
2. **The bundle stays `ai-dnd-adventure-v3`** unless a brief makes the case under
|
||||
M9's semantic test and the owner agrees. Adding a key inside stored evidence
|
||||
is not a format change when an absent key unambiguously means "not
|
||||
recorded".
|
||||
3. **The v1 contract is the regression floor.** No REQUIRED FOR V1 test is
|
||||
retired, relaxed or reclassified.
|
||||
4. **Evidence discipline** (M11 report §O). Any product change made after a
|
||||
black-box run was taken invalidates that run for release purposes. Harness
|
||||
defects and product defects are reported separately. A check that cannot fail
|
||||
is not evidence.
|
||||
5. **Each package has a report** (`planning/reports/V1.1-WP-<id>-REPORT.md`).
|
||||
It records what was built, the acceptance results with measurements, what
|
||||
was not verified, and the as-implemented documentation changes. The report
|
||||
rotation for `reports/` is decided when the first one is written
|
||||
(`planning/README.md`).
|
||||
6. **Real-model runs longer than a few minutes on the GPU host** require the
|
||||
power, link and kernel logging in `DEVELOPMENT.md`, started before the run.
|
||||
7. **No real hostnames, addresses or people's names** in committed files,
|
||||
fixtures or reports. Evidence stays under `$HOME`, never `/tmp`.
|
||||
8. **Commits and tags are the owner's.** A package ends with its changes staged
|
||||
for a signed commit.
|
||||
|
||||
---
|
||||
|
||||
## 6. Backlog triage
|
||||
|
||||
Severity reflects user risk: **High** means story corruption or lost continuity
|
||||
that the reader is not told about. **Medium** means silent degradation or a
|
||||
real verification gap. **Low** means visible, rare or cosmetic.
|
||||
|
||||
### 6.1 The recorded post-v1 backlog
|
||||
|
||||
| # | Item | Source | Current evidence | Severity | Disposition |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| 1 | Context-window safety margin | M11 §P risk 3, §N; post-v1 backlog | The largest prompts left 23, 34 and 42 real tokens of headroom in the GPU runs. The application's `cl100k_base` count ran 16 tokens below the narrator's on every prompt measured. The only slack is `OUTPUT_SAFETY_MARGIN = 64` (`context/builder.py`), which is fixed and must also absorb the section separators. Ollama 0.34 cut an over-window synthetic prompt to 8,194 tokens with no error. Truncation drops the oldest tokens, which are the narrator's rules and the canon. The server's reported usage is already stored per turn (`snapshot["usage"]`, `routers/adventures/turns.py`) and read by nothing. | **High**: silent, invisible, and reachable by a model whose tokenizer diverges further | **v1.1, WP-A1** |
|
||||
| 2 | Narrator restating prompt and state text | §P risk 4, §G.4 | The state's fact line was restated as a sentence on 74 of 104 turns in the evidence run. Phrases such as "Scene set, continue your adventure." and lines opening "Memory:" are model-invented, and none is an application string (verified by search). The restatements sit inside story sentences. | **Medium**: narration quality, and it is the path by which M04's fact reached memory, which masks item 5 | **v1.1, WP-A2** measures it. Reducing in-sentence restatement is **v1.2**, because removing it means judging prose |
|
||||
| 3 | Protocol leakage: event-call syntax, stray headings, repeated length hints | §P risk 6, §S.6 | On 4 of 10 turns of the closeout identity run (4,096 window) and 0 of 104 in the 100-turn runs. Causes verified in code. `events.vocabulary_for_prompt` shows every event as `name(field, …)`, which is the notation copied. `_is_echoed_instruction` requires both "state block" and "events list", and the length hint's tail says only "state block". A lone trailing `Scene:` is not removed. Stored text is replayed as history (TECHNICAL-DESIGN §15.4). | **Medium-high**: silent and self-reinforcing, but rare at a full window | **v1.1, WP-A2** |
|
||||
| 4 | Genre-specific example in the fixed state rule | §P risk 16; `narrative/extract.py` `EMIT_RULE` | The example names `mara`, `silver-key`, `old-abbey` and `aldric`. In the office-meeting identity run the narrator proposed giving `silver-key` to Alice, and the validator refused it. No state was corrupted. | **Medium**: every non-fantasy campaign gets fantasy nouns to copy | **v1.1, WP-A2** |
|
||||
| 5 | Independent long-term-memory retention | §P risk 5, §G.4; F02 and M04 qualifications | No run showed a planting-era memory carrying the planted fact. Recovery ran from state to the narrator's restatement to memory. Code facts that bear on it, none yet proven to be the cause: a memory is at most 50 words for a 6-action block. The summariser sees the block's *last* 2,000 tokens (`truncate_to_last_tokens`), so an early fact in a long block can be cut. Eviction is least-recently-used above `memory_bank_capacity` (default 80), so a never-retrieved early memory is the first to go once a campaign passes about 480 actions; the 33-memory evidence run never reached that. Ranking is cosine similarity plus a pin (CONTEXT-AND-MEMORY §20). | **High** for long campaigns: lost continuity, seen only as the story forgetting | **v1.1, WP-B** |
|
||||
| 6 | Browser automation gaps: Retry, Save Point, state correction, narration length, failed generation, export download | §S.3, §T qualifications, §P risk 11 | Proved through the API, the component suite and the 100-turn campaign, but never driven in a browser. The export control is exercised only as far as the click, because a `blob:` download does not leave the headless snap Firefox. No defect is known. | **Medium**: a verification gap on the release path | **v1.1, WP-C** |
|
||||
| 7 | Practical export and bundle-size ceiling | §P risk 12, §K; M9 debt | The import body limit is 20 MB (`limits.MAX_IMPORT_BODY_BYTES`). M9's conservative ceiling is about 279 turns. The real 100-turn campaign was 2.74 MB, about 13 kB per action, or roughly 1,600 actions. A larger campaign exports and is then refused on import, and nothing says so at export. | **Medium** impact, low frequency | **v1.1, WP-D**: warn at export. Raising the limit or a streaming import is **v1.2** |
|
||||
| 8 | Identity-confusion diagnostic follow-up | §P risk 8, §S.5 | Not reproduced. The diagnostic cannot see a stale scene, has never run with memory on, and has never run at a 16,384 window. It found no product deficiency. | **Low**: no occurrence since the original, whose campaign is gone | **No v1.1 package.** The diagnostic is re-run as a WP-A2 regression and in the v1.1 release gate, with memory on. A stale-scene detector is **v1.2**. The next real occurrence is classified with `tools/m11_identity.py` |
|
||||
| 9 | WCAG 1.4.11 control-boundary contrast | §P risk 10, §L | Measured at 1.33:1 resting and 1.75:1 on hover (`--border` and `--border-bright` against `--bg-panel`). M11's argument that the label identifies the control is weakest for text inputs, where the boundary shows where to type. | **Low-medium** | **v1.1, WP-E** |
|
||||
| 10a | Backup integrity: `integrity_check` | §P risk 14; M9 | `backup.py` runs `PRAGMA quick_check`, which skips index-content verification. | **Low** | **v1.1, WP-D** |
|
||||
| 10b | Scheduled backups | §P risk 14; M9 | None exist; backups are manual. They need owner policy on interval, retention, location and disk use, and must not contend with a turn's write lock (§O.7). | **Medium** impact, but a new feature | **v1.2** |
|
||||
| 11 | Real media-provider adapters | §P risk 13; M10; K05 and K06 (FUTURE) | The seam has no consumer. Adding one brings dependencies, a stricter endpoint rule (`DEVELOPMENT.md`) and a new UI. | n/a: a feature, not a risk to existing stories | **Future / optional**, not v1.1 |
|
||||
|
||||
### 6.2 Other recorded v1 residual risks and debt
|
||||
|
||||
| # | Item | Source | Disposition |
|
||||
| --- | --- | --- | --- |
|
||||
| 12 | The state lags the narration at 4,096 with the 3B narrator (1 proposal in 10 applied) | §P risk 17, §S.5 | Model behaviour, not a defect. WP-A2's real-model runs record the proposal outcome counts. No package |
|
||||
| 13 | Every realistic observation is a 3B model's | §P risk 2 | No package. v1.1 keeps the reference narrator for comparability. A comparative model recommendation stays open (`planning/README.md`, *Still open*) |
|
||||
| 14 | A campaign that never chose a narration length keeps the pre-M11 hint | §P risk 15 | Deliberate. No change. WP-C drives the control |
|
||||
| 15 | The GPU host dropped off the PCIe bus after the evidence run | §P risk 7, §E.1 | Operational. Rule 6 in §5 applies to every long run |
|
||||
| 16 | Release timings are two hosts' | §P risk 1 | No action. No performance requirement exists or is invented |
|
||||
| 17 | Cross-layer duplication: one fact in state, memory, history and a passage at once | CONTEXT-AND-MEMORY §22; open since M6 and M7 | **v1.2.** It costs tokens (item 1) and is related to item 2. WP-B may touch ranking only where its diagnostic requires |
|
||||
| 18 | Memory-ranking factors beyond similarity and pin not implemented | CONTEXT-AND-MEMORY §20 | Inside WP-B's bounded fix menu, only if WP-B's diagnostic places the failure at ranking |
|
||||
| 19 | M8 carried debt: no whole-transcript copy or story search (§77, §78), no discarded-history recovery screen (§63), tablet untuned, RPG world state read-only | `BUILD-MILESTONES.md` M8 and M9 | **Future backlog.** Features, not reliability |
|
||||
| 20 | M10 debt: `ambience` is an empty shape; a deleted visual profile is unrecoverable in the app; profiles are API-only | M11 §D | **Future**, with media |
|
||||
| 21 | A hostile host on the trusted LAN; the DNS-rebinding interval | `SECURITY-THREAT-MODEL.md` §10A | Accepted for v1. **Future.** No v1.1 package |
|
||||
| 22 | A stale `chunk_id` in a restored snapshot | M9 | No action. It is a React key, not a live pointer |
|
||||
| 23 | Developer-venv residue (`quickjs`, `psycopg`) | §M | No action. It does not ship |
|
||||
| 24 | A Settings warning comparing the real window to the budget | M8 debt, "unowned" | **Closed by M11**: Settings reports the window and warns |
|
||||
| 25 | Root `README.md` has stale facts: "1,191 backend tests" (1,421 at closeout), and the Screenshots paragraph calls screenshots "a job for the UI pass in M8" | found in this review | Documentation debt, not release status, so it is not changed in v4.1. It is fixed in WP-A1's documentation update |
|
||||
|
||||
---
|
||||
|
||||
## 7. Grouping decisions
|
||||
|
||||
**The suggested "narrator boundary hardening" package is split into A1 and A2.**
|
||||
They share a theme but not a mechanism or a test method.
|
||||
|
||||
- **A1**, the context reserve, is budget arithmetic in `context/builder.py` and
|
||||
`contextwindow.py`. It is proved deterministically with a scripted provider
|
||||
that reports token counts.
|
||||
- **A2**, the protocol echo, is the contract between the prompt's wording
|
||||
(`extract.EMIT_RULE`, `events.vocabulary_for_prompt`, the length hint) and the
|
||||
extractor. It is proved with a replay corpus of real narration and
|
||||
real-model runs.
|
||||
|
||||
Bundling them would make one review carry two unrelated risk profiles.
|
||||
|
||||
A2's three items belong together. The example slugs and the call notation are
|
||||
the *source* of the copied text, and the extractor is its *sink*. Changing one
|
||||
without the other means proving the change twice. Both also alter the stored
|
||||
prompt, so they invalidate the same evidence.
|
||||
|
||||
**B stands alone.** It is the only package whose success depends on model output
|
||||
quality. Its test design, a fact that nothing but memory carries, is the hard
|
||||
part. It follows A2, so that it measures the prompt v1.1 will ship.
|
||||
|
||||
**D is narrowed to recovery honesty**, meaning `integrity_check` and the export
|
||||
warning. Both are small and deterministic, and both are recovery-path changes
|
||||
with no policy question. Scheduled backups need owner decisions and have real
|
||||
design risk, including write-lock contention. Raising or streaming the import
|
||||
limit changes a DoS guard. Both move to v1.2.
|
||||
|
||||
**E stays focused** on control boundaries. The same measurement found no other
|
||||
accessibility defect (M11 §L).
|
||||
|
||||
**F is not a package.** Nothing concrete in the product is deficient. The
|
||||
diagnostic's known blind spots are a stale scene, memory off, and one window
|
||||
size. The first two can be addressed by running it differently: the v1.1 gate
|
||||
runs it with memory on. The stale-scene detector is new diagnostic work with no
|
||||
occurrence to justify it yet, so it is v1.2.
|
||||
|
||||
**G is not in v1.1.** It adds no reliability or quality to existing stories, it
|
||||
brings dependencies and a new endpoint surface, and its tests (K05, K06) are
|
||||
FUTURE. `MEDIA-EXTENSION-CONTRACT.md` and the M10 seam remain authoritative for
|
||||
whenever it is taken up.
|
||||
|
||||
---
|
||||
|
||||
## 8. Work packages
|
||||
|
||||
### WP-A1 — Context-window safety reserve
|
||||
|
||||
**Objective.** An assembled prompt plus its reply reserve stays at least a
|
||||
documented safety reserve below the effective window, whatever tokenizer the
|
||||
narrator uses. Where the server reports its own prompt count, a turn that
|
||||
exceeded the window, or was evidently truncated, is recorded and shown, never
|
||||
silent.
|
||||
|
||||
**Rationale.** §15.2 closed "budget larger than the window". It did not close
|
||||
"our count is not the server's count". Measured headroom was 23-42 tokens
|
||||
(item 1), and the fixed 64-token margin absorbs separators as well as tokenizer
|
||||
drift. A narrator whose tokenizer runs a few percent heavier than
|
||||
`cl100k_base` overflows a full 16,384 window. The failure deletes the canon at
|
||||
the front of the prompt with every request returning 200. The data to detect it
|
||||
is already stored with each turn and unread.
|
||||
|
||||
**Scope.**
|
||||
- A safety reserve that scales with the effective budget and has a floor. It
|
||||
replaces or supplements `OUTPUT_SAFETY_MARGIN`, is defined once, and applies
|
||||
to verified windows, declared overrides and an unverified configured budget
|
||||
alike. The brief states the tokenizer divergence it is sized to absorb.
|
||||
- Reading the server-reported prompt-token count from the turn's usage. It is
|
||||
recorded beside the application's count in the turn's window provenance.
|
||||
Recorded states:
|
||||
- `fits`;
|
||||
- `exceeded`: reported count plus reply reserve is over the window;
|
||||
- `truncation_suspected`: the reported count is materially below the
|
||||
application's count;
|
||||
- `unknown`: no usage was reported.
|
||||
|
||||
`exceeded` and `truncation_suspected` are surfaced in the context inspector,
|
||||
and in the turn's response so the reader sees a notice.
|
||||
- **Optional, decided in the brief:** calibrating the reserve from observed
|
||||
server-to-application ratios per endpoint and model. It must be bounded and
|
||||
quantised so that it does not re-price the prompt prefix every turn (§15.3).
|
||||
Detection is required either way.
|
||||
- `ContextOverflow` remains the explicit failure when protected context plus
|
||||
both reserves exceeds the budget.
|
||||
- `TECHNICAL-DESIGN.md` §15.2 and `DEVELOPMENT.md`'s context-window section, as
|
||||
implemented. The root `README.md` stale facts from item 25.
|
||||
|
||||
**Non-scope.** The history block trim mechanism, section order, the knowledge
|
||||
budget share, a per-model tokenizer or any tokenizer download, a hard-coded
|
||||
window, raising any window, a provider abstraction, a Settings redesign, and
|
||||
failing or discarding a turn after the fact.
|
||||
|
||||
**Likely affected.** `backend/app/context/builder.py`,
|
||||
`backend/app/contextwindow.py`, `backend/app/routers/adventures/turns.py` and
|
||||
`insights.py`, `backend/app/providers/openai_compatible.py` (usage, read-only),
|
||||
the context inspector panel in `frontend/src/pages/Play/`. Tests:
|
||||
`test_m11_context_window.py`, `test_m11_declared_window.py`,
|
||||
`test_history_block_trim.py`.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. **Per configuration.** The application's count plus the output reserve plus
|
||||
the safety reserve is no more than the effective budget, and the safety
|
||||
reserve is at least its documented value. This holds for each of: a
|
||||
verified 4,096 window, a verified 16,384 window, a declared override with no
|
||||
probe, and an unverified configured budget. There is a test per
|
||||
configuration.
|
||||
2. **Divergence.** A scripted provider reports prompt counts at 1.00×, and at
|
||||
the brief's stated divergence, of the application's count, across a campaign
|
||||
that fills the window. At both, no turn's reported count plus reply reserve
|
||||
exceeds the window. At 1.25×, either no turn exceeds (calibration built) or
|
||||
every exceeding turn is recorded `exceeded`. **An unflagged overflow fails.**
|
||||
3. **Truncation.** A scripted provider reports 8,194 tokens against an
|
||||
application count near 15,700. The turn is recorded `truncation_suspected`
|
||||
in stored provenance, shown in the inspector, and reported in the turn's
|
||||
response. A provider that reports no usage records `unknown`, never `fits`.
|
||||
4. **Explicit overflow.** Protected context that fits without the safety
|
||||
reserve but not with it raises `ContextOverflow` with an actionable message.
|
||||
The story and state are unchanged (A05, L01).
|
||||
5. **Canon survives.** `test_the_canon_at_the_front_survives_a_window_far_too_small`
|
||||
and `test_without_the_cap_the_same_prompt_would_have_overflowed` both pass
|
||||
with the reserve in place.
|
||||
6. **Cache stability.** At steady state the history floor still moves in blocks
|
||||
(`test_history_block_trim.py`). If calibration is built, a test shows the
|
||||
budget changes only when the observed ratio crosses a stated step.
|
||||
7. **Real model.** A long run of at least 50 turns runs at a 16,384 window with
|
||||
memory on, using the reference narrator on the GPU host. Its ten largest
|
||||
stored prompts are re-counted by the server, as §N did. Every one leaves at
|
||||
least the documented reserve, and every turn's provenance reads `fits`. The
|
||||
report carries the headroom table beside v1's.
|
||||
|
||||
**Regression requirements.** The full backend and frontend suites; F03, F04 and
|
||||
M03 coverage; the I-series (a new provenance key must import and export); the
|
||||
offline container; the browser harness's F05 inspector checks.
|
||||
|
||||
**Dependencies.** None. First.
|
||||
|
||||
**Compatibility.**
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Databases | None expected. Provenance lives in the turn snapshot's JSON, so no migration |
|
||||
| Bundles | None. Snapshots carry new keys, and absence means "not recorded" |
|
||||
| Save Points, branches, state, memories, knowledge | None |
|
||||
| Model settings | None preferred. If a setting is added, it takes a default that reproduces the documented reserve, plus an additive migration |
|
||||
| Behaviour | At the largest prompts the history window holds slightly fewer actions. Nothing stored changes |
|
||||
| Docker and local-only | None |
|
||||
|
||||
**Test modes.** Deterministic: yes. Real model: yes. Browser: inspector display,
|
||||
via the harness. Offline: regression only. Long run: yes, at least 50 turns.
|
||||
|
||||
**Risk.** Low implementation risk and high value. Its blast radius is the
|
||||
budget arithmetic of every turn. It changes prompts at full windows, so it
|
||||
invalidates the v1 long-run evidence for v1.1 (§10).
|
||||
|
||||
---
|
||||
|
||||
### WP-A2 — Protocol echo hardening and genre-neutral instructions
|
||||
|
||||
**Objective.** Stored narration no longer keeps the application-owned protocol
|
||||
shapes a narrator copies, and the fixed instructions stop supplying
|
||||
genre-specific identifiers. Recognition uses only strings and syntax the
|
||||
application itself owns, so story prose is not removed.
|
||||
|
||||
**Rationale.** Items 3 and 4. The call notation in the prompt is copied, the
|
||||
length hint's echo escapes `_is_echoed_instruction`, and a lone trailing
|
||||
`Scene:` survives. Each leak is replayed as history, so it teaches the next turn.
|
||||
The example slugs were copied into an office scene.
|
||||
|
||||
**Scope.**
|
||||
- A genre-neutral `EMIT_RULE` example and slug examples. No identifier in the
|
||||
fixed instructions names a fixture entity or a genre noun.
|
||||
- A decision, by measurement, on whether the vocabulary is shown in a
|
||||
notation that does not look callable. The extractor handles call lines
|
||||
regardless.
|
||||
- The length hint's wording becomes a named constant shared by the builder and
|
||||
the extractor, in the way `render.SECTION_HEADINGS` is shared. The extractor
|
||||
then recognises:
|
||||
- a line consisting only of a call to an event named in `events.SPECS`
|
||||
(case-insensitive), optionally `>`-quoted;
|
||||
- the length hint echoed as a trailing bracket, closed or cut off;
|
||||
- a renderer heading that is the reply's final line with nothing after it.
|
||||
- A replay tool that runs the old and new extractors over a corpus of stored AI
|
||||
turns and writes a per-turn diff report.
|
||||
- Measurement only: per-turn counts of state fact-line restatements,
|
||||
model-invented headings, and proposal outcomes (applied, empty, unparseable,
|
||||
refused), added to `tools/m11_long_run.py`.
|
||||
|
||||
**Non-scope.** Removing a fact restated inside a sentence; removing
|
||||
model-invented headings that are not the renderer's; the event vocabulary itself,
|
||||
the validator, the fence protocol, or history replay; any model-based cleanup
|
||||
pass; rewriting stored narration of existing campaigns.
|
||||
|
||||
**Likely affected.** `backend/app/narrative/extract.py`, `events.py` (prompt
|
||||
rendering only), `render.py` (constants), `backend/app/context/builder.py`
|
||||
(length-hint constant), `tests/test_narrative_state.py`, `tests/test_m11_scifi.py`
|
||||
(the J03 vocabulary check), `tools/m11_long_run.py`. Also
|
||||
`TECHNICAL-DESIGN.md` §15.4 and ADR 013's implementation note.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. **Known shapes are removed.** There is a committed case per shape, cut down
|
||||
from the four §S.6 turns and anonymised:
|
||||
- `> set_possession(silver-key, "alice")`;
|
||||
- `> Create_entity(...)`;
|
||||
- the length hint echoed as `[Hard limit: …append the state block well
|
||||
inside the limit.]`, closed and cut off;
|
||||
- a lone trailing `Scene:`.
|
||||
|
||||
The four stored §S.6 texts, replayed, lose exactly those lines. Every other
|
||||
sentence is intact, asserted as set equality of the remaining sentences.
|
||||
2. **Adversarial prose is untouched.** Each of these is a test:
|
||||
- dialogue that mentions `set_possession` mid-sentence;
|
||||
- `create_entity(ship)` inside a story's own ```python fence;
|
||||
- a call-shaped line whose name is not in the vocabulary, such as
|
||||
`> open_door(north)`;
|
||||
- an in-world bracket, `[Hard limit of the reactor: three hours]`;
|
||||
- `Scene:` followed by prose;
|
||||
- `Memory: she remembered the bells`;
|
||||
- a fact restated inside a sentence;
|
||||
- a trailing bracket about a state block that lacks the hint's wording.
|
||||
|
||||
§O.8's existing negative controls all still pass.
|
||||
3. **Replay corpus.** Every stored AI turn available from the v1 evidence runs
|
||||
is replayed through both extractors: §O.8's 443 and the identity run's 10,
|
||||
kept under `$HOME`, not committed. The package passes when all of these
|
||||
hold:
|
||||
- no turn changes except by removing a criterion-1 shape;
|
||||
- every changed turn appears in the diff report, and the WP report reviews
|
||||
it;
|
||||
- the review classifies zero changes as removed story.
|
||||
|
||||
The report gives the counts.
|
||||
4. **False-positive detection.** The replay tool flags any removal that is not
|
||||
the reply's final segment or a whole line matching a criterion-1 shape, and
|
||||
any removal of more than a stated share of a turn's prose. An unreviewed
|
||||
flag fails the package.
|
||||
5. **Genre-neutral instructions.** A test asserts that `EMIT_RULE`,
|
||||
`EMIT_REMINDER`, the length hints and the vocabulary text contain no
|
||||
identifier from either acceptance fixture (Westhaven, Persephone) and none of
|
||||
J03's genre nouns.
|
||||
6. **Real model.**
|
||||
- **Identity diagnostic.** The office fixture at 4,096 on the CPU/HTTPS host
|
||||
gives 0 proposals naming an example identifier (baseline 1 of 10),
|
||||
protocol shapes in 0 of 10 stored turns (baseline 4 of 10), and 0 identity
|
||||
signals. Its `--scripted --inject` self-test still fires.
|
||||
- **Long run.** At least 50 turns at 16,384 with memory on stores protocol in
|
||||
0 turns (baseline 0 of 104). Restatement counts are recorded against the
|
||||
74 of 104 baseline, as a measurement, not a gate.
|
||||
|
||||
**Regression requirements.** The full backend suite, including the 25 §O.8
|
||||
cases; C06 and H05 (refusals are still recorded and shown to the model); J01-J03;
|
||||
the I-series; the browser harness's hostile-narration checks (H04, H06, H07).
|
||||
|
||||
**Dependencies.** After A1. Both change the prompt, and one real-model run can
|
||||
then cover both. The brief can be written while A1 is under review.
|
||||
|
||||
**Compatibility.**
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Databases | None. No migration |
|
||||
| Stored narration of v1.0.0 campaigns | Not rewritten |
|
||||
| Bundles | None |
|
||||
| Save Points, branches | None |
|
||||
| State | None; the validator is unchanged |
|
||||
| Memories and summaries | Future ones summarise cleaner prose. Existing ones are untouched |
|
||||
| Docker and local-only | None |
|
||||
|
||||
**Test modes.** Deterministic: yes. Replay corpus: yes. Real model: yes. Browser:
|
||||
regression only. Offline: regression only. Long run: yes, at least 50 turns, and
|
||||
this run may be the same one as A1's criterion 7 if A1 is already merged.
|
||||
|
||||
**Risk.** Medium. The failure to fear is silent removal of prose. Constants-only
|
||||
recognition, the corpus and the detector are the mitigations.
|
||||
|
||||
---
|
||||
|
||||
### WP-B — Independent long-term memory retention
|
||||
|
||||
**Objective.** An important fact planted early is recoverable at depth 100 or
|
||||
more from the memory bank itself, with provenance to a memory whose source range
|
||||
covers the planting turn. This must hold when neither authoritative state, later
|
||||
narration, the summary nor the history window carries the fact, and lineage
|
||||
safety must be unchanged.
|
||||
|
||||
**Rationale.** Item 5. F02 and M04 pass on state-based recovery, which the owner
|
||||
accepted. Memory's own retention is unproven, and in a long campaign whose facts
|
||||
are not all state-shaped it is the only continuity there is.
|
||||
|
||||
**Scope.** Diagnosis first, then the smallest sufficient fix.
|
||||
- **B.1, the diagnostic.** A retention harness that reports, for a planted fact,
|
||||
each stage with ids and depths:
|
||||
- **created:** does a memory whose `source_start`..`source_end` covers the
|
||||
planting turn contain the fact?
|
||||
- **retained:** is it not `forgotten`?
|
||||
- **ranked:** where is it in similarity order for the recall query?
|
||||
- **injected:** is it in the memories section?
|
||||
|
||||
The harness runs in two modes. The deterministic mode uses a scripted
|
||||
narrator, summariser and embedder. The real-model mode uses the reference
|
||||
narrator and embedder. It also adds a `recovered_through_memory_independent`
|
||||
verdict to `tools/m11_long_run.py`, keeping the existing verdicts and their
|
||||
meanings.
|
||||
- **B.2, the fix,** only at the stages B.1 shows failing, from this bounded menu:
|
||||
- how the summariser's excerpt is chosen, instead of keeping only the last
|
||||
2,000 tokens;
|
||||
- the memory prompt's instruction to keep named facts and objects;
|
||||
- an eviction rule that does not throw out a never-retrieved early memory
|
||||
first;
|
||||
- one additional ranking term (lexical or entity overlap, CONTEXT-AND-MEMORY
|
||||
§20).
|
||||
|
||||
Anything outside the menu needs the owner's agreement in the brief.
|
||||
|
||||
**Non-scope.**
|
||||
- Memory becoming authoritative, or outranking state (F07).
|
||||
- Any change to lineage attachment or filtering (`tree.attach_memory`,
|
||||
`forget_node`, E02) or summary lineage (E03).
|
||||
- Merging memory with imported knowledge, or new embedding models or
|
||||
dependencies.
|
||||
- Automatic re-summarisation of existing memories.
|
||||
- Redesigning cross-layer duplication (§22).
|
||||
|
||||
**Likely affected.** `backend/app/memorybank.py` (creation, eviction, ranking);
|
||||
`backend/app/context/builder.py` (memories section, only if the query changes);
|
||||
`backend/app/models.py`, only if a field is unavoidable. Tools: `tools/memory_ab.py`,
|
||||
`tools/m11_long_run.py`. Tests: `test_context_memory.py`, `test_memory_nodes.py`,
|
||||
`test_m11_leakage.py`, `test_m11_long_run_memory.py`.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. **Deterministic retention test**, committed. Fact F is planted at depth 3 or
|
||||
less through narration only, with no state correction and no knowledge source.
|
||||
These are asserted for the whole run:
|
||||
- F is never in authoritative state;
|
||||
- no turn after the planting block contains F's sentinel tokens;
|
||||
- the summary section never contains F;
|
||||
- the planting turn is outside the history window at recall.
|
||||
|
||||
At a recall depth of 100 or more, the memories section holds a memory that
|
||||
contains F and whose source range includes the planting depth, and the
|
||||
context report lists its id and similarity. **The test must fail on the
|
||||
`432f041` tree**, and the report must name the stage at which it fails.
|
||||
2. **Past capacity.** The same test runs with more memories written than
|
||||
`memory_bank_capacity`. The planting-era memory is still active and
|
||||
retrieved, or its eviction follows a documented, tested rule that the WP
|
||||
report justifies. No pinned memory is evicted. The "frozen bank" regression
|
||||
(a new memory evicted at once) still passes.
|
||||
3. **Lineage.** Fact G is planted only on a line later abandoned by Undo and
|
||||
divergence. G's memory stays on disk and is absent from every active-line
|
||||
prompt and from `memories.used`. All 14 tests in `test_m11_leakage.py` and the
|
||||
E02 tests pass unchanged.
|
||||
4. **Authority.** A memory contradicting state loses, and memory is still framed
|
||||
as non-canon (F07).
|
||||
5. **Real model.** A run of at least 100 turns at 16,384 with memory on, on the
|
||||
GPU host with logging. The fact is chosen so that the harness can verify the
|
||||
criterion-1 preconditions on real output. **Pass:** one run meets every
|
||||
precondition and returns `recovered_through_memory_independent`, with the
|
||||
memory's provenance. A run whose precondition fails reports which one, and
|
||||
counts as neither pass nor failure. The report gives the stage-by-stage
|
||||
diagnostic for every run.
|
||||
6. **Budget.** The memories section stays inside its budget (F03), and the
|
||||
steady-state prompt is no larger than A1's arithmetic allows.
|
||||
7. `CONTEXT-AND-MEMORY.md` §15, §20 and §21 are updated as implemented.
|
||||
|
||||
**Regression requirements.** F01-F08, E01-E04, M04 verdicts (the new verdict is
|
||||
added and the old ones are unchanged), the I-series (memories travel in the
|
||||
bundle), L04 (`test_memory_rewrite.py`), §O.7's no-write-lock and
|
||||
failure-recording tests.
|
||||
|
||||
**Dependencies.** After A2, so that it measures the prompt v1.1 ships and a
|
||||
narrator no longer handed protocol to restate. B.1's deterministic harness may
|
||||
be built earlier.
|
||||
|
||||
**Compatibility.**
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Databases | No migration preferred. If a memory field is unavoidable, it is additive with a default, tested from a real v1.0.0 database, with schema parity |
|
||||
| Bundles | Stay v3. Any new memory field is optional on import, and its absence means "unknown". v1.0.0 bundles import |
|
||||
| Existing memories | Valid and used as they are. A rebuild stays opt-in (`tools/rewrite_memories.py`) |
|
||||
| Branch history, Save Points, state, knowledge | None |
|
||||
| Settings | Any change to capacity semantics keeps v1 defaults |
|
||||
| Docker and local-only | None |
|
||||
|
||||
**Test modes.** Deterministic: yes. Real model: yes. Browser: no. Offline:
|
||||
regression only. Long run: yes, at least 100 turns, possibly several.
|
||||
|
||||
**Risk.** High uncertainty, because the outcome depends on the model, and medium
|
||||
implementation risk. The blast radius is the memory subsystem.
|
||||
|
||||
---
|
||||
|
||||
### WP-C — Browser release coverage
|
||||
|
||||
**Objective.** Every reader-facing behaviour that v1 proved only through the API
|
||||
or the component suite is driven in a real browser, including an export that
|
||||
actually leaves the browser as a file.
|
||||
|
||||
**Rationale.** Item 6. §T carries two qualifications that exist only because the
|
||||
harness stops short.
|
||||
|
||||
**Scope.** New scenarios in `tools/m11_browser.py`, or a sibling that reuses
|
||||
`tools/m11_webdriver.py`: Retry; Save Point create and restore; state correction;
|
||||
narration length; failed generation; export download. Also the Firefox profile
|
||||
preferences that direct a download to a harness-owned directory under `$HOME`.
|
||||
The package is harness-only, unless it finds a product defect or a control with
|
||||
no accessible name. Either is fixed with a regression test and reported as a
|
||||
product change.
|
||||
|
||||
**Non-scope.** New UI, frontend refactors, Selenium or any new dependency,
|
||||
screenshot diffing, tablet layout, CI.
|
||||
|
||||
**Likely affected.** `backend/tools/m11_browser.py`, `backend/tools/m11_webdriver.py`,
|
||||
the `DEVELOPMENT.md` harness section.
|
||||
|
||||
**Acceptance criteria.** The run ends with 0 failed and 0 skipped. It runs against
|
||||
the built SPA served by FastAPI, with turns from the reference narrator over
|
||||
trusted-LAN HTTPS. Every assertion reads the rendered DOM or a file on disk.
|
||||
1. **Retry.** Retry on the newest turn yields a second take, and the indicator
|
||||
reads 2/2. Stepping to 1/2 shows the original narration unchanged. The takes
|
||||
persist after a reload.
|
||||
2. **Save Point.** Create a named Save Point through the UI, play two more
|
||||
turns, then restore it through the UI. The transcript ends at the named
|
||||
moment, the position indicator says later story is ahead, and Redo walks
|
||||
into the later turns. After a reload the Save Point is still listed.
|
||||
3. **State correction.** A correction submitted through the State panel applies
|
||||
and is still shown after a reload. A partly refused correction shows its
|
||||
refusal and reason to the reader.
|
||||
4. **Narration length.** After changing the control to brief and playing a turn,
|
||||
then to long and playing a turn, each turn's context inspector shows its
|
||||
band's word range.
|
||||
5. **Failed generation.** With the model set through Settings to a name the
|
||||
server does not serve, a submitted turn shows an error, adds no narration and
|
||||
keeps the typed input. Setting the model back, the next turn succeeds and the
|
||||
earlier story is intact.
|
||||
6. **Export download.** A real click on Export, from both the campaign library
|
||||
and campaign settings, writes a file with no manual step. The file:
|
||||
- exists and is not empty;
|
||||
- parses as `ai-dnd-adventure-v3`;
|
||||
- imports into a fresh data directory with the same action count, head
|
||||
position and Save Points.
|
||||
|
||||
If the snap Firefox cannot be made to download, the check runs on a non-snap
|
||||
Firefox under `$HOME`, and `DEVELOPMENT.md` says so. **Exercising only the
|
||||
click does not pass.**
|
||||
7. The existing 38 checks pass in the same run.
|
||||
|
||||
**Regression requirements.** The existing browser checks; the frontend suite
|
||||
where a product fix is made.
|
||||
|
||||
**Dependencies.** None. It can run at any point. Its final run is repeated on the
|
||||
v1.1 release candidate.
|
||||
|
||||
**Compatibility.** None, unless a product fix is made, and then per §5.
|
||||
|
||||
**Test modes.** Browser: yes. Real model: yes, over the HTTPS reference host.
|
||||
Deterministic: no. Offline: no. Long run: no.
|
||||
|
||||
**Risk.** Low product risk. The medium risk is harness flakiness, so every wait is
|
||||
on a DOM condition, never a sleep used as an assertion.
|
||||
|
||||
---
|
||||
|
||||
### WP-D — Recovery honesty
|
||||
|
||||
**Objective.** A backup is kept only after a full integrity check. A campaign
|
||||
whose export exceeds what import accepts is exported with a warning saying so,
|
||||
not silently.
|
||||
|
||||
**Rationale.** Items 7 and 10a. Both are recovery gaps that are known, measured
|
||||
and cheap to close. Neither needs a policy decision.
|
||||
|
||||
**Scope.**
|
||||
- `backup.py` runs `PRAGMA integrity_check` on the finished copy, and the time
|
||||
it takes is measured.
|
||||
- Export compares the serialised bundle's size with
|
||||
`limits.MAX_IMPORT_BODY_BYTES`. When it is over, the file is still delivered
|
||||
and the reader sees a warning naming the limit and what it means.
|
||||
- The import refusal for an oversized bundle names the limit.
|
||||
- `DEVELOPMENT.md` documents the ceiling as measured: about 13 kB per action on
|
||||
a real campaign, and M9's conservative figure of about 279 turns.
|
||||
|
||||
**Non-scope.** Scheduled backups; raising the import limit or streaming import; a
|
||||
bundle format change or further compression; a restore button.
|
||||
|
||||
**Likely affected.** `backend/app/backup.py`, `backend/app/routers/adventures/bundle_io.py`,
|
||||
`frontend/src/pages/Campaigns.jsx`,
|
||||
`frontend/src/pages/Play/panels/CampaignSettingsPanel.jsx`,
|
||||
`frontend/src/pages/backup.test.jsx`, and the backup and bundle tests.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. A backup of a healthy database reports `integrity_check` ok and is kept.
|
||||
2. A copy with damage that `integrity_check` detects and `quick_check` does not,
|
||||
such as an index inconsistent with its table, is rejected and not kept.
|
||||
Existing backups are still never overwritten.
|
||||
3. The time for `integrity_check` is recorded on the 100-turn evidence database
|
||||
and on a synthetic database of 100 MB or more, and the backup completes
|
||||
through the UI on both.
|
||||
4. Exporting a fixture campaign over 20 MB succeeds, delivers the file, and
|
||||
shows a warning naming the import limit, asserted in the API response and in
|
||||
a component test. A campaign under the limit shows no warning.
|
||||
5. Importing that bundle is refused with a message naming the limit.
|
||||
6. A normal campaign's export is byte-identical before and after the package,
|
||||
apart from timestamp fields.
|
||||
|
||||
**Regression requirements.** I01-I07, L01-L04, the backup case in
|
||||
`test_m11_migration.py`, `backup.test.jsx`, and the offline container's
|
||||
export and import.
|
||||
|
||||
**Dependencies.** None.
|
||||
|
||||
**Compatibility.** Databases and bundles: no change. Backups: a stricter check,
|
||||
with the same file.
|
||||
|
||||
**Test modes.** Deterministic: yes. Browser: the warning, optionally in WP-C's
|
||||
harness. Offline: regression only. Real model: no. Long run: no.
|
||||
|
||||
**Risk.** Low.
|
||||
|
||||
---
|
||||
|
||||
### WP-E — Control-boundary contrast
|
||||
|
||||
**Objective.** Control boundaries meet WCAG 1.4.11's 3:1 against their panel, at
|
||||
rest and on hover.
|
||||
|
||||
**Rationale.** Item 9.
|
||||
|
||||
**Scope.**
|
||||
- The boundary token values in `frontend/src/styles/tokens.css`, and any
|
||||
component that overrides them.
|
||||
- `tools/contrast_audit.py` treats boundary pairs below 3:1 as failures, not
|
||||
advisories.
|
||||
- The browser harness measures rendered boundary contrast.
|
||||
- Before-and-after screenshots for the owner's approval.
|
||||
|
||||
**Non-scope.** A palette redesign, typography, layout, tablet work, a
|
||||
screen-reader audit, and other WCAG criteria. A defect the same measurement finds
|
||||
is recorded, not taken on.
|
||||
|
||||
**Likely affected.** `frontend/src/styles/tokens.css`, `backend/tools/contrast_audit.py`,
|
||||
`backend/tools/m11_browser.py`, `frontend/src/a11y.test.jsx`.
|
||||
|
||||
**Acceptance criteria.**
|
||||
1. `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and
|
||||
exits 0 on the package tree.
|
||||
2. Every text pair still clears 4.5:1 (1.4.3). The rendered text contrasts
|
||||
measured by the harness do not fall below v1's (14.57, 5.48, 13.57 and
|
||||
5.88:1) without a stated reason.
|
||||
3. The rendered boundary of the story input and of a primary control, at rest
|
||||
and on hover, is at least 3:1. The focus indicator is still visible.
|
||||
4. The owner approves the before-and-after screenshots, and the WP report records
|
||||
the approval.
|
||||
|
||||
**Regression requirements.** The frontend suite and lint; the browser harness's
|
||||
accessibility checks.
|
||||
|
||||
**Dependencies.** None. Its browser checks go into WP-C's harness if WP-C has
|
||||
landed, and into `m11_browser.py` otherwise.
|
||||
|
||||
**Compatibility.** None.
|
||||
|
||||
**Test modes.** Deterministic: yes. Browser: yes. Everything else: no.
|
||||
|
||||
**Risk.** Low.
|
||||
|
||||
---
|
||||
|
||||
## 9. Order and dependencies
|
||||
|
||||
```text
|
||||
A1 context safety reserve ──► A2 protocol echo + neutral instructions ──► B memory retention
|
||||
(B.1 harness may start early)
|
||||
C browser coverage ── independent
|
||||
D recovery honesty ── independent
|
||||
E boundary contrast ── independent (checks land in C's harness if C is first)
|
||||
│
|
||||
▼
|
||||
v1.1 release validation (§10)
|
||||
```
|
||||
|
||||
**Recommended order: A1, A2, B, C, D, E.**
|
||||
|
||||
- **A1 first.** It is the only item that can silently remove canon from a
|
||||
prompt. It is deterministic to test, small in blast radius, and it
|
||||
establishes the provenance that A2's and B's real-model runs will read.
|
||||
- **A2 second.** Its leak is silent and compounds through replay. It must
|
||||
precede B, because it changes what the memory pass summarises.
|
||||
- **B third.** It carries the highest continuity value and the most uncertainty.
|
||||
Measuring it before A1 and A2 settle would measure a prompt v1.1 does not ship.
|
||||
- **C, D and E** close verification, recovery and accessibility gaps with no
|
||||
known story risk. They depend on nothing, so the owner may move any of them
|
||||
earlier. While B waits on long runs is a natural slot. Each still has its own
|
||||
brief and review.
|
||||
|
||||
| Property | A1 | A2 | B | C | D | E |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| Can begin independently | yes | after A1 | after A2 | yes | yes | yes |
|
||||
| Schema or migration | no | no | avoid; additive if unavoidable | no | no | no |
|
||||
| Export format | no (additive evidence key) | no | no; optional field at most | no | no | no |
|
||||
| Changes acceptance tests (`V1-ACCEPTANCE-TESTS.md`) | no | no | no | no | no | no |
|
||||
| Changes the stored prompt | yes | yes | possibly | no | no | no |
|
||||
| Real-model validation | yes | yes | yes | yes, for turns | no | no |
|
||||
| Browser testing | regression | regression | no | yes | optional | yes |
|
||||
| Offline / no-network testing | regression | regression | regression | no | regression | no |
|
||||
| Long-run testing | yes, 50 turns or more | yes, 50 turns or more | yes, 100 turns or more | no | no | no |
|
||||
| Risk | low | medium | high uncertainty | low | low | low |
|
||||
|
||||
## 10. Compatibility summary
|
||||
|
||||
| Area | A1 | A2 | B | C | D | E |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| v1.0.0 campaign databases | open unchanged | unchanged | unchanged; an additive migration only if unavoidable | — | unchanged | — |
|
||||
| Export and import bundles | v3; new evidence key | — | v3; an optional field at most | — | v3; export warns | — |
|
||||
| Save Points | — | — | — | — | — | — |
|
||||
| Branch history | — | — | lineage rules unchanged | — | — | — |
|
||||
| Narrative state | — | validator unchanged | memory never outranks state | — | — | — |
|
||||
| Memories and summaries | — | new ones from cleaner prose | creation, retention and ranking change; existing rows kept | — | — | — |
|
||||
| Knowledge sources | — | — | — | — | — | — |
|
||||
| Model settings | none preferred | — | capacity defaults kept | — | — | — |
|
||||
| Docker and local-only | — | — | — | — | — | — |
|
||||
|
||||
## 11. v1.1 release criteria
|
||||
|
||||
v1.1 is not called v1.1.0 until all of the following hold on one release
|
||||
candidate tree:
|
||||
|
||||
1. **Every package in scope is accepted**, each with its report, and with its
|
||||
acceptance criteria passing on the candidate or on a tree whose product code
|
||||
the candidate carries unchanged.
|
||||
2. **The v1 contract still passes.** All 82 REQUIRED FOR V1 tests hold, with H09
|
||||
not applicable on the same condition. None is relaxed.
|
||||
3. **Suites and builds:** the backend suite, the frontend suite and lint, the
|
||||
production build, and `docker build --no-cache`, with the image's SPA
|
||||
identical to the local build.
|
||||
4. **Offline:** `tools/m11_offline.py` passes all checks with no network and a
|
||||
fresh volume.
|
||||
5. **Browser:** the existing 38 checks plus WP-C's and WP-E's pass with 0 failed
|
||||
and 0 skipped, on the candidate, over trusted-LAN HTTPS.
|
||||
6. **One v1.1 long run** on the candidate's product code: 100 turns or more at a
|
||||
16,384 window with memory on, on the GPU host with logging. It passes M01-M04.
|
||||
A1's headroom table shows the documented reserve on every re-counted prompt,
|
||||
and every turn is `fits`. A2's leak count is 0. B's independent-retention
|
||||
verdict is recorded.
|
||||
12. **WP-B's memory limitation is reported, not summarised away.** The v1.1
|
||||
release report states each of these, and never shortens them to "WP-B
|
||||
passed":
|
||||
- deterministic independent-memory recovery: **PASS**;
|
||||
- reference-model independent-memory recovery: **FAIL** on the
|
||||
precondition-valid attempt;
|
||||
- the failing stage: **memory creation**, the summariser's content
|
||||
selection;
|
||||
- the owner's decision to accept that limitation for v1.1.
|
||||
|
||||
The release long run's independent-retention verdict (item 6) is read
|
||||
against it. A recovery there is reported as evidence, not as a reversal of
|
||||
the limitation, unless it meets every isolation precondition.
|
||||
13. **Carried residuals are listed with their status:**
|
||||
- the mid-reply narrator instruction echo that A2's trailing cleanup does
|
||||
not remove (WP-B.1 §K);
|
||||
- the doubled full stop in the memory-search scene text (WP-B.2 §R 10).
|
||||
7. **Identity diagnostic** on the candidate with memory on: 0 signals, and 0
|
||||
protocol shapes in stored narration.
|
||||
8. **Recovery:** `tools/m11_recovery.py` on the v1.1 long run's bundle passes all
|
||||
checks.
|
||||
9. **Upgrade from a real v1.0.0 database.** A database is created by the
|
||||
`v1.0.0` tree and played. It has retained history, an undone head, Save
|
||||
Points, memories, summaries, imported knowledge and a narration-length
|
||||
choice, and uses loopback or placeholder settings with no real hostnames.
|
||||
Opened by the candidate, its transcript, head, Redo availability, state, Save
|
||||
Points, memories, knowledge and settings compare identical, and any migration
|
||||
is forward-only with schema parity. A v1.0.0 export imports into v1.1.
|
||||
Because the format stays v3, a v1.1 export of that campaign is checked for
|
||||
import into v1.0.0, and the result is reported.
|
||||
10. **Documentation:** `README.md`, `DEVELOPMENT.md`, the as-implemented sections
|
||||
named by each package, and `VERSION.md`.
|
||||
11. **Owner events**, each separate: the signed release commit, `main`, and a
|
||||
`v1.1.0` tag.
|
||||
|
||||
## 12. Scope recommendation
|
||||
|
||||
**Recommended: Option 1, a focused v1.1.**
|
||||
|
||||
| Ship in v1.1 | Defer to v1.2 | Future / optional |
|
||||
| --- | --- | --- |
|
||||
| WP-A1 context safety reserve | Scheduled backups (10b) | Real media-provider adapters (11; K05, K06) |
|
||||
| WP-A2 protocol echo and neutral instructions | Raising or streaming the import limit (7) | Whole-transcript copy, story search (§77, §78) |
|
||||
| WP-B independent memory retention | Reducing in-sentence restatement (2) | Discarded-history recovery screen (§63) |
|
||||
| WP-C browser release coverage | Stale-scene and derived-contamination detectors in the identity diagnostic (8) | Tablet layout |
|
||||
| WP-D recovery honesty | Cross-layer duplication suppression (17) | Editable RPG world state |
|
||||
| WP-E control-boundary contrast | | Media debt: `ambience`, visual-profile recovery and UI (20) |
|
||||
| | | Trusted-LAN residual limits: address pinning against DNS rebinding (21) |
|
||||
| | | Comparative narrator-model recommendation (13) |
|
||||
|
||||
**Why focused.** Four of the six packages close silent failure modes or
|
||||
verification gaps that the v1 evidence itself named. The other two are small and
|
||||
deterministic. Everything deferred is either a new feature, or needs a policy
|
||||
decision the owner has not been asked for, or rests on an occurrence that has not
|
||||
happened. A broader v1.1 that took scheduled backups and media would add the two
|
||||
packages with the most new surface. It would also push the long-run
|
||||
re-validation, which every prompt change already requires, further from the
|
||||
changes it validates.
|
||||
|
||||
**If B's real-model criterion cannot be met.** Precondition-valid runs may recover
|
||||
nothing even after B.2. In that case the owner chooses between shipping v1.1 with
|
||||
B's deterministic criteria met and the real-model result recorded as a residual
|
||||
risk, or holding v1.1 for B. The plan does not pre-decide this.
|
||||
|
||||
## 13. Coding briefs needed next
|
||||
|
||||
These briefs are not written here, and none is to be executed from this document.
|
||||
|
||||
1. **WP-A1 — context-window safety reserve. Write this one first.** It carries
|
||||
the highest user risk, has no dependencies, is deterministically testable,
|
||||
and produces the provenance that later real-model runs read. The brief must
|
||||
settle three things: the tokenizer divergence the reserve is sized for, whether
|
||||
calibration is built or detection alone, and whether any setting is added
|
||||
(recommended: none).
|
||||
2. WP-A2 — protocol echo hardening and genre-neutral instructions. It can be
|
||||
drafted while A1 is under review.
|
||||
3. WP-B — independent memory retention: B.1 diagnostic, then B.2 fix. Two briefs
|
||||
are an option if B.1's findings should be reviewed before a fix is chosen.
|
||||
4. WP-C — browser release coverage.
|
||||
5. WP-D — recovery honesty.
|
||||
6. WP-E — control-boundary contrast.
|
||||
7. v1.1 release validation, written only after every in-scope package is
|
||||
accepted.
|
||||
|
||||
## 14. Decisions for the owner
|
||||
|
||||
- Confirm Option 1, the focused v1.1.
|
||||
- A1: the divergence the reserve must absorb; calibration or detection only; that
|
||||
a detected truncation is flagged, not turned into a failed turn.
|
||||
- B: whether B.1 and B.2 are one brief or two; the choice in §12 if the
|
||||
real-model criterion is not met.
|
||||
- D: that the 20 MB import limit stays in v1.1.
|
||||
- E: approval of the visual change.
|
||||
- Whether v1.1 real-model validation stays on the reference narrator, as this plan
|
||||
recommends.
|
||||
- The report naming and rotation in `planning/reports/` (§5 rule 5).
|
||||
+146
-3
@@ -1,8 +1,151 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v4.0
|
||||
- **Revision date:** 2026-09-14
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M11 complete**; **M11 accepted at its closeout (2026-09-14), and the v1 release gate passed** on the release-candidate tree. All 82 REQUIRED FOR V1 tests pass, with H09 not applicable. There is no further planned milestone. The signed closeout commit and any `v1.0.0` tag are the repository owner's, and neither exists as of this revision.
|
||||
- **Package:** Adventure Storyteller Planning Package v4.6
|
||||
- **Revision date:** 2026-09-16
|
||||
- **Status:** **v1.0.0 released on 2026-09-14**: the signed tag `v1.0.0` and `main` both point at the signed release commit `432f041`. Milestones M1-M11 complete and closed; all 82 REQUIRED FOR V1 tests pass, with H09 not applicable. **v1.1 is in progress** on `v1.1-development` (`V1.1-PLAN.md`): WP-A1/A2 (`d63804f`), WP-B.1 (`beb17ad`) and WP-B.2 (`0c1ba83`, accepted with a documented real-model memory limitation) are committed. WP-C is signed as `59b5ebc`, and **WP-D (recovery honesty) and WP-E (control-boundary contrast) are signed as `87a4032`**. **Integrated v1.1 release validation has run on candidate `87a4032` and PASSED** (`reports/v1.1/V1.1-RELEASE-REPORT.md`): 82 REQUIRED v1 tests hold (81 PASS, H09 NOT APPLICABLE), backend 1,723 / frontend 175 / lint 0 errors, a `--no-cache` image whose SPA is file-for-file identical to the local build, offline 23/23, browser 101/0/0 over trusted-LAN HTTPS, a 102-turn 16,384-window run passing M01-M04 with every turn `fits` and 0 protocol leaks, identity 0/0, recovery 16/16, a real v1.0.0 upgrade identical on all 15 fields with both bundle directions importing, and a release smoke of 15/15. WP-B's reference-model memory limitation remains an accepted, documented residual. **v1.0.0 is still the released version: no release commit, no `main` update and no `v1.1.0` tag exist** — those are the owner's events.
|
||||
|
||||
## v4.6 — v1.1 integrated release validation (2026-09-16)
|
||||
|
||||
Release validation of candidate `87a4032`, not a work package: no requirement,
|
||||
acceptance test, schema, bundle format or product code changed. The evidence is
|
||||
`reports/v1.1/V1.1-RELEASE-REPORT.md`, sections A-W.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/v1.1/V1.1-RELEASE-REPORT.md` | **New.** The frozen candidate, the 82-row v1 acceptance matrix, every gate's result, the A1 headroom table, the A2 leak count, the WP-B verdict kept in both halves, the residual classification, and the final decision | release report |
|
||||
| `reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL` **PENDING → APPROVED**, sourced and dated to the owner's release-validation brief; the signed commit predated the review | correction of record |
|
||||
| `V1.1-PLAN.md`, `planning/README.md`, `VERSION.md` | Status: WP-D/WP-E signed `87a4032`; validation passed; the three owner events still outstanding | status |
|
||||
| `README.md` | v1.0.0 **remains** released; v1.1 implemented and validated but untagged; schema figure corrected to 94 | product docs |
|
||||
| `backend/tools/v11_upgrade_check.py`, `backend/tools/v11_release_smoke.py` | **New**, harness only: the real-v1.0.0 upgrade gate and the release-shaped smoke test | tooling |
|
||||
| `backend/tools/m11_identity.py` | Reads `AIDND_TEST_EMBED_MODEL` and enables the memory bank, so the diagnostic can run with memory on as the gate requires | tooling |
|
||||
|
||||
**Requirement changes: zero. Product-code changes: zero.**
|
||||
|
||||
**Outcome:** `V1.1 RELEASE VALIDATION: PASS`. Carried residuals: WP-B's
|
||||
reference-model memory limitation, the mid-reply instruction echo (still
|
||||
reproducible on the stored fixture, absent from release evidence), and the
|
||||
doubled full stop. K1 is classified v1.2 backlog. **No release commit, no `main`
|
||||
update, no `v1.1.0` tag.**
|
||||
|
||||
## v4.5 — WP-D recovery honesty and WP-E control-boundary contrast (2026-09-15)
|
||||
|
||||
The last two planned v1.1 packages, implemented and reported separately. No
|
||||
requirement, acceptance test, schema or bundle format changed. WP-D's evidence is
|
||||
in `reports/v1.1/V1.1-WP-D-REPORT.md`, WP-E's in `V1.1-WP-E-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/v1.1/V1.1-WP-D-REPORT.md` | **New.** Full `integrity_check` on the finished backup copy, proved against a fixture `quick_check` calls healthy; an export that says when this version could not import it back, carried in headers because the response body *is* the bundle | work-package report |
|
||||
| `reports/v1.1/V1.1-WP-E-REPORT.md` | **New.** Control boundaries raised to clear WCAG 1.4.11 (3:1), the contrast audit turned from a report into a gate, rendered before/after boundary measurements, and the `.slice-7` finding the change itself created | work-package report |
|
||||
| `V1.1-PLAN.md` | Status: WP-C signed `59b5ebc`; WP-D and WP-E complete and staged | status |
|
||||
| `planning/README.md` | Current state | index |
|
||||
| `DEVELOPMENT.md` | How large an export can get: the 20 MB import ceiling, ~13 kB per action, M9's ~279-turn figure and why it is not a turn limit, and the four export headers | developer docs |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
Two owner decisions are outstanding, both recorded rather than assumed: WP-E's
|
||||
before/after screenshots await approval (`OWNER SCREENSHOT APPROVAL: PENDING`),
|
||||
and WP-D records that no real-browser click was made on *Back up now* — the
|
||||
endpoint that button calls was driven instead.
|
||||
|
||||
## v4.4 — WP-C browser release coverage (2026-09-15)
|
||||
|
||||
Harness work, with one narrow product fix it found. No requirement, acceptance
|
||||
test, schema or bundle format changed. The evidence is in
|
||||
`reports/v1.1/V1.1-WP-C-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/v1.1/V1.1-WP-C-REPORT.md` | **New.** The six reader workflows driven in a real browser, the existing 38 checks, the download environment, harness defects J1-J7, and product defects K1 (open) and K2 (fixed) | work-package report |
|
||||
| `V1.1-PLAN.md` | Status: WP-B.2 committed; WP-C complete and staged | status |
|
||||
| `planning/README.md` | Current state | index |
|
||||
| `DEVELOPMENT.md` | The browser harness command and flags; Firefox download preferences and the `$HOME` rule; what counts as a finished download; waiting on conditions, never sleeping | developer docs |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v4.3 — WP-B.2 independent memory retention, accepted with a documented limitation (2026-09-15)
|
||||
|
||||
WP-B.1's diagnostic (`beb17ad`) placed three memory deficiencies; WP-B.2 corrects
|
||||
exactly those, one at a time, each verified before the next. No requirement or
|
||||
acceptance test changed, and no schema, bundle format or setting default
|
||||
changed. Evidence and the WP-B decision are in
|
||||
`reports/v1.1/V1.1-WP-B2-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1 WP-B.2)**: a block longer than 2,000 tokens is shown to the summariser as its opening and its end with an omission marker, inside the same budget; a block that fits is unchanged; the marker is never stored. | as-implemented record |
|
||||
| `CONTEXT-AND-MEMORY.md` §18 | **As implemented**: the retrieval query is the player's input plus a bounded scene context, recorded per turn. | as-implemented record |
|
||||
| `CONTEXT-AND-MEMORY.md` §20 | **As implemented**: `final = semantic + 0.15 × lexical`, the rarity-weighted lexical term over the input, the sweep that chose the weight, pins and ties. | as-implemented record |
|
||||
| `CONTEXT-AND-MEMORY.md` §15 | **As implemented (v1.1)**: the memory prompt is unchanged from v1.0.0, and the accepted limitation: the reference summariser can omit or misattribute a fact from a block it was given whole. | as-implemented record |
|
||||
| `DEVELOPMENT.md` | The GPU-host kernel/Ollama watch command corrected: `-k -u ollama` matched nothing; the OR form records both. | developer docs |
|
||||
| `CONTEXT-AND-MEMORY.md` §21 | **As implemented**: selection unchanged in shape; eviction ordered by coverage first (boundaries kept, smallest hole first), recency second, v1.0.0 order as fallback. | as-implemented record |
|
||||
| `V1.1-PLAN.md` | Status: WP-B accepted with a documented real-model limitation. §11 release criteria 12 and 13: the v1.1 release report must state the deterministic PASS and reference-model FAIL at memory creation, and list the carried residuals. | status, release gate |
|
||||
| `planning/README.md` | Current state. | index |
|
||||
| `reports/v1.1/V1.1-WP-B2-REPORT.md` | **New.** B2.1-B2.3 designs and evidence, full deterministic acceptance, real-model attempts, the rejected B2.4 prompt experiment, compatibility, offline, and the WP-B disposition. | work-package report |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v4.2 — WP-A1 and WP-A2 implemented (2026-09-14)
|
||||
|
||||
Two v1.1 work packages, implemented in sequence. No requirement or acceptance
|
||||
test changed, and no schema or bundle format changed. Evidence is in
|
||||
`reports/v1.1/V1.1-WP-A1-A2-REPORT.md`.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `TECHNICAL-DESIGN.md` §15.2 | **As implemented (v1.1 WP-A1)**, covering: the safety reserve, `max(256, ceil(5%))`; what replaced M6's 64-token margin; the server's own count; the four accounting states; and keeping the turn. | as-implemented record |
|
||||
| `TECHNICAL-DESIGN.md` §15.4 | **As implemented (v1.1 WP-A2)**: the vocabulary shown in the wire format, genre-neutral placeholders, and extractor rules R1-R4, each anchored to application-owned text. | as-implemented record |
|
||||
| `DECISIONS/013-authoritative-narrative-state-document.md` | An implementation note for v1.1. No event type, field, validation rule or proposal record changed. | as-implemented note |
|
||||
| `V1.1-PLAN.md` | Status: A1 and A2 implemented and staged. | status |
|
||||
| `planning/README.md` | Current state. | index |
|
||||
| `reports/v1.1/V1.1-WP-A1-A2-REPORT.md` | **New.** The combined review package, with A1 and A2 kept separate. | work-package report |
|
||||
| `README.md`, `DEVELOPMENT.md` | The reserve and the accounting. The stale backend test count and the Screenshots paragraph are corrected. | developer docs |
|
||||
|
||||
**Corrective work before commit (owner review, 2026-09-14).** Report §R.
|
||||
|
||||
| Document | Change |
|
||||
| --- | --- |
|
||||
| `TECHNICAL-DESIGN.md` §15.2 | A cold model is loaded once before its turn is built (`contextwindow.ensure_window`). |
|
||||
| `TECHNICAL-DESIGN.md` §15.4 | R5: the echoed continue hint is recognised by its own sentence, and the application-opened tail above it is removed. |
|
||||
| `DEVELOPMENT.md`, `README.md` | The cold-model load, in operator terms. |
|
||||
|
||||
**Requirement changes: zero.**
|
||||
|
||||
## v4.1 — Post-release correction, and the v1.1 plan (2026-09-14)
|
||||
|
||||
Documentation only. No product code, no requirement, no acceptance test and no
|
||||
schema changed.
|
||||
|
||||
**Part 1 — post-release correction.** v4.0 was written before three owner
|
||||
events that have since happened: the closeout commit was signed (`432f041`),
|
||||
`main` was fast-forwarded to it, and the signed tag `v1.0.0` was created on it
|
||||
and pushed. The current-state wording that said otherwise is corrected. The M11
|
||||
report is not edited: its §T records the state at closeout, which was true when
|
||||
written.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `README.md` | Status: v1.0.0 released; tag and `main` at `432f041`; v1.1 on `v1.1-development`. | developer docs |
|
||||
| `planning/README.md` | Current state, the release events, the map (v1.0.0 released, v1.1 planning), the stop rule restated for work packages, `V1.1-PLAN.md` in the document table and reading order. | index |
|
||||
| `planning/BUILD-MILESTONES.md` | Header status, **stale since M8** ("M9 is next"), now says complete and closed. A dated post-release note under M11's status, leaving the closeout paragraph as it stood. The *Post-v1 backlog* points at its triage. | milestone status |
|
||||
| `planning/VERSION.md` | This header and entry. | package version |
|
||||
|
||||
**Part 2 — the v1.1 plan.**
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `planning/V1.1-PLAN.md` | **New.** Every post-v1 backlog item and every other recorded v1 residual risk, triaged. Six v1.1 work packages in order (A1, A2, B, C, D, E), each with objective, rationale, scope, non-scope, affected subsystems, acceptance criteria, regression requirements, dependencies, compatibility and risk. The v1.1 release criteria. A v1.2 list and future work. | planning |
|
||||
|
||||
**Recommended scope: a focused v1.1.** Context-window safety, protocol-echo
|
||||
hardening, independent memory retention, browser coverage, recovery honesty and
|
||||
control-boundary contrast. Scheduled backups, identity-diagnostic extensions and
|
||||
media adapters are deferred, with the reason for each.
|
||||
|
||||
**First brief to write:** WP-A1, the context-window safety reserve.
|
||||
|
||||
**Requirement changes: zero.** Every v1.1 package improves the implementation of
|
||||
an existing requirement, so `SPECIFICATION.md` and `V1-ACCEPTANCE-TESTS.md` are
|
||||
unchanged.
|
||||
|
||||
## v4.0 — M11 closeout: v1 release validation accepted (2026-09-14)
|
||||
|
||||
|
||||
@@ -0,0 +1,881 @@
|
||||
# Adventure Storyteller v1.1 — Integrated Release Validation
|
||||
|
||||
**Status:** COMPLETE — **V1.1 RELEASE VALIDATION: PASS**. The decision, and what
|
||||
it deliberately does not cover, is in §W.
|
||||
|
||||
This report answers one question: **does this exact candidate preserve the
|
||||
complete v1 contract and satisfy every accepted v1.1 package on one integrated
|
||||
release tree?** It is release validation, not a work package. Nothing here adds
|
||||
a feature, and no release tag is created by it.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository / provenance
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **Candidate SHA** | **`87a40326a29533c8d52c9f9f41022e7b499b1de7`** |
|
||||
| Branch | `v1.1-development`, up to date with `origin/v1.1-development` |
|
||||
| Working tree at freeze | **clean** — nothing modified, nothing staged |
|
||||
| Commit | *v1.1: harden recovery and control boundaries* (WP-D + WP-E) |
|
||||
| **Owner signature** | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, made 2026-09-16 05:37:13 EDT |
|
||||
| Tag at HEAD | **none** — no `v1.1.0` tag exists |
|
||||
| v1.0.0 baseline | `432f04100b9a67198bcdc46c6ff8ee0f181e1667`, **an ancestor** |
|
||||
| Package ancestry | `d63804f` (WP-A1/A2), `beb17ad` (WP-B.1), `0c1ba83` (WP-B.2), `59b5ebc` (WP-C) — **all ancestors** |
|
||||
| Diff v1.0.0..HEAD | 61 files, +14,833 / −366 |
|
||||
| LICENSE / PROVENANCE | **unchanged since v1.0.0** (empty diff) |
|
||||
|
||||
### A.1 Frozen candidate identity
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Dependency locks | `backend/requirements.txt` `sha256:ed28bc0f8970cf4e…`, `frontend/package-lock.json` `sha256:355cb370837ade01…`, `frontend/package.json` `sha256:2016580ddfa176a9…`, `backend/requirements-dev.txt` `sha256:06d7695816b201e9…` |
|
||||
| Schema | `LATEST_VERSION` **94**, 93 migrations (`PRAGMA user_version`) |
|
||||
| Bundle format | **`ai-dnd-adventure-v3`** |
|
||||
| Import ceiling | 20 MB (`MAX_IMPORT_BODY_BYTES`), unchanged |
|
||||
| Frontend build | `dist` built 2026-09-16T05:39:55, 16 files, `sha256(dist) = ea2753ad24f61959fe084f4674911acc` |
|
||||
| Docker image | `storyteller:release-87a4032`, `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398`, 312 MB |
|
||||
| Firefox / geckodriver | 155.0.1 / 0.37.1 (2026-09-04) |
|
||||
| Docker | 29.8.0, build 88096ef |
|
||||
| CPU HTTPS reference host | Ollama **0.33.0**, serving `qwen2.5:3b-instruct`, `qwen2.5:3b-instruct-16k`, `nomic-embed-text:latest`; certificate verifies through the machine's CA store with no bypass |
|
||||
| GPU inference host | Ollama **0.34.0**; `qwen2.5:3b-instruct-16k` digest `21ff8cc52f375f19`, `nomic-embed-text:latest` digest `0a109f422b47e3a3`; no model resident at start |
|
||||
| Evidence root | `$HOME/v11-evidence/release-87a4032/` — never `/tmp`, and no real hostname appears in any committed file |
|
||||
|
||||
---
|
||||
|
||||
## B. Package acceptance inventory
|
||||
|
||||
| Package | Status | Source |
|
||||
| --- | --- | --- |
|
||||
| **WP-A1** context-window safety reserve | **ACCEPTED** | `V1.1-WP-A1-A2-REPORT.md`, signed `d63804f` |
|
||||
| **WP-A2** protocol-echo cleanup, genre-neutral prompting | **ACCEPTED** | same report and commit |
|
||||
| **WP-B** independent memory | **ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION** | `V1.1-WP-B1-REPORT.md`, `V1.1-WP-B2-REPORT.md` §S, signed `beb17ad` / `0c1ba83` |
|
||||
| **WP-C** browser release coverage | **ACCEPTED** | `V1.1-WP-C-REPORT.md`, signed `59b5ebc` |
|
||||
| **WP-D** recovery honesty | **ACCEPTED** | `V1.1-WP-D-REPORT.md`, signed `87a4032` |
|
||||
| **WP-E** control-boundary contrast | **ACCEPTED** | `V1.1-WP-E-REPORT.md`, signed `87a4032` |
|
||||
|
||||
### B.1 WP-B's qualification, carried whole
|
||||
|
||||
The WP-B disposition is **not** shortened to "WP-B passed" anywhere in this
|
||||
report. Its own §S records:
|
||||
|
||||
```text
|
||||
B2.1 RANKING: PASS
|
||||
B2.2 EVICTION: PASS
|
||||
B2.3 EXCERPT CREATION: PASS
|
||||
B2.4 SUMMARIZER-PROMPT EXPERIMENT: FAIL — REVERTED
|
||||
|
||||
DETERMINISTIC WP-B: PASS
|
||||
REAL-MODEL WP-B: FAIL
|
||||
|
||||
WP-B OVERALL:
|
||||
ACCEPTED WITH DOCUMENTED REAL-MODEL LIMITATION
|
||||
```
|
||||
|
||||
In the release contract's own words (§11 items 12), that is:
|
||||
|
||||
```text
|
||||
DETERMINISTIC WP-B: PASS
|
||||
REFERENCE-MODEL INDEPENDENT MEMORY: FAIL
|
||||
OWNER ACCEPTED THE LIMITATION FOR v1.1
|
||||
```
|
||||
|
||||
The failing stage is **memory creation** — the summariser's content selection —
|
||||
not ranking, eviction or injection, each of which passes deterministically.
|
||||
|
||||
### B.2 WP-E screenshot approval — a correction of record
|
||||
|
||||
The committed WP-E report read `OWNER SCREENSHOT APPROVAL: PENDING`, because it
|
||||
was written before the owner reviewed the images. The owner's release-validation
|
||||
brief (2026-09-16) states the before/after screenshots are approved and
|
||||
instructs this validation to record it. The WP-E report is updated to
|
||||
`APPROVED` as part of this closeout (§V), sourced to that brief and dated. No
|
||||
visual code changed during release validation, so the approval stands (§Q).
|
||||
|
||||
---
|
||||
|
||||
## C. v1 acceptance matrix
|
||||
|
||||
Every test marked **REQUIRED FOR V1** — there are **82** — against evidence taken
|
||||
on **this candidate**. Evidence types follow M11's: `browser` (the 101-check run,
|
||||
§G), `campaign` (the 102-turn integrated run, §H), `container` (the offline run
|
||||
on the candidate image, §F), `process` (spawned server processes — recovery §M,
|
||||
upgrade §N), `suite` (the 1,723-test backend suite, §D). No historical result
|
||||
from different product code is used where the contract asks for candidate
|
||||
evidence.
|
||||
|
||||
**Result: 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified.**
|
||||
|
||||
### A — Local-first operation
|
||||
|
||||
| ID | Result | Evidence on this candidate |
|
||||
| --- | --- | --- |
|
||||
| A01 Start application offline | **PASS** | container: first page load, fresh volume, no route and no DNS |
|
||||
| A02 Storyteller loopback default | **PASS** | suite; every harness reached it on `127.0.0.1`; `docker-compose.yml` publishes `127.0.0.1:8000:8000` |
|
||||
| A03 No cloud API key | **PASS** | suite; container: no secret in an export |
|
||||
| A04 Campaign survives restart | **PASS** | campaign: **3 process restarts, 4 process starts**, state compared across each; container: campaigns survive a container restart |
|
||||
| A05 Failed model call does not corrupt story | **PASS** | campaign: a real `failed_call` at turn 69 against an unserved model, play resumed; container: same with no model reachable |
|
||||
| A06 Trusted-LAN Ollama inference | **PASS** | browser: the whole 101-check run over **trusted-LAN HTTPS with a private CA**, verification on, no bypass, storyteller loopback-bound |
|
||||
|
||||
### B — Core play
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| B01 Natural language action | **PASS** | browser (real turns through the UI) + campaign (102 accepted) |
|
||||
| B02 Dialogue input | **PASS** | campaign: dialogue beats in the turn list |
|
||||
| B03 Continue | **PASS** | suite; browser: the Continue control present and enabled |
|
||||
|
||||
### C — Story authority and state
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| C01 Campaign canon is preserved | **PASS** | campaign: canon present in the prompt on **102 of 102** turns |
|
||||
| C02 Possession state | **PASS** | campaign (the silver key) + suite |
|
||||
| C03 Character knowledge is not invented | **PASS** | suite |
|
||||
| C04 Manual state correction | **PASS** | campaign: **2 state corrections**; browser: C3's accepted and refused corrections; suite |
|
||||
| C05 Canon beats reference | **PASS** | suite |
|
||||
| C06 Structured state matches accepted narrative consequence | **PASS** | campaign: real extraction across 102 turns, every event validated or refused; suite |
|
||||
|
||||
### D — Non-destructive history
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| D01 Undo one turn | **PASS** | browser + campaign (`undo`) |
|
||||
| D02 Minimum five undos | **PASS** | suite; campaign (`undo_redo`) |
|
||||
| D04 Redo | **PASS** | browser + campaign |
|
||||
| D05 Redo invalidated by new continuation | **PASS** | campaign: `diverged`, after which Redo is gone |
|
||||
| D06 Retry narrator response | **PASS** | campaign: **2 retries** |
|
||||
| D07 Select prior retry take | **PASS** | campaign: `take_selected` |
|
||||
| D08 Retry does not delete prior take | **PASS** | campaign + suite |
|
||||
| D09 Edit earlier user input | **PASS** | suite |
|
||||
| D10 Edit narrator output | **PASS** | suite; browser (hostile-Markdown plants through the narrator-edit path) |
|
||||
| D11 Named checkpoint | **PASS** | campaign: **2 Save Points**; browser; suite |
|
||||
| D12 Restore checkpoint | **PASS** | campaign: `save_point_restored`; browser |
|
||||
| D13 Restore does not delete later history | **PASS** | campaign: retained actions after the restore; recovery §M |
|
||||
| D14 Delete checkpoint | **PASS** | suite; browser: the delete confirmation dialog |
|
||||
|
||||
### E — Branch and derived-data isolation
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| E01 Abandoned future cannot affect active state | **PASS** | suite `test_m11_leakage.py`, with a positive control |
|
||||
| E02 Abandoned memory cannot leak | **PASS** | as above |
|
||||
| E03 Abandoned summary cannot leak | **PASS** | as above |
|
||||
| E04 Scene state is lineage-safe | **PASS** | as above |
|
||||
|
||||
### F — Long-term memory and context
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| F01 Recent turns remain coherent | **PASS** | campaign: history populated every turn, newest always included |
|
||||
| F02 Old important event retrieval | **PASS** | campaign §K: the planting turn outside the window and the fact recovered — **through authoritative state**, not independent memory (§K states which) |
|
||||
| F03 Prompt remains bounded | **PASS** | campaign: 1,602–14,982 tokens against a 16,384 budget across 102 turns |
|
||||
| F04 Output token reserve | **PASS** | campaign: `output_reserve` 500 present and subtracted on every turn |
|
||||
| F05 Prompt inspector | **PASS** | browser: the context panel shows the assembled prompt |
|
||||
| F06 Retrieval provenance | **PASS** | campaign: knowledge and memory provenance per turn; suite |
|
||||
| F07 Heuristic memory is not canon | **PASS** | suite |
|
||||
| F08 Memory failure is non-fatal | **PASS** | suite; container: derived work fails with no model and turns still commit; campaign: **0 post-turn failures**, no database-lock errors |
|
||||
|
||||
### G — Imported knowledge
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| G01 Import local text | **PASS** | browser: through the real file input; container: offline |
|
||||
| G02 Import local Markdown | **PASS** | campaign: **3 sources imported**; browser |
|
||||
| G03 Classification | **PASS** | campaign: all three classes; recovery §M confirms them after a move |
|
||||
| G04 Disable knowledge source | **PASS** | suite |
|
||||
| G05 Canon retrieval | **PASS** | campaign: canon passages in stored prompts; suite |
|
||||
| G06 Reference retrieval | **PASS** | suite |
|
||||
| G07 Inspiration is low authority | **PASS** | suite |
|
||||
| G08 No automatic URL fetch | **PASS** | suite `test_egress.py`; container: no network at all, import still works |
|
||||
| G09 Remote Markdown image does not auto-load | **PASS** | browser: no remote image src in the rendered story |
|
||||
| G10 Prompt injection in source is treated as data | **PASS** | suite; browser: injection text rendered as text |
|
||||
|
||||
### H — Security
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| H01 No unexpected outbound connections | **PASS** | container (no network at all) + suite `test_egress.py` |
|
||||
| H02 No telemetry | **PASS** | suite |
|
||||
| H03 No cloud provider required | **PASS** | container: a full campaign offline |
|
||||
| H04 Model output cannot execute shell | **PASS** | browser + suite |
|
||||
| H05 Invalid state event rejected | **PASS** | suite; campaign: refusals recorded |
|
||||
| H06 Stored XSS protection | **PASS** | browser: `onerror` and `<script>` in accepted narration, neither executed |
|
||||
| H07 JavaScript URL protection | **PASS** | browser: no `javascript:` href in the DOM |
|
||||
| H08 Path traversal import rejected | **PASS** | suite |
|
||||
| **H09 ZIP Slip protection** | **NOT APPLICABLE** | the product extracts no archives, and a test enforces it — the same condition v1 recorded |
|
||||
| H10 Restrictive CORS and local API behaviour | **PASS** | browser: an unknown API path is a 404 with a non-HTML body; suite |
|
||||
| H11 No first-use runtime asset download | **PASS** | container: every asset local with no network; suite: the tokenizer table is vendored |
|
||||
| H12 Inference endpoint enforcement | **PASS** | suite; smoke §R: a public endpoint refused by the running image; the container's refusal of an unresolvable name is the same policy (§R) |
|
||||
|
||||
### I — Export, import and recovery
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| I01 Export campaign | **PASS** | campaign: the 102-turn campaign exported (3,071,683 bytes) |
|
||||
| I02 Import exported campaign | **PASS** | process §M: imported into a database and directory that never existed |
|
||||
| I03 Branch/disposable history export | **PASS** | §M: **88 actions retained beyond the active line** after the move |
|
||||
| I04 Checkpoint export | **PASS** | §M: both Save Points restore after the move |
|
||||
| I05 Knowledge provenance export | **PASS** | §M: all three classes with their content |
|
||||
| I06 Database/export contains no API secrets | **PASS** | §M: the bundle carries no secret; suite; container |
|
||||
| I07 Export/import preserves an undone active head | **PASS** | §N: the v1.0.0 campaign's undone head and `can_redo` survive upgrade and both bundle directions; suite `test_m9_portability.py`. *(The long-run bundle ended head-at-tip, so §M exercises the other case — stated in §M rather than implied)* |
|
||||
|
||||
### J — Genre neutrality
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| J01 Science-fiction campaign | **PASS** | suite `test_m11_scifi.py` |
|
||||
| J02 Generic entity support | **PASS** | as above |
|
||||
| J03 Genre profiles are configuration | **PASS** | as above |
|
||||
|
||||
### K — Future media architecture
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| K01 Scene snapshot exists | **PASS** | suite; container: a scene packet builds offline |
|
||||
| K02 Visual character profile | **PASS** | suite |
|
||||
| K03 Visual location profile | **PASS** | suite |
|
||||
|
||||
### L — Data integrity
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| L01 Atomic turn commit | **PASS** | container + campaign: a real induced failure, no narration accepted, no half-written state |
|
||||
| L02 State reconstruction | **PASS** | suite; campaign: state compared across 3 restarts |
|
||||
| L03 Checkpoint reconstruction after restart | **PASS** | campaign + §M |
|
||||
|
||||
### M — Long-run
|
||||
|
||||
| ID | Result | Evidence |
|
||||
| --- | --- | --- |
|
||||
| **M01** 100-turn campaign | **PASS** | **102 accepted turns**, every scheduled operation exercised, **0 post-turn failures** (§H) |
|
||||
| **M02** Restart during long campaign | **PASS** | **3 genuine process restarts** (4 process starts); everything crossed as bytes on disk |
|
||||
| **M03** Long-run context stability | **PASS** | the prompt held **13,492–14,982** tokens over the last 70 turns against a 16,384 budget; window verified **102/102**; canon present **102/102** |
|
||||
| **M04** Long-run memory recall | **PASS**, qualified | the planted clue was outside the history window (planted depth 1, floor 72) and reached the prompt: `m04_verdict: **recovered_through_state_only**`. Recovery was **through authoritative state**, not independent memory — §K states this distinction and does not relabel it |
|
||||
|
||||
### SHOULD and FUTURE
|
||||
|
||||
Not counted as REQUIRED. **SHOULD:** B04, D03, K04, L04 — all still pass on the
|
||||
candidate (suite; browser for B04's direction toggle). **FUTURE:** K05, K06 —
|
||||
deliberately not run; both need a media provider this release does not build.
|
||||
|
||||
## D. Backend / frontend suites
|
||||
|
||||
| Suite | Result |
|
||||
| --- | --- |
|
||||
| **Backend** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,210.6 s) |
|
||||
| **Frontend** (`npm test`) | **175 passed**, 15 files, 0 failed |
|
||||
| **Lint** (`npm run lint`, oxlint) | **exit 0 — 0 errors**, 15 warnings |
|
||||
| **Production build** (`npm run build`) | succeeded |
|
||||
|
||||
**Every skip explained — one category, and it is the expected one.** All 17 are
|
||||
environment-gated real-model tests, skipped because `AIDND_TEST_*` is
|
||||
deliberately unset for the deterministic suite:
|
||||
|
||||
| File | Skipped | Gate |
|
||||
| --- | --- | --- |
|
||||
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
|
||||
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` + `AIDND_TEST_MODEL` |
|
||||
| `test_narrative_realistic.py` | 3 | same |
|
||||
| `test_m11_real_window.py` | 3 | same; one needs `AIDND_TEST_WIDE_MODEL` |
|
||||
| `test_provider_wiring.py` | 1 | same |
|
||||
|
||||
**0 xfailed.** No deficiency WP-B fixed remains parked as an expected failure —
|
||||
B.1's two strict xfails became ordinary passes in B.2 and stayed that way.
|
||||
|
||||
**Lint warnings are the documented, unchanged set:** the pre-existing
|
||||
`only-export-components` and unused-import warnings recorded at WP-C, WP-D and
|
||||
WP-E. None is in a file this candidate changed relative to those packages.
|
||||
|
||||
## E. Production build and Docker image
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Command | the repository's documented production build (`DEVELOPMENT.md`), with `--no-cache` |
|
||||
| Image | `sha256:fe10e8a2511395958882e489bfdf53fa00d981597cc81aa5587235f7b7836398` (312 MB) |
|
||||
| Dependencies | **installed, not reused** — `npm ci` and `pip install --no-cache-dir` both executed in the log |
|
||||
| `CACHED` steps | **2**, and both are `WORKDIR` metadata (`/build`, `/app`) — no dependency or source layer was cached |
|
||||
| **Image SPA vs local build** | **file-for-file identical**: 16 files each, `diff -r` clean, combined `sha256 = ea2753ad24f61959fe084f4674911acc` on both sides |
|
||||
|
||||
The image is built from the candidate tree for this validation. No earlier
|
||||
work-package image was reused.
|
||||
|
||||
## F. Offline / no-network
|
||||
|
||||
`tools/m11_offline.py` against **the candidate image**, `--network none`, fresh
|
||||
volume. Evidence: `…/release-87a4032/offline/offline-report.json`.
|
||||
|
||||
**23 checks, 23 passed, 0 failed.** Including: no route to the public Internet;
|
||||
no external DNS; first page load; no remote origin named; CSP served; every
|
||||
referenced asset local; campaign creation; state extraction; local file import;
|
||||
prompt assembly; knowledge search; a turn with no model reachable reported as a
|
||||
failure with no narration accepted, the player's words kept and state unchanged;
|
||||
export; import with state; no secret in the export; the media module inert with
|
||||
no provider; campaigns surviving a container restart.
|
||||
|
||||
**No first-use download occurred**, which is what the container's absent network
|
||||
makes unfalsifiable rather than merely unobserved.
|
||||
|
||||
## G. Browser release validation
|
||||
|
||||
`tools/m11_browser.py` on the candidate, production build served by FastAPI,
|
||||
Firefox 155.0.1 / geckodriver 0.37.1, narrator over **trusted-LAN HTTPS with the
|
||||
private CA** (`endpoint_class: trusted-LAN HTTPS`), storyteller on loopback.
|
||||
`kind: release regression` — not partial, not `--only`. 693 s.
|
||||
|
||||
| Suite | Passed | Failed | Skipped |
|
||||
| --- | --- | --- | --- |
|
||||
| **M11** (the v1 release regression) | **38** | **0** | **0** |
|
||||
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
|
||||
| **WP-E** (control boundaries) | **10** | **0** | **0** |
|
||||
| **Total** | **101** | **0** | **0** |
|
||||
|
||||
**A1 accounting:** 8 narrator turns, **`fits` on every one**; `turns_not_clean`
|
||||
**0**; `protocol_shapes_in_narration` **0**. The verified window on this host was
|
||||
4,096 (`source: loaded`); the 16,384 evidence is the long run's (§I).
|
||||
|
||||
WP-C's proofs are all present in the 53: Retry and alternate takes, Save Point
|
||||
create/restore/Redo, state correction including a refusal shown as a refusal,
|
||||
narration length reaching the prompt, failed generation and recovery, and **real
|
||||
export downloads** from both the library and campaign settings, each file landing
|
||||
on disk and importing into a fresh application. WP-E's ten are the rendered
|
||||
boundary measurements of §Q.
|
||||
|
||||
## H. Integrated 100+ turn long run
|
||||
|
||||
One new campaign on the candidate's product code, GPU inference host, with the
|
||||
owner's power/link/kernel logging running before the first turn.
|
||||
Evidence: `…/release-87a4032/long-run/`.
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Narrator | **`qwen2.5:3b-instruct-16k`**, digest `21ff8cc52f375f19` |
|
||||
| Embeddings | **`nomic-embed-text`**, digest `0a109f422b47e3a3` |
|
||||
| Window | **16,384**, `window_verified` on **102 of 102** turns |
|
||||
| Memory / summaries | **on** — 19 memories in the bank, **12 summaries** written |
|
||||
| **Accepted turns** | **102** (target 100) |
|
||||
| Restarts | **3** genuine process restarts, 4 process starts |
|
||||
| Elapsed | 1,330 s |
|
||||
| Export | 3,071,683 bytes; database 2,613,248 bytes |
|
||||
| Status | `complete`; `aborted_reason` null, `failed_reason` null |
|
||||
|
||||
**Not a repeated-turn benchmark.** Every scheduled operation fired and is in the
|
||||
timeline: 3 restarts, 1 undo, 1 undo→redo, 2 retries, 1 take selection, 2 Save
|
||||
Points, 1 Save Point restore, 1 divergence, 2 state corrections, 3 knowledge
|
||||
imports, memory activation, 1 deliberate failed call, 1 export, the planted clue
|
||||
and the planted independent fact, and the recall probe.
|
||||
|
||||
### H.1 M01–M04
|
||||
|
||||
| | Verdict | What decides it |
|
||||
| --- | --- | --- |
|
||||
| **M01** 100-turn campaign | **PASS** | 102 accepted turns, each with a committed action and state document |
|
||||
| **M02** Restart during long campaign | **PASS** | 3 genuine `uvicorn` restarts; everything that survived crossed as bytes on disk |
|
||||
| **M03** Long-run context stability | **PASS** | prompt 13,492–14,982 tokens over the last 70 turns against a 16,384 budget; canon present 102/102; window verified 102/102 |
|
||||
| **M04** Long-run memory recall | **PASS**, and qualified | the clue was planted at depth 1, the history floor reached depth 72, and it was **not** in the recent window; it reached the prompt through **authoritative state**. `m04_verdict: recovered_through_state_only`. §K keeps the distinction the criterion was written around |
|
||||
|
||||
### H.2 State and derived-work integrity
|
||||
|
||||
**0 post-turn failures across 102 turns, and no `database is locked` error.** The
|
||||
only failure-shaped events in the whole timeline are the two the run creates on
|
||||
purpose: the scheduled `failed_call` at turn 69 (a model name the server does not
|
||||
serve — A05/L01 evidence) and the independent-fact precondition notes (§K).
|
||||
Narrative-state proposals were recorded, applied or refused as designed across
|
||||
the run, and no accepted narration carried an unresolved protocol block (§J).
|
||||
|
||||
## I. Context-window / A1 evidence
|
||||
|
||||
**Every turn `fits`.** Across all 102 accepted turns the accounting status was
|
||||
`fits` — **0 `exceeded`, 0 `truncation_suspected`** — and the window was verified
|
||||
at 16,384 on every one.
|
||||
|
||||
**The ten largest stored prompts**, re-counted against what the server itself
|
||||
reported:
|
||||
|
||||
| Turn | App estimate | Server count | Difference | Window | Output reserve | Safety reserve | Observed margin | Status |
|
||||
| ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | :--- |
|
||||
| 86 | 14,990 | 15,005 | +15 | 16,384 | 500 | 820 | 879 | fits |
|
||||
| 83 | 14,978 | 14,993 | +15 | 16,384 | 500 | 820 | 891 | fits |
|
||||
| 97 | 14,965 | 14,980 | +15 | 16,384 | 500 | 820 | 904 | fits |
|
||||
| 41 | 14,960 | 14,975 | +15 | 16,384 | 500 | 820 | 909 | fits |
|
||||
| 66 | 14,951 | 14,966 | +15 | 16,384 | 500 | 820 | 918 | fits |
|
||||
| 33 | 14,941 | 14,956 | +15 | 16,384 | 500 | 820 | 928 | fits |
|
||||
| 35 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
|
||||
| 77 | 14,940 | 14,955 | +15 | 16,384 | 500 | 820 | 929 | fits |
|
||||
| 38 | 14,927 | 14,942 | +15 | 16,384 | 500 | 820 | 942 | fits |
|
||||
| 45 | 14,922 | 14,937 | +15 | 16,384 | 500 | 820 | 947 | fits |
|
||||
|
||||
**The reserve is preserved on every re-counted prompt.** The estimate runs
|
||||
exactly **15 tokens below** the server's own count on all ten — a constant,
|
||||
known offset rather than drift — and the smallest observed margin anywhere in the
|
||||
run is **879 tokens**, against the documented safety reserve of 820.
|
||||
|
||||
**Against v1.** M11's closeout recorded a remaining margin of **23–42 tokens**.
|
||||
The same measurement on this candidate is **879 at its tightest** — roughly
|
||||
twenty to thirty times the headroom, which is what WP-A1 was for.
|
||||
|
||||
## J. Protocol-leak / A2 evidence
|
||||
|
||||
**Release-gate result: 0.** Across the run's **105 stored AI actions**,
|
||||
`protocol_leaks` reports **0 leaking**, `example_ids` empty.
|
||||
|
||||
Measured separately, by the detector's own four rules:
|
||||
|
||||
| Shape | Count in this run |
|
||||
| --- | --- |
|
||||
| State/section heading with an indented entry | **0** |
|
||||
| Event list (`"events"`) | **0** |
|
||||
| Event-call syntax (`create_entity(`, `set_scene(`, …) | **0** |
|
||||
| Hard-limit / continue-hint echo | **0** |
|
||||
|
||||
The browser run agrees independently: `protocol_shapes_in_narration` **0** across
|
||||
its narrated turns (§G).
|
||||
|
||||
**The known mid-reply echo did not recur — and is still not fixed.** WP-B.1
|
||||
recorded one stored reply (action 153, depth 143) where the narrator echoed the
|
||||
length hint mid-reply and then continued the story, which A2's trailing cleanup
|
||||
does not remove. Replaying **that stored fixture through this candidate's
|
||||
detector** still flags it — 1 of 105 AI actions, matched by the hard-limit/hint
|
||||
rule alone. So the residual is live (§T.2); what this release run shows is that
|
||||
no equivalent shape occurred in **its** 105 replies. This report does not claim
|
||||
protocol leakage is solved.
|
||||
|
||||
Ordinary fact and state restatement in prose was not counted: it is measurement,
|
||||
not application-owned protocol, and no rule treats it as a leak.
|
||||
|
||||
## K. Memory / WP-B evidence
|
||||
|
||||
### K.1 At the final recall point
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **Created** | 19 memories in the bank; 12 summaries |
|
||||
| **Retained** | the planting turn's era survived to the end of a 102-turn run |
|
||||
| **Ranked** | 4 memories were selected into the prompt at the recall point |
|
||||
| **Injected** | the memory section reached the prompt (14,395 tokens that turn) |
|
||||
| **Independent-memory verdict** | **not demonstrated** — `precondition_failed: absent_from_later_narration` |
|
||||
|
||||
### K.2 The independent-fact probe, precondition by precondition
|
||||
|
||||
| Precondition | Held? |
|
||||
| --- | --- |
|
||||
| The planted turn is outside the history window (planted depth 3, floor 72) | **yes** |
|
||||
| Absent from authoritative state | **yes** |
|
||||
| Absent from the summary | **yes** |
|
||||
| Absent from imported knowledge | **yes** |
|
||||
| Absent from later narration | **no** — 6 violations, the first at turn 6 |
|
||||
|
||||
The narrator restated the planted fact in later narration, so the probe could not
|
||||
isolate memory as the only path. The run therefore **records no independent
|
||||
recovery**, and nothing here is relabelled as one.
|
||||
|
||||
### K.3 The two statements the contract requires, kept apart
|
||||
|
||||
```text
|
||||
deterministic independent-memory recovery: PASS
|
||||
reference-model independent-memory limitation: ACCEPTED RESIDUAL
|
||||
```
|
||||
|
||||
- **Deterministic** (`WP-B.2` §I, and the suite on this candidate): the
|
||||
`independent_full` scenario fails on v1.0.0 at creation and returns
|
||||
`recovered_through_memory_independent` on the candidate, with isolation
|
||||
asserted every turn and provenance resolving to the planting turn.
|
||||
- **Reference model:** failed on the precondition-valid attempt in WP-B.2, and
|
||||
in this release run the attempt was not precondition-valid at all. The failing
|
||||
stage remains **memory creation** — the summariser's content selection.
|
||||
- **No new regression.** B2.1 ranking, B2.2 eviction and B2.3 excerpt creation
|
||||
all pass deterministically in the 1,723-test suite on this candidate, and the
|
||||
bank behaved normally through the run (19 memories, ranked and injected). What
|
||||
this run shows is the known summariser-quality limitation, not a fault in
|
||||
ranking, eviction or injection.
|
||||
|
||||
## L. Identity diagnostic
|
||||
|
||||
**Scripted half — complete.** `tools/m11_identity.py --scripted` on the
|
||||
candidate: 10 identity stresses, **0 signals**, verdict *no objective identity
|
||||
defect detected*.
|
||||
|
||||
**The detector's negative control fires.** With `--inject`, the same harness on
|
||||
the same candidate raises **7 signals** — `shared_display_name` once and
|
||||
`duplicate_character_creation` on turns 5–10 — and preserves each turn's
|
||||
evidence. A clean run therefore means something: the check is capable of
|
||||
failing.
|
||||
|
||||
**Model-backed half — complete, on the candidate, with memory on.**
|
||||
Evidence: `…/release-87a4032/identity-memory/`.
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Model | **`qwen2.5:3b-instruct-16k`** (the reference narrator) |
|
||||
| Context window | **16,384** |
|
||||
| Memory status | **on** — `memoryBankEnabled: true`, embedding model `nomic-embed-text` |
|
||||
| Summary status | **on** — `autoSummarize: true`; **1 summary written** |
|
||||
| **Identity signals** | **0** across all 10 stresses |
|
||||
| **Stored protocol shapes** | **0** of 10 AI actions, by the release-gate detector |
|
||||
| Fixture | accepted successfully — five entities kept distinct (`bill`, `alice`, `roger`, `john`, `office`) |
|
||||
| State proposals | recorded and applied with no shared display name and no duplicate creation |
|
||||
| Scripted detector self-test | still fires (7 signals under `--inject`) |
|
||||
|
||||
**A harness correction made during validation, and what it does not invalidate.**
|
||||
The diagnostic shipped with `embedding_model=""` and the memory bank switched
|
||||
off, so a release run of it would have reported a clean identity result with
|
||||
memory never taking part — which is not what the gate asks for. Two lines now
|
||||
read `AIDND_TEST_EMBED_MODEL` and enable the bank. **Harness-only: no product
|
||||
code changed**, so no product evidence became stale; the corrected harness
|
||||
repeated its own check, which is the run reported above.
|
||||
|
||||
**One honest observation:** with memory enabled the bank still wrote **0
|
||||
memories** in this 10-beat campaign — 21 actions is enough to pass
|
||||
`MEMORY_START`, but the summary pass is what ran and produced the single
|
||||
summary. Memory was configured and active; it was not meaningfully *exercised*
|
||||
here. The bank's real exercise is the 102-turn run (§K), which wrote 19.
|
||||
|
||||
**No claim about the historical root cause.** The post-M8 identity finding's
|
||||
campaign was destroyed and its cause cannot be established. A clean run here is
|
||||
evidence that the product does not do the things it can be blamed for on this
|
||||
fixture — not a discovery of what happened then.
|
||||
|
||||
## M. Recovery
|
||||
|
||||
`tools/m11_recovery.py` against **the integrated run's own bundle**
|
||||
(3,071,683 bytes), imported into a database file that never existed, in a
|
||||
directory that never existed, by a second server process — so migrations ran
|
||||
from nothing and this is the fresh-install path as well as the import path.
|
||||
|
||||
**16 checks, 16 passed, 0 failed.**
|
||||
|
||||
| Claim | Result |
|
||||
| --- | --- |
|
||||
| The destination database did not exist beforehand | PASS |
|
||||
| The bundle imports into a clean directory | PASS — 209 actions in the file, 121 on the active line |
|
||||
| The active transcript is not empty | PASS |
|
||||
| Authoritative state came across | PASS — 6 entities, 1 fact |
|
||||
| The campaign's own canon came across | PASS |
|
||||
| The narration-length choice came across | PASS |
|
||||
| Redo availability matches what the file said | PASS |
|
||||
| Retained (undone) history came across | PASS — **88 actions retained beyond the active line** |
|
||||
| Both Save Points restore | PASS — *On the ridge*, *Before the ridge* |
|
||||
| Every imported class came across, with content | PASS — 3 sources |
|
||||
| The moved campaign accepts a new change, unrefused | PASS |
|
||||
| The bundle carries no secret | PASS |
|
||||
| The moved campaign exports again, same story length | PASS |
|
||||
|
||||
**One thing this does not prove, stated rather than implied.** The long run
|
||||
ended with its head at the tip, so the file's head was at the tip and Redo was
|
||||
correctly unavailable after import. The **undone-head** case (I07) is proved by
|
||||
§N's upgrade campaign, which ends on an Undo with Redo available and survives
|
||||
both bundle directions, and by `test_m9_portability.py` in the suite — not by
|
||||
this bundle.
|
||||
|
||||
## N. v1.0.0 upgrade compatibility
|
||||
|
||||
A campaign **built and played by the `432f041` application** in its own
|
||||
worktree, then opened by the candidate. Evidence:
|
||||
`…/release-87a4032/upgrade/upgrade-report.json`.
|
||||
|
||||
**Phase 1 — v1.0.0 builds it.** 20 turns accepted, 39 actions, canon knowledge
|
||||
imported, a Save Point taken, one Undo so the head is not at the tip, narration
|
||||
length chosen, memory bank and auto-summary on. It contains what §11 item 9
|
||||
names: retained history, an undone head with Redo available, a Save Point,
|
||||
**6 memories**, **2 summaries**, imported knowledge, a narration-length choice
|
||||
and narrative state. Settings hold a **loopback placeholder**
|
||||
(`http://127.0.0.1:11434/v1`) — no real hostname is in the evidence database.
|
||||
|
||||
**Phase 2 — the candidate opens the same file.** Nothing was copied; the
|
||||
candidate's migrations ran against it.
|
||||
|
||||
| Field | Before | After |
|
||||
| --- | --- | --- |
|
||||
| transcript | 39 actions | **identical** |
|
||||
| newest action (the head) | — | **identical** |
|
||||
| total / can_undo / can_redo | 39 / true / **true** | **identical** |
|
||||
| checkpoints | 1 | **identical** |
|
||||
| narrative state | 7 keys | **identical** |
|
||||
| memories | **6** | **identical** |
|
||||
| summaries | **2** | **identical** |
|
||||
| knowledge sources | 1 | **identical** |
|
||||
| narration length | `brief` | **identical** |
|
||||
| memory_bank_enabled / auto_summarize | true / true | **identical** |
|
||||
| settings | loopback placeholder | **identical** |
|
||||
| **schema `user_version`** | **94** | **94** |
|
||||
|
||||
**All 15 census fields compared identical; 0 differ.** Schema parity is exact:
|
||||
v1.0.0 and the candidate both stamp `user_version` 94 with 93 migrations, so the
|
||||
upgrade required no migration at all, and nothing was rewritten in passing. A
|
||||
fresh-install database from the candidate carries the same 94 (§M's import ran
|
||||
migrations from nothing).
|
||||
|
||||
**A first pass at this gate was discarded.** It played 5 turns, which is below
|
||||
the memory and summary thresholds, so it compared **0 memories against 0
|
||||
memories** and proved nothing about two of the criterion's required contents;
|
||||
its settings also carried the live hostname rather than a placeholder. Both were
|
||||
corrected and the gate was rerun — the run reported above.
|
||||
|
||||
## O. Bundle compatibility
|
||||
|
||||
| Direction | Result |
|
||||
| --- | --- |
|
||||
| **v1.0.0 export → imported by v1.1** | **PASS** — imported, id 2 |
|
||||
| **v1.1 export → offered to v1.0.0** | **ACCEPTED** — v1.0.0 imported it, id 1 |
|
||||
|
||||
Both directions were **executed**, not inferred from the unchanged format
|
||||
string. The format is `ai-dnd-adventure-v3` on both sides, and it did not change
|
||||
during release validation.
|
||||
|
||||
Backward import succeeding means the compatibility question the brief raised —
|
||||
whether optional or additive v1.1 evidence data would break a v1.0.0 importer —
|
||||
is answered in the negative for this campaign's contents: v1.0.0 accepted the
|
||||
candidate's bundle whole. No bundle-format change was made or needed.
|
||||
|
||||
## P. WP-D regression
|
||||
|
||||
Reconfirmed on the candidate; no backup-affecting product code changed after
|
||||
WP-D, so its 117 MB browser measurement is not repeated (the owner's brief
|
||||
permits this).
|
||||
|
||||
| Claim | Result |
|
||||
| --- | --- |
|
||||
| The completed backup copy is verified with `PRAGMA integrity_check` | **PASS** — `app/backup.py:193`, docstring at 198–201 records why the full check replaced `quick_check` |
|
||||
| The corruption fixture still separates the two pragmas | **PASS** — `quick_check` → `ok`, `integrity_check` → `row 145 missing from index i_t_k` |
|
||||
| Oversized export still succeeds and is delivered | **PASS** |
|
||||
| Importability metadata names the effective ceiling | **PASS** — "This export is larger than this version's **20 MB** import limit (… bytes). The file was exported successfully, but this version cannot import it." |
|
||||
| Normal export unchanged in content | **PASS** |
|
||||
| Oversized import still refused | **PASS** — 413 naming the limit |
|
||||
| Suite | **12 passed** (`tests/test_v11_d_recovery.py`) |
|
||||
|
||||
## Q. WP-E regression
|
||||
|
||||
| Claim | Result |
|
||||
| --- | --- |
|
||||
| Contrast audit exit code | **0** |
|
||||
| All applicable control boundaries ≥ 3:1 | **PASS** — every boundary pair clears 3:1 (1.4.11) |
|
||||
| Applicable text contrast still compliant | **PASS** — every text pair clears 4.5:1 (1.4.3); baselines 14.57 / 13.57 / 5.48 / 5.88 unchanged |
|
||||
| Browser boundary checks | **PASS** — the 10 WP-E rows in §G |
|
||||
| Focus visibility | **PASS** — M11's visible-focus check inside the 38, plus WP-E's focused-edge measurement |
|
||||
| Gate tests | **11 passed** (`tests/test_v11_e_contrast.py`), including 2.99:1 failing and 3.00:1 passing |
|
||||
|
||||
```text
|
||||
OWNER SCREENSHOT APPROVAL: APPROVED
|
||||
```
|
||||
|
||||
Approved by the owner in the release-validation brief of 2026-09-16. **No visual
|
||||
code changed during release validation**, so that approval remains valid; had any
|
||||
changed, it would have been void and new screenshots would have been required.
|
||||
|
||||
## R. Release-shaped smoke test
|
||||
|
||||
The **final no-cache candidate image**, a fresh volume, published on loopback,
|
||||
with the private CA installed into the container's own trust store. Evidence:
|
||||
`…/release-87a4032/smoke/`.
|
||||
|
||||
**15 checks, 15 passed, 0 failed.**
|
||||
|
||||
| Claim | Result |
|
||||
| --- | --- |
|
||||
| The container starts | PASS |
|
||||
| The application answers on loopback | PASS |
|
||||
| The port is published on **loopback only** | PASS — `8000/tcp -> 127.0.0.1:…` |
|
||||
| This machine's **LAN address does not serve** the application | PASS |
|
||||
| The first page loads | PASS — HTTP 200 |
|
||||
| The shell references no remote origin | PASS — none found |
|
||||
| A CSP is served | PASS |
|
||||
| **The approved HTTPS narrator verifies through its private CA** | PASS — HTTP 200 through `tlstrust.ssl_context()`, **no bypass** |
|
||||
| **A public endpoint is refused** | PASS — HTTP 400 |
|
||||
| A campaign is created | PASS |
|
||||
| **One real narrator turn is accepted** | PASS |
|
||||
| The container restarts and serves again | PASS |
|
||||
| The transcript survived the restart | PASS |
|
||||
| The narrative state survived the restart | PASS |
|
||||
| **Firefox renders the reopened campaign** | PASS — 470 characters of story |
|
||||
|
||||
**A finding worth recording, and it is not a product defect.** The first attempt
|
||||
failed at `PUT /api/settings` with **HTTP 400**. The cause: a `.local` name is
|
||||
mDNS, a Docker container has no mDNS resolver, and `endpoints.py` correctly
|
||||
refuses an endpoint whose address it cannot classify — the policy behaving
|
||||
exactly as designed. The fix is to resolve the name inside the container
|
||||
(`--add-host`), **not** to substitute the IP address, because the certificate is
|
||||
issued for the hostname and substituting the address would have quietly bypassed
|
||||
the hostname verification this test exists to prove. Harness-only; no product
|
||||
code changed.
|
||||
|
||||
This is supplemental evidence, not a substitute for the gates above.
|
||||
|
||||
## S. Security / local-only review
|
||||
|
||||
| Claim | Evidence on the candidate |
|
||||
| --- | --- |
|
||||
| Served on loopback only | every harness reached the application on `127.0.0.1`; the documented container run publishes loopback |
|
||||
| Endpoint policy | `app/endpoints.py` admits loopback (v4 and v6), the three RFC1918 ranges, link-local, IPv6 unique-local and CGNAT, and refuses the public Internet; a name resolving to both a private and a public address is refused |
|
||||
| Inference actually used | trusted-LAN **HTTPS** with a private CA for the browser gate (§G); plain HTTP to a LAN GPU host for the long run, which `SECURITY-THREAT-MODEL.md` §83 permits and which is not A06 evidence |
|
||||
| No secret in exports | offline gate, WP-D tests and the M11 suite |
|
||||
| CSP served | `default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; font-src 'self'; img-src 'self' data:; connect-src 'self'; object-src 'none'; base-uri 'none'; form-action 'self'; frame-ancestors 'none'` |
|
||||
| Other headers | `x-content-type-options: nosniff`, `referrer-policy: same-origin`, `x-frame-options: DENY` |
|
||||
| No remote origin in the shell | offline gate: the page names none, and every asset is local |
|
||||
|
||||
`'unsafe-inline'` remains on `style-src` only, because React writes inline
|
||||
`style` attributes; it is deliberately absent from `script-src`.
|
||||
|
||||
## T. Known residual risks
|
||||
|
||||
Each is classified, and none is collapsed into another category.
|
||||
|
||||
### T.1 WP-B reference-model memory limitation — ACCEPTED RESIDUAL
|
||||
|
||||
```text
|
||||
deterministic independent memory: PASS
|
||||
reference-model independent memory: FAIL
|
||||
failing stage: memory creation — the summariser's content selection
|
||||
owner decision: accepted for v1.1
|
||||
```
|
||||
|
||||
**Status on this candidate:** unchanged, and no broader regression. The release
|
||||
long run did not demonstrate independent recovery, but it also could not: its
|
||||
probe failed the `absent_from_later_narration` precondition because the narrator
|
||||
restated the fact (§K.2). Ranking, eviction and excerpt creation all pass
|
||||
deterministically in the 1,723-test suite, and the bank worked normally through
|
||||
102 turns (19 memories, ranked and injected). **Accepted residual**, carried
|
||||
visibly, not a blocker.
|
||||
|
||||
### T.2 A2 mid-reply application-instruction echo — ACCEPTED RESIDUAL, still live
|
||||
|
||||
The known occurrence is WP-B.1's action 153 (depth 143): the narrator echoed the
|
||||
length hint **mid-reply** and then continued the story, which A2's trailing
|
||||
cleanup does not remove.
|
||||
|
||||
**Reproduced on this candidate.** Replaying that stored fixture through the
|
||||
release-gate detector still flags it — **1 of 105** AI actions, matched by the
|
||||
hard-limit/hint rule alone. The extractor still leaves it. This report therefore
|
||||
does **not** claim protocol leakage is solved.
|
||||
|
||||
**But it did not recur in release evidence.** This run's own 105 stored replies
|
||||
leak **0** (§J), and the browser run's narrated turns leak 0. The brief's
|
||||
stop-and-report condition — *an equivalent shape occurring in this final run* —
|
||||
was **not** triggered, so validation continues. No broader sanitizer was written;
|
||||
that remains for owner review.
|
||||
|
||||
### T.3 Doubled full stop in memory-search scene text — ACCEPTED RESIDUAL
|
||||
|
||||
`"rain outside.."` when the state's scene summary already ends in punctuation.
|
||||
Only the embedding query sees it; the effect is one stray token. Not changed
|
||||
during release validation, deliberately: cleanliness is not a reason to alter
|
||||
product behaviour after evidence is taken.
|
||||
|
||||
### T.4 K1 — "Correct" on an Important Facts row is always refused
|
||||
|
||||
**Reproduced, unchanged on this candidate**, deterministically and without a
|
||||
browser: Correct on a Characters row (`subject='mara'`) applies, **201**; Correct
|
||||
on an Important Facts row (`subject='f1'`) is refused, **400 — "add_fact names
|
||||
subject='f1', which does not exist."** The cause is frontend-side: the panel
|
||||
sends the row key as `add_fact.subject`, and the validator checks `subject` as an
|
||||
entity reference.
|
||||
|
||||
**Classification: v1.2 backlog, not a release blocker.** It blocks no v1
|
||||
REQUIRED test — C04 passes through the working correction paths (§C) — and WP-C's
|
||||
State-panel correction coverage passes. It is a narrow bug with an obvious fix
|
||||
(offer Correct only against entities, or send facts without a subject), but
|
||||
fixing it during release validation would change product code after the evidence
|
||||
above was taken, which the brief forbids without a product brief. **Left for the
|
||||
owner.**
|
||||
|
||||
### T.5 Import ceiling and scheduled backups — INTENTIONAL, not unfinished work
|
||||
|
||||
The 20 MB import ceiling is deliberate and unchanged; WP-D made it honest rather
|
||||
than raising it. Scheduled backups remain unbuilt by design. Neither is a
|
||||
residual defect.
|
||||
|
||||
### T.6 Harness corrections made during validation — no product evidence invalidated
|
||||
|
||||
Three, all harness-only, each named where it occurred: the identity diagnostic
|
||||
did not enable memory (§L); the smoke test needed hostname resolution inside the
|
||||
container (§R); and the first upgrade campaign was too short to write memories
|
||||
and carried a live hostname (§N). **No product code changed at any point during
|
||||
release validation**, so no black-box or long-run evidence became stale. Each
|
||||
corrected harness repeated its own affected check.
|
||||
|
||||
## U. Deferred v1.2 / future work
|
||||
|
||||
| Item | Why it is deferred |
|
||||
| --- | --- |
|
||||
| Raising the import ceiling, or a streaming import | The plan assigns it to v1.2; WP-D's scope was honesty about the limit, not the limit |
|
||||
| Scheduled backups; a restore button | Explicitly out of WP-D's scope |
|
||||
| K1's Correct-on-a-fact-row fix (§T.4) | A narrow frontend bug needing a product brief |
|
||||
| A broader mid-reply protocol sanitizer (§T.2) | Needs owner review; A2's cleanup is deliberately trailing-only |
|
||||
| The doubled full stop (§T.3) | Cosmetic, embedding-query only |
|
||||
| K05 generate local image, K06 multi-turn video | FUTURE tests; both need a media provider this release does not build |
|
||||
| Reference-model independent memory (§T.1) | Needs a stronger summariser or a different creation strategy — a v1.2 investigation, not a v1.1 fix |
|
||||
|
||||
## V. Documentation changes
|
||||
|
||||
Current documents were brought up to date. **No historical milestone report was
|
||||
rewritten, and no failed WP-B real-model evidence was turned into success.**
|
||||
|
||||
| Document | Change |
|
||||
| --- | --- |
|
||||
| `README.md` | Status now says v1.0.0 **remains** the released version, that all six v1.1 packages are complete and accepted, that release validation passed on candidate `87a4032`, that WP-B ships with a documented limitation, and that **no `v1.1.0` tag exists and `main` is unchanged`**. Also corrected a stale figure: the schema is versioned at **94**, not "92 and counting" |
|
||||
| `planning/V1.1-PLAN.md` | WP-D/WP-E recorded as signed `87a4032`; the release-validation outcome summarised with its residuals; the three owner events named as still outstanding |
|
||||
| `planning/VERSION.md` | Same status correction, plus a new revision entry for the closeout |
|
||||
| `planning/README.md` | Current-state paragraph rewritten for the same facts |
|
||||
| `planning/reports/v1.1/V1.1-WP-E-REPORT.md` | `OWNER SCREENSHOT APPROVAL: PENDING` → **`APPROVED`**, with the source (the release-validation brief) and date recorded, and a note that the signed commit predated the review |
|
||||
| `planning/reports/v1.1/V1.1-RELEASE-REPORT.md` | **New** — this document |
|
||||
| `DEVELOPMENT.md` | Unchanged: its WP-C/WP-D sections already describe the candidate as built |
|
||||
|
||||
**New tools committed with this closeout** (harness only, no product code):
|
||||
`backend/tools/v11_upgrade_check.py` (Gate 9) and
|
||||
`backend/tools/v11_release_smoke.py` (§R), plus two narrow corrections to
|
||||
`backend/tools/m11_identity.py` (read `AIDND_TEST_EMBED_MODEL`; enable the memory
|
||||
bank) so the diagnostic can run with memory on.
|
||||
|
||||
## W. Final release decision
|
||||
|
||||
**The question this validation set out to answer:** does candidate `87a4032`
|
||||
preserve the complete v1 contract and satisfy every accepted v1.1 package on one
|
||||
integrated release tree?
|
||||
|
||||
| Gate | Result |
|
||||
| --- | --- |
|
||||
| 1 Package acceptance | **PASS** — six packages accepted; WP-B's qualification carried whole (§B.1) |
|
||||
| 2 v1 acceptance contract | **PASS** — 81 PASS, H09 NOT APPLICABLE, 0 waived, 0 weakened, 0 reclassified (§C) |
|
||||
| 3 Suites, lint, build, Docker | **PASS** — 1,723 / 175 / 0 errors; image SPA file-for-file identical (§D, §E) |
|
||||
| 4 Offline / no-network | **PASS** — 23/23 on the candidate image (§F) |
|
||||
| 5 Browser release run | **PASS** — 101/0/0 over trusted-LAN HTTPS, every turn `fits` (§G) |
|
||||
| 6 Integrated long run | **PASS** — 102 turns, M01–M04, 0 post-turn failures (§H) |
|
||||
| 7 Identity diagnostic | **PASS** — 0 signals, 0 protocol shapes, memory on (§L) |
|
||||
| 8 Recovery | **PASS** — 16/16 on the run's own bundle (§M) |
|
||||
| 9 v1.0.0 upgrade | **PASS** — 15/15 identical, schema parity at 94 (§N) |
|
||||
| 10 WP-D regression | **PASS** (§P) |
|
||||
| 11 WP-E regression | **PASS**, screenshots approved (§Q) |
|
||||
| Bundle compatibility | **Both directions execute and import** (§O) |
|
||||
| Release smoke | **PASS** — 15/15 from the shipped image (§R) |
|
||||
|
||||
**A1** holds its reserve on every re-counted prompt, with a smallest margin of
|
||||
**879 tokens** against v1's 23–42. **A2**'s release-gate leak count is **0**.
|
||||
**WP-B**'s deterministic independent-memory recovery passes and its
|
||||
reference-model limitation remains an accepted, documented residual — stated in
|
||||
§K.3 in both halves, never shortened to "WP-B passed".
|
||||
|
||||
No product code was changed at any point during release validation, so no
|
||||
evidence was invalidated. Three harness corrections were made and each corrected
|
||||
harness repeated its own check (§T.6).
|
||||
|
||||
```text
|
||||
V1.1 RELEASE VALIDATION:
|
||||
PASS
|
||||
```
|
||||
|
||||
### What this decision is not
|
||||
|
||||
These are separate, and only the first is done:
|
||||
|
||||
```text
|
||||
WP-A-E accepted: YES
|
||||
release validation passed: YES
|
||||
release candidate prepared: YES (87a4032, with this report)
|
||||
release commit signed: NO
|
||||
main updated to v1.1: NO
|
||||
v1.1.0 tagged: NO
|
||||
```
|
||||
|
||||
The closeout changes are **staged and uncommitted**. Nothing was committed,
|
||||
pushed, merged or tagged by this validation. The owner's next decision is to
|
||||
review this evidence, resolve anything they disagree with, then sign the v1.1
|
||||
release commit and publish `v1.1.0`.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,623 @@
|
||||
# v1.1 WP-B.1 — Independent Long-Term Memory Retention Diagnostic
|
||||
|
||||
**Status:** COMPLETE. Diagnostic only: no memory behaviour was changed. The final decision is in §S.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD at start | `d63804f22ecbaed80741241a154cdaef82f7b2ed` — *v1.1: harden context window and narrator protocol boundary* (WP-A1/A2), signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||||
| Its parents | `ac465ed` (planning v4.1, signed) and `432f041` (the signed v1.0.0 release commit; tag `v1.0.0`; `main`) |
|
||||
| Working tree at start | clean |
|
||||
| `git diff --stat v1.0.0..HEAD` | 27 files, +4,880 / −88: the planning commit and WP-A1/A2 |
|
||||
| Comparison baseline | `432f041` (v1.0.0), run from a throwaway worktree (§D) |
|
||||
|
||||
`app/memorybank.py`, `app/summaries.py`, `app/context/lineage.py`, `app/tree.py` and
|
||||
`app/vectors.py` are unchanged between `v1.0.0` and HEAD. A1 and A2 did not touch
|
||||
the memory pipeline, apart from adding `accounting` to `attempts.ATTEMPT_KEYS`.
|
||||
|
||||
---
|
||||
|
||||
## B. Existing memory pipeline
|
||||
|
||||
Answered from the code at HEAD. Nothing was changed.
|
||||
|
||||
| # | Question | Answer |
|
||||
| --- | --- | --- |
|
||||
| 1 | How are memories generated? | `memorybank.run_post_turn`, a fire-and-forget task after each accepted turn (`schedule_post_turn`), runs `_create_due_memories` when the campaign has `auto_summarize`. It writes one memory per block of `MEMORY_INTERVAL` = 6 story actions past the memory cursor. It starts once the story has `MEMORY_START` = 12 actions, and only when `SETTLE_SLACK` = 1 action sits past the block. At most `MAX_MEMORIES_PER_RUN` = 5 memories are written per run. |
|
||||
| 2 | What range does a memory cover? | The block's first and last action depths: `Memory.source_start`, `Memory.source_end`. |
|
||||
| 3 | How are source depth and lineage stored? | `tree.attach_memory` sets `Memory.branch_id` and `Memory.depth` from the block's **last** node, so a memory is visible exactly on paths that contain that node. |
|
||||
| 4 | Memory text length limit | Prompt-only: `MEMORY_MAX_WORDS` = 50 in `MEMORY_SYSTEM_PROMPT` ("1-2 plain sentences"). Nothing truncates the stored text. |
|
||||
| 5 | What input does the summariser receive? | `summarize_block`: a cast brief (`cast_brief`, which reads story cards, persona and plot essentials), then `"Story excerpt:\n\n{excerpt}\n\nMemory:"`. The excerpt is the block's action texts joined by blank lines. |
|
||||
| 6 | Where does the 2,000-token truncation happen? | `summarize_block`: `excerpt = truncate_to_last_tokens(raw, MEMORY_EXCERPT_TOKENS)`, with `MEMORY_EXCERPT_TOKENS` = 2,000. It keeps the block's **last** 2,000 `cl100k_base` tokens. The cast brief is matched against the untruncated block. |
|
||||
| 7 | How does capacity and eviction work? | `_evict_over_capacity`, at the end of every `run_post_turn`. It counts non-forgotten memories in the whole adventure (not lineage-scoped). Above `settings.memory_bank_capacity` (default 80), it marks `forgotten = True` on the overflow unpinned memories, ordered by `coalesce(last_used_at, created_at)` ascending, then `use_count`. Forgotten rows are kept. |
|
||||
| 8 | How is `last_accessed` updated? | The field is `Memory.last_used_at`. `record_use` sets it, and increments `use_count`, for every memory in the turn's `memories.used`. That write is part of the turn's single commit (§O.7 of the M11 report). Dry runs never count. |
|
||||
| 9 | How are memories ranked? | `retrieve_memories`. The query is the text of the newest `RETRIEVAL_WINDOW_ACTIONS` = 4 actions, truncated to the last `RETRIEVAL_WINDOW_TOKENS` = 600 tokens. It is embedded with the configured embedding model. Every eligible memory is scored by `vectors.cosine`. Pinned memories are taken first, the rest fill `memory_top_k` (default 5) in score order, and `_drop_redundant` skips a candidate at cosine ≥ 0.93 to one already chosen, within the same authority. **There is no lexical term, recency term, importance term or similarity floor** (CONTEXT-AND-MEMORY §20, as implemented). |
|
||||
| 10 | How do pins affect ranking and eviction? | Ranking: a pinned memory is always selected, and counts toward `memory_top_k`. Eviction: pinned memories are never evicted, and if every active memory is pinned, capacity is exceeded. |
|
||||
| 11 | How does the active-lineage clause filter memories? | `lineage.path_of(db, adventure).clause(models.Memory)` matches `(branch_id, depth)` against the head's path entries, capped at each fork depth. Retrieval also requires `forgotten = False` and `embedded = True`. |
|
||||
| 12 | How do retrieved memories enter `build_context`? | Through the `memory_bank` argument. The builder renders `"Memories from earlier in the story. Lines marked [inferred] are interpretation, not established fact — do not treat them as settled truth:"` plus one `- [inferred]? text` line per memory, as section `used_memories`. It is a live section priced into protected context, placed after history and before `narrative_state`. |
|
||||
| 13 | Where is the provenance recorded? | `context_snapshot["memories"]`: the whole retrieval result, meaning `used` (id, text, similarity, pinned, authority, source range), `considered` and `suppressed`. It is stored per turn, so a past turn's selection is inspectable. |
|
||||
| 14 | How does memory survive export and import? | `bundle._exported_memory` carries text, pinned, forgotten, `sourceStart`, `sourceEnd`, `useCount`, authority, branch and depth. **It does not carry the vector, `last_used_at` or `created_at`.** On import the memories are re-embedded by the post-turn pass, and their recency restarts from the import. |
|
||||
|
||||
**A finding from the inspection itself.** The v1 long-run harness (`tools/m11_long_run.py`)
|
||||
planted M04's clue as an accepted **state correction** (`CLUE_FACT`, via
|
||||
`add_fact`). The player turn mentioning it uses the sentinel code, but the fact the
|
||||
recall looks for was established in state. Memories are written from **story
|
||||
text**. So every v1 M04 recovery could only ever run through state, or through the
|
||||
narrator restating state, and none could have tested memory on its own. That is why
|
||||
§P risk 5 of the M11 report could say "no run showed memory keeping a planted fact"
|
||||
without any run having given memory the chance.
|
||||
|
||||
---
|
||||
|
||||
## C. Diagnostic design
|
||||
|
||||
### C.1 What is built
|
||||
|
||||
| File | Kind | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `backend/tools/memory_diagnostic.py` | tool, new | See below: the fact spec, isolation checks, four-stage diagnosis, deterministic stubs and scenario runner. |
|
||||
| `backend/tools/v11_b1_memory.py` | tool, new | CLI. `scenarios` runs the deterministic campaigns against an isolated database and writes JSON. `diagnose` runs the four stages against a **copy** of a finished real campaign's database. |
|
||||
| `backend/tools/m11_long_run.py` | tool, extended | `--independent-fact`: see §C.3. |
|
||||
| `backend/tests/test_v11_b1_memory_diagnostic.py` | tests, new | The deterministic diagnostic. |
|
||||
| `backend/tests/test_v11_b1_long_run_verdict.py` | tests, new | The new long-run verdict. |
|
||||
|
||||
What `tools/memory_diagnostic.py` holds:
|
||||
- **`Fact`:** a planted fact, with whole-word carry and leak matching.
|
||||
- **`isolation()`:** every non-memory layer checked.
|
||||
- **`diagnose()`:** the four stages and the verdict.
|
||||
- **`rank_bank()`:** production ranking recomputed for every eligible memory.
|
||||
- **The stubs:** `BestCaseSummariser`, `ConceptEmbedder` and `ScriptNarrator`.
|
||||
- **`run_scenario()`:** a campaign played through the real turn route.
|
||||
|
||||
**No application file is changed.** No column, table, migration or setting is
|
||||
added. The diagnostic's extra fields are computed at report time.
|
||||
|
||||
### C.2 How each stage is judged
|
||||
|
||||
| Stage | Judged from | Output (real field names where they exist) |
|
||||
| --- | --- | --- |
|
||||
| **created** | Memories on the active lineage with `source_start ≤ plant_depth ≤ source_end` whose text carries F (both "sundial" and "teapot"). For every covering memory, the block is re-read (`memorybank.source_block`) and cut exactly as `summarize_block` does, to report whether F was in the block and whether it was in **the excerpt the summariser saw**. | `memory_id`, `source_start`, `source_end`, `memory_text`, `covering_memories[].{block_tokens, fact_in_block, fact_in_summariser_excerpt}` |
|
||||
| **retained** | That memory's row, plus the eviction order production would use (the same `ORDER BY`, read-only) | `forgotten`, `pinned`, `embedded`, `on_active_lineage`, `use_count`, `last_used_at`, `created_at`, `active_memories`, `memory_bank_capacity`, `eviction_position`, `reason` |
|
||||
| **ranked** | `rank_bank`: the same catalogue clause, cosine, pin rule and `memorybank._drop_redundant`, over **every** eligible memory, with the recall turn's own query (`history.tail(4, exclude=recall AI node)`, cut to 600 tokens). It is checked against the recall turn's stored `memories.used` (`replica_matches_stored_selection`). | `semantic_score` (`similarity`), `lexical_score` (always `None`: no such term exists), `final_score`, `rank`, `of`, `top_k_cutoff`, `selected`, `suppressed_as_duplicate_of`, `query` |
|
||||
| **injected** | The recall turn's stored `context_snapshot`: `memories.used` names the memory, and its text is in the `used_memories` section | `context_component`, `in_stored_memories_used`, `text_in_section`, `token_count` |
|
||||
|
||||
Verdicts, in order: `not_created`, `created_but_evicted`, `retained_but_not_ranked`,
|
||||
`ranked_but_not_selected`, `selected_but_not_injected`, `injected`. The fifth is
|
||||
added to the brief's list, so that "the retrieval picked it" and "the narrator was
|
||||
shown it" stay distinguishable.
|
||||
|
||||
### C.3 Isolation (precondition) checks
|
||||
|
||||
A result counts only if every check holds.
|
||||
|
||||
| Check | How |
|
||||
| --- | --- |
|
||||
| `state_document` | `adventure.narrative_state`: entities, facts, relationships, threads, scene and possessions, plus the whole document |
|
||||
| `state_snapshots` | `narrative_state_after` of every node on the active lineage |
|
||||
| `later_narration` | every AI turn deeper than the planting block's end, and before the recall turn |
|
||||
| `summary` | `summaries.current` and the recall prompt's `story_summary` section |
|
||||
| `knowledge` | every `KnowledgeSource.content`, and the recall prompt's imported-knowledge sections |
|
||||
| `recent_history` | the recall prompt's `history.floor_depth` is greater than the planting depth, and F is not in the `history` or `recent_history` sections |
|
||||
| `state_section` | the recall prompt's `narrative_state` section |
|
||||
|
||||
The negative control for the precondition itself is
|
||||
`test_the_isolation_check_fails_when_another_layer_carries_the_fact`.
|
||||
|
||||
### C.4 The deterministic stubs, and what they model
|
||||
|
||||
- **`BestCaseSummariser`.** An ideal memory writer. A memory keeps every sentence of
|
||||
its excerpt that carries a planted fact, plus one sentence naming the block's own
|
||||
place so that memories differ. Summary updates never mention a planted fact. **A
|
||||
creation failure under this stub is the application's, not a model's.**
|
||||
- **`ConceptEmbedder`.** A 96-dimension deterministic embedding. Words in a small
|
||||
concept table ("sundial", "dial", "hour", "clock", …) share a dimension, other
|
||||
words are hashed, and the vector is normalised. It models a paraphrase landing
|
||||
near the original. **It says nothing about `nomic-embed-text`.**
|
||||
- **`ScriptNarrator`.** Narration that names only filler places and never a planted
|
||||
fact, with an empty state block, so state never records F.
|
||||
|
||||
Scenarios are played through the real `POST /actions` route. Automatic post-turn
|
||||
scheduling is replaced by an explicit `run_post_turn` settle after every turn, so
|
||||
eviction happens at a known turn. Depths: the opening is 0, turn *n*'s player
|
||||
action is 2n−1 and its reply 2n. The planted fact is a `story` action.
|
||||
|
||||
### C.4.1 Scenarios
|
||||
|
||||
| Scenario | Turns | Capacity | top_k | History budget | Prose per reply | Planted at |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| `independent_default` | 52 (recall at depth 106) | 80 | 5 | 4,096 | ~60 words | depth 1 |
|
||||
| `past_capacity` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1 |
|
||||
| `past_capacity_pinned` | 52 | 6 | 5 | 4,096 | ~60 words | depth 1; first other memory pinned |
|
||||
| `past_capacity_low_top_k` | 52 | 8 | 2 | 4,096 | ~60 words | depth 1 |
|
||||
| `long_block_fact_early` | 10 | 80 | 5 | 16,384 | ~850 words | depth 1, early in a long block |
|
||||
| `long_block_fact_late` | 10 | 80 | 5 | 16,384 | ~850 words | depth 5, late in the same-sized block |
|
||||
| `lineage_control` | 52 | 80 | 5 | 4,096 | ~60 words | F at depth 1; G on line A, then Undo × 9 and divergence at turn 30; Save Points before and after G |
|
||||
|
||||
### C.5 The long-run verdict
|
||||
|
||||
`tools/m11_long_run.py --independent-fact` plants a second, **story-only** fact
|
||||
at depth 3, right after M04's own plant, and never corrects it into state. Every
|
||||
existing M04 behaviour and verdict is unchanged.
|
||||
|
||||
**Every accepted turn** records the first failure of:
|
||||
- `absent_from_state`;
|
||||
- `absent_from_summary`;
|
||||
- `absent_from_later_narration`, meaning narration deeper than the planting depth
|
||||
plus 6.
|
||||
|
||||
The plant and those first failures survive `--resume`.
|
||||
|
||||
**At recall** a dedicated question is played. `_independent_recall` reads the
|
||||
recall turn's stored context and the campaign database (read-only), and
|
||||
`_independent_memory_verdict` returns one of:
|
||||
|
||||
- `recovered_through_memory_independent`, only when every one of
|
||||
`planted_turn_outside_history`, `absent_from_state`, `absent_from_summary`,
|
||||
`absent_from_knowledge` and `absent_from_later_narration` holds, and a memory
|
||||
covering the planting turn carries the fact and was injected;
|
||||
- `precondition_failed:<name>` or `precondition_unknown:<name>`, never a recovery;
|
||||
- `not_recovered:not_created`, `not_recovered:evicted` or
|
||||
`not_recovered:not_injected`.
|
||||
|
||||
Ranking is recomputed afterwards by `tools/v11_b1_memory.py diagnose`, against a
|
||||
copy of the run's database.
|
||||
|
||||
---
|
||||
|
||||
## D. v1.0.0 baseline
|
||||
|
||||
**How it was run.**
|
||||
- A throwaway worktree was checked out at `432f041` (`git describe`: `v1.0.0`).
|
||||
- Only the three B.1 files were copied in: `tools/memory_diagnostic.py`,
|
||||
`tools/v11_b1_memory.py` and `tests/test_v11_b1_memory_diagnostic.py`.
|
||||
- The imported `app` was confirmed to come from the worktree.
|
||||
- The worktree was removed afterwards, and the `v1.0.0` tag and commit were not
|
||||
touched.
|
||||
- `app/memorybank.py` is byte-identical between `v1.0.0` and HEAD.
|
||||
- Evidence: `$HOME/v11-evidence/b1/v100/`, with HEAD's in
|
||||
`$HOME/v11-evidence/b1/head/`.
|
||||
|
||||
| | v1.0.0 (`432f041`) | HEAD (`d63804f` + B.1 files) |
|
||||
| --- | --- | --- |
|
||||
| `test_v11_b1_memory_diagnostic.py` | **24 passed, 2 xfailed (strict)** | 24 passed, 2 xfailed (strict) |
|
||||
| `independent_default` | `injected` | `injected` |
|
||||
| `past_capacity` (capacity 6, top_k 5) | **`created_but_evicted`** | `created_but_evicted` |
|
||||
| `past_capacity_pinned` | `created_but_evicted` (the pinned memory kept) | same |
|
||||
| `past_capacity_low_top_k` (capacity 8, top_k 2) | **`created_but_evicted`** | `created_but_evicted` |
|
||||
| `long_block_fact_early` | **`not_created`** | `not_created` |
|
||||
| `long_block_fact_late` | `injected` | `injected` |
|
||||
| `lineage_control` | `injected`; G never injected after the divergence | same |
|
||||
|
||||
**The two independent-retention acceptance criteria fail on v1.0.0, and each names
|
||||
the stage.**
|
||||
|
||||
```text
|
||||
criterion: an early fact is recalled from memory past capacity
|
||||
created: yes
|
||||
retained: no
|
||||
FAILURE STAGE: retention (capacity eviction)
|
||||
|
||||
criterion: a fact early in a long block is remembered
|
||||
created: no (the fact was in the block, not in the summariser's excerpt)
|
||||
FAILURE STAGE: creation (input truncation)
|
||||
```
|
||||
|
||||
Under the best-case summariser, with blocks shorter than 2,000 tokens and a bank
|
||||
under capacity, v1.0.0 carries the fact all the way to injection. The two failure
|
||||
stages above are therefore **application mechanisms**, reached under conditions
|
||||
a long campaign meets:
|
||||
- a block of long narration;
|
||||
- more memories than `memory_bank_capacity`.
|
||||
|
||||
Which of them a real campaign meets first is §K's question.
|
||||
|
||||
---
|
||||
|
||||
## E. Creation results
|
||||
|
||||
`long_block_fact_early` and `long_block_fact_late` use the same block geometry: the
|
||||
opening (depth 0) and turns 1-3. Each reply is about 850 words, and the block
|
||||
`source_start` 0 … `source_end` 5 is **2,079 tokens**, 79 over
|
||||
`MEMORY_EXCERPT_TOKENS`.
|
||||
|
||||
| Shape | Planted at | Fact in block | Fact in summariser excerpt | Memory written | Verdict |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| F early in the block | depth 1 | yes | **no** | "The travellers spent time at the ferry landing." | **`not_created`** |
|
||||
| F late in the same-sized block | depth 5 | yes | yes | "Mara slipped the amber sundial inside the cracked teapot on the tavern's top shelf. The travellers spent time at the ferry landing." | `injected` |
|
||||
|
||||
`summarize_block` keeps the **last** 2,000 tokens. An early-block fact is cut off
|
||||
before the summariser reads it, even by an overflow of only 79 tokens. No summariser
|
||||
quality can recover what it was never given.
|
||||
|
||||
`independent_default`: 60-word replies, a 4,096 budget, and a block under 2,000
|
||||
tokens. Memory 1 covers depths 0-5, and its text carries F. Both the block and the
|
||||
excerpt contain F.
|
||||
|
||||
---
|
||||
|
||||
## F. Retention / capacity results
|
||||
|
||||
Default: the bank holds 17 active memories at depth 106, against capacity 80. F's
|
||||
memory is retained, has `use_count` 16, and sits at eviction position 14 of 17.
|
||||
|
||||
**Past capacity** (the brief's capacity test). 17 memories are written over 52 turns.
|
||||
|
||||
| Scenario | F's uses before eviction | F's last retrieval | First eviction | F evicted | F first evicted? | Created and evicted in the same pass | Pinned kept |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| `past_capacity` (6 / top_k 5) | 12 | turn 18 | turn 21, memory 1 | **turn 21** | **yes** | none | — |
|
||||
| `past_capacity_pinned` (6 / top_k 5, memory 2 pinned) | 12 | turn 18 | turn 21, memory 1 | turn 21 | yes | none | **memory 2 never evicted** |
|
||||
| `past_capacity_low_top_k` (8 / top_k 2) | 4 | turn 24 | turn 27, memory 3 | **turn 36** (4th eviction) | no | none | — |
|
||||
|
||||
**What the trace shows about the rule.**
|
||||
- **"Never retrieved" is not the mechanism.** F was retrieved, 12 or 4 times, while
|
||||
it still ranked in the top-k for the recent-narration query.
|
||||
- **Eviction orders by `coalesce(last_used_at, created_at)`.** An early fact that
|
||||
recent narration never mentions stops being retrieved once newer memories fill
|
||||
the top-k. It then ages out:
|
||||
- at capacity 6, it is gone 3 turns after its last use, as the first eviction;
|
||||
- at capacity 8 with `top_k` 2, 12 turns after its last use, as the fourth.
|
||||
- **Recall itself cannot rescue it.** Retrieval is driven by recent narration, and
|
||||
a fact nobody mentions is exactly the one that loses recency.
|
||||
- **Pinned memories stay protected** (`past_capacity_pinned`).
|
||||
- **No memory is evicted by the pass that created it.** The frozen-bank regression
|
||||
fix still holds in all three scenarios.
|
||||
|
||||
---
|
||||
|
||||
## G. Ranking results
|
||||
|
||||
The `independent_default` recall turn is at depth 106. Its query is the newest 4
|
||||
actions, cut to 600 tokens: three narration turns and the one-line question.
|
||||
|
||||
| | Value |
|
||||
| --- | --- |
|
||||
| F's `similarity` (semantic score) | **0.241** |
|
||||
| Lexical score | none: memory ranking has no lexical term |
|
||||
| Pin effect | none (not pinned) |
|
||||
| Final rank | **2 of 17** |
|
||||
| `memory_top_k` cutoff | 5 |
|
||||
| Selected | yes |
|
||||
| Replica agrees with the turn's stored `memories.used` | **yes** |
|
||||
|
||||
The same memory against three stand-alone queries, `ConceptEmbedder`:
|
||||
|
||||
| Query | Rank | Similarity | Selected |
|
||||
| --- | --- | --- | --- |
|
||||
| direct: "I ask Mara where she hid the amber sundial." | 1 of 17 | 0.708 | yes |
|
||||
| paraphrase: "… the little brass dial that tells the hour." | 1 of 17 | 0.636 | yes |
|
||||
| unrelated: "… what rope costs at the landing this season." | 5 of 17 | **0.064** | **yes** |
|
||||
|
||||
Two diagnostic observations. Neither is a failure in this scenario.
|
||||
1. **The production query dilutes the question.** The question alone scores 0.708;
|
||||
inside the four-action window it scores 0.241. F survives at rank 2 of 17. In a
|
||||
bank where more memories share the recent narration's vocabulary, the same
|
||||
dilution would push it past `top_k`.
|
||||
2. **There is no relevance floor.** With more memories than `memory_top_k`, five are
|
||||
injected whatever their similarity. An unrelated query still injects F at 0.064.
|
||||
This matters to F only in the other direction: it can ride along even when it is
|
||||
not relevant.
|
||||
|
||||
---
|
||||
|
||||
## H. Injection results
|
||||
|
||||
`independent_default` injects F: `memories.used` in the recall turn's stored snapshot
|
||||
names memory 1, its text is in the `used_memories` section, and that section is 97
|
||||
tokens. The same holds in `long_block_fact_late` and `lineage_control`.
|
||||
**`selected_but_not_injected` never occurred**: whatever retrieval selected, the
|
||||
builder rendered.
|
||||
|
||||
---
|
||||
|
||||
## I. Lineage negative control
|
||||
|
||||
`lineage_control`:
|
||||
- G is planted on line A at depth 41, turn 21.
|
||||
- G's memory 7 is written.
|
||||
- A Save Point is placed on line A after G.
|
||||
- Undo ×9, then divergent writing at turn 30. The last action before the
|
||||
divergence has id 59.
|
||||
|
||||
| Assertion | Result |
|
||||
| --- | --- |
|
||||
| G's memory stays stored | **yes** (memory 7 present) |
|
||||
| G is not eligible on the active lineage | **yes** (the path clause returns nothing) |
|
||||
| G is never injected after the divergence | **yes** (no turn with id > 59 names it or carries its text) |
|
||||
| `memories.used` does not report it after the divergence | **yes** |
|
||||
| Returning to line A (Save Point restore) makes it eligible again | **yes** (memory 7 eligible) |
|
||||
| F on the active line is unaffected | `injected`, isolation holds |
|
||||
|
||||
Before the divergence, G's memory was legitimately used on line A (turn 22). The
|
||||
first version of this check counted that as a leak. That was a defect in the
|
||||
diagnostic, and the scan now starts after the divergence.
|
||||
|
||||
No lineage code was touched. `test_m11_leakage.py`, 14 tests, passes unchanged (§P).
|
||||
|
||||
---
|
||||
|
||||
## J. Authority negative control
|
||||
|
||||
`test_a_memory_that_contradicts_state_loses_and_changes_nothing` sets up the
|
||||
conflict like this:
|
||||
- a state correction adds the fact "the tavern lamp is lit";
|
||||
- a hand-written memory says "The tavern lamp was never lit that night.";
|
||||
- the memory is pinned, so it is injected;
|
||||
- a turn is played with an empty proposal.
|
||||
|
||||
| Assertion | Result |
|
||||
| --- | --- |
|
||||
| The narrative state document is unchanged by retrieval and the turn | **yes** (identical before and after) |
|
||||
| The state fact is in the prompt's `narrative_state` section | yes |
|
||||
| The memory is in `used_memories`, under "Memories from earlier in the story …" | yes, framed as historical and non-canon context |
|
||||
| `narrative_state` comes after `used_memories`, so state is read last and settles the conflict | yes |
|
||||
|
||||
F07 semantics are unchanged: memory never writes state.
|
||||
|
||||
---
|
||||
|
||||
## K. Real-model attempts
|
||||
|
||||
**Setup.**
|
||||
- **Command:** `tools/m11_long_run.py --turns 100 --independent-fact`.
|
||||
- **Host and models:** the GPU inference host (Ollama 0.34.0), `qwen2.5:3b-instruct-16k` at a verified 16,384 window, embeddings by `nomic-embed-text:latest`.
|
||||
- **Memory:** the memory bank and auto-summarise on.
|
||||
- **Harness settings:** `memory_top_k` 4 and `context_token_budget` 16,384.
|
||||
- **Logging:** the owner's power, link and kernel logging was running on the host before the run started.
|
||||
- **Permission:** inference was used only with the owner's explicit approval.
|
||||
|
||||
**Evidence:**
|
||||
- `$HOME/v11-evidence/b1/real-1/`: `summary.json`, `recall-independent.json`, `timeline.jsonl` and `campaign.db`;
|
||||
- `$HOME/v11-evidence/b1/real-1-diagnosis/diagnosis.json`: the four stages, recomputed on a copy of the database with the same embedding model.
|
||||
|
||||
### K.1 Attempt 1 — PRECONDITION FAILED
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Run status | `complete`: 102 accepted turns (100 requested), 3 restarts (4 process starts), 12 summaries, 33 active memories (19 eligible on the recall line), 1,308 s elapsed |
|
||||
| Window / accounting | verified 16,384 on every turn; `fits` |
|
||||
| M04 verdict (unchanged) | `recovered_through_state_only` |
|
||||
| Independent fact | planted at depth 3 by the player turn *"I watch Mara slip the amber sundial inside the cracked teapot on the tavern's top shelf, and she makes me promise to tell no one."* It was never corrected into state |
|
||||
| **Verdict** | **`precondition_failed:absent_from_summary`** (the verdict names the first failed precondition in its fixed order) |
|
||||
|
||||
| Precondition | Result | First failure |
|
||||
| --- | --- | --- |
|
||||
| planted turn outside recent history | **held**: the history window started at depth 66 | — |
|
||||
| absent from authoritative state | **held**: the document, every node's snapshot and the recall prompt's state section | — |
|
||||
| absent from imported knowledge | **held**: 3 sources | — |
|
||||
| **absent from later narration** | **failed**: the narrator mentions the fact at depths 6, 8, 10, 14, 16, 18, 20, 22, 24, 26 and later | accepted turn **5** |
|
||||
| **absent from the active summary** | **failed** | accepted turn **9** |
|
||||
| fact text in the recall prompt's history sections | **present** (restated narration inside the window) | — |
|
||||
|
||||
**Why it did not qualify.** The 3B narrator took the planted detail up as a motif
|
||||
and restated it for the rest of the campaign ("Mara's amber sundial flickered
|
||||
softly, a silent reminder of their shared history"). The summariser, whose prompt
|
||||
asks it to "preserve important established facts", folded it into the running
|
||||
summary. Neither is a defect in the harness. They are the other layers doing what
|
||||
they do with a salient fact, which is exactly what makes a clean memory-only
|
||||
measurement hard to obtain with a real narrator.
|
||||
|
||||
**The four stages, diagnosed anyway.** They are not evidence for independent
|
||||
retention, but they are evidence for the mechanism.
|
||||
|
||||
| Stage | Result |
|
||||
| --- | --- |
|
||||
| **created** | **yes.** Memory 1 covers depths 0-5. The block is 688 tokens, so there was no truncation: the fact was in the block and in the summariser's excerpt. The memory text: *"Aldric taps the silver key against his chest … Mara slipped the amber sundial into the cracked teapot, her fingers tightening on the silver key. …"* |
|
||||
| **retained** | **yes.** Not forgotten, `use_count` 53, 33 active memories against capacity 80, eviction position 28 |
|
||||
| **ranked** | **no.** Eligible and embedded. For the recall turn's own query it scored `similarity` **0.866** and ranked **10 of 19**, against `top_k_cutoff` 4. It was not selected and was not suppressed as a duplicate. The replica matched the stored selection |
|
||||
| **injected** | not reached. The recall turn's `memories.used` = [30, 29, 31, 5] |
|
||||
|
||||
**What the narrator was given instead.** Five **later** memories also carry the
|
||||
fact, all written from narration that restated it: memories 12 (depths 66-71), 28
|
||||
(81-86), 30 (93-98), 17 (96-101) and 18 (102-107). Memory 30 was injected at
|
||||
recall. The fact reached the narrator through memory, but through a
|
||||
**restatement's** memory, not the planting-era one.
|
||||
|
||||
**The same memory 1 against reference queries** (`nomic-embed-text`):
|
||||
|
||||
| Query | Rank | Similarity | Selected |
|
||||
| --- | --- | --- | --- |
|
||||
| the recall turn's production query | 10 / 19 | 0.866 | no |
|
||||
| paraphrase: "… the little brass dial that tells the hour" | 6 / 19 | 0.593 | no |
|
||||
| unrelated: "… what rope costs at the landing this season" | 12 / 19 | 0.435 | no |
|
||||
|
||||
Two properties of the real embedder matter to B.2:
|
||||
1. **A high floor.** An unrelated query still scores 0.44.
|
||||
2. **A crowded top.** The bank is full of near-identical "Aldric and Mara step out,
|
||||
the silver key's weight in his pocket" memories, so 0.866 was not enough to reach
|
||||
the top 4.
|
||||
|
||||
*A side finding outside B.1's scope.* The export's protocol-leak counter flags
|
||||
**1 of 105** stored turns: action 153, depth 143. Mid-reply, the narrator echoed the
|
||||
length hint and the state reminder, with an `Events: [...]` line, and then continued
|
||||
the story. WP-A2's cleanup rules act only at the end of a reply, so an instruction
|
||||
echo with story after it stays. It is recorded here for the v1.1 backlog. B.1 did not
|
||||
touch it.
|
||||
|
||||
---
|
||||
|
||||
## L. First failing stage
|
||||
|
||||
**Deterministic, with a best-case summariser and a concept embedder, identical on
|
||||
v1.0.0 and HEAD:**
|
||||
|
||||
| Condition | First failing stage |
|
||||
| --- | --- |
|
||||
| a bank under capacity, blocks under 2,000 tokens | **none**: created, retained, ranked (2 of 17) and injected |
|
||||
| a bank past `memory_bank_capacity` | **retention**: evicted by least-recently-used order after recent narration stops retrieving it (§F) |
|
||||
| the fact early in a block over 2,000 tokens | **creation**: the fact is in the block, but not in the summariser's excerpt (§E) |
|
||||
|
||||
**Real model, attempt 1** (isolation not met, mechanism only): created yes, retained
|
||||
yes, **ranking failed first**. The planting-era memory scored 0.866 but ranked 10th of
|
||||
19 behind later, near-identical memories, and outside `top_k` 4.
|
||||
|
||||
**Across all the evidence, the first stage an early fact fails at a real campaign's
|
||||
length is ranking.** In both the deterministic default and the real run, the memory
|
||||
exists and is retained at 100 turns, where the bank is below capacity. What decides
|
||||
whether the narrator is shown it is its rank against a recent-narration query in a
|
||||
bank of similar memories.
|
||||
- **Retention (eviction)** is a second, later failure. It is certain once a campaign
|
||||
outgrows capacity: about 480 actions at the defaults.
|
||||
- **Creation (truncation)** is a third, conditional one. It needs blocks longer than
|
||||
2,000 tokens, which the 16,384 window with 500-token replies did not produce
|
||||
(688 tokens).
|
||||
|
||||
---
|
||||
|
||||
## M. Evidence for likely root cause
|
||||
|
||||
1. **The retrieval query is recent narration, not the question** (§B.9, §G). The
|
||||
query is the last 4 actions cut to 600 tokens, so a one-line recall question is
|
||||
outweighed by three turns of prose. The effect is deterministic: the question's
|
||||
own similarity of 0.708 fell to 0.241 in the production query. In the real run,
|
||||
recent prose about the same tavern, key and people made every memory look
|
||||
similar, and ten ranked above the planting-era one.
|
||||
2. **Ranking has no term that favours the planting-era record** (§B.9;
|
||||
CONTEXT-AND-MEMORY §20). There is cosine only: no lexical match on the question's
|
||||
rare terms ("sundial", "teapot"), no importance, and no preference for the earliest
|
||||
or a coverage-distinct source. Later restatement memories carry the same words in
|
||||
more familiar company, and outrank the original.
|
||||
3. **Eviction is purely least-recently-used** (§B.7, §F). Retrieval is driven by
|
||||
recent narration, so exactly the facts nothing recent mentions lose recency and go
|
||||
first. Recall itself cannot rescue them, because they are no longer retrieved.
|
||||
4. **The summariser reads only the last 2,000 tokens of a block** (§B.6, §E). This is
|
||||
proven deterministically. It did not bite at the real run's block sizes.
|
||||
5. **v1's M04 never tested memory** (§B finding). The clue was planted as state, so
|
||||
the long-standing "memory does not keep the fact" observation was never a
|
||||
measurement of memory.
|
||||
|
||||
---
|
||||
|
||||
## N. What B.2 is allowed to change
|
||||
|
||||
B.2 is allowed only the smallest changes the evidence supports, one mechanism at a
|
||||
time, each with a failing test first. In order of the evidence:
|
||||
|
||||
1. **Ranking** (the first failing stage). Candidates:
|
||||
- build the retrieval query so the player's newest input is not drowned out, for
|
||||
example by giving the newest player action its own weight or its own query;
|
||||
- and/or add one inspectable ranking term from CONTEXT-AND-MEMORY §20, most
|
||||
directly a lexical match on the query's rare terms.
|
||||
|
||||
Either must keep `replica_matches_stored_selection` meaningful: a
|
||||
diagnostic-visible score, recorded in `memories.used`.
|
||||
2. **Eviction** (the certain second failure). Stop least-recently-used eviction from
|
||||
discarding a never-again-retrieved early memory first. For example, weight eviction
|
||||
by coverage, keeping the only memory of a story range, or by age, instead of recency
|
||||
alone. The frozen-bank protection must be kept.
|
||||
3. **Creation** (conditional). Choose the summariser's excerpt so that a fact early in
|
||||
a long block is not cut. For example, the head and tail, or the whole block up to a
|
||||
larger bound.
|
||||
|
||||
Each change turns one of B.1's diagnostics into a passing result:
|
||||
- `past_capacity` and `long_block_fact_early` flip their strict xfails;
|
||||
- a real-model re-run shows `ranked: yes` for the planting-era memory.
|
||||
|
||||
## O. What B.2 must not change
|
||||
|
||||
- **Lineage safety.** `tree.attach_memory`, the path clause, `forget_node` and E02
|
||||
stay as they are. An abandoned line's memory stays stored and ineligible (§I).
|
||||
- **Summary lineage** (E03) and summary content policy.
|
||||
- **Authority.** Memory never writes state and is never framed as canon (F07, §J).
|
||||
- **Imported-knowledge authority and retrieval.**
|
||||
- **Pins.** Pinned memories stay always-selected and never evicted.
|
||||
- **The frozen-bank fix.** A memory is never evicted by the pass that created it.
|
||||
- **The single-commit use counter** (M11 §O.7).
|
||||
- **F01-F08, E01-E04, the M04 verdicts, the bundle format and the schema**, unless a
|
||||
migration is separately justified.
|
||||
- **The deterministic diagnostic itself.** B.2 flips the strict xfails. It does not
|
||||
weaken the scenarios or the isolation checks.
|
||||
|
||||
---
|
||||
|
||||
## P. Tests / regression
|
||||
|
||||
| Run | Result |
|
||||
| --- | --- |
|
||||
| `test_v11_b1_memory_diagnostic.py` on HEAD | **24 passed, 2 xfailed (strict)** |
|
||||
| `test_v11_b1_memory_diagnostic.py` on v1.0.0 (§D) | **24 passed, 2 xfailed (strict)**, identical |
|
||||
| `test_v11_b1_long_run_verdict.py`, `test_m11_long_run_memory.py`, `test_m11_long_run_resume.py` | **52 passed** |
|
||||
| **Full backend suite**, HEAD plus the B.1 files | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,323 s). The 17 skips are the tests that need a real model, the same 17 as before. The 2 strict xfails are the two diagnosed retention criteria (§D). |
|
||||
| **Full backend suite, final re-run at staging** (after the real-model attempt; the staged tree) | **1,578 passed, 17 skipped, 2 xfailed, 0 failed** (1,912 s), identical |
|
||||
|
||||
No frontend file was changed, so the frontend suite, lint and build are not affected.
|
||||
|
||||
---
|
||||
|
||||
## Q. Compatibility
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| Application code changed | **none.** `git diff --stat HEAD -- backend/app frontend` is empty |
|
||||
| Schema migration | none |
|
||||
| Bundle format | unchanged |
|
||||
| Stored campaign behaviour | unchanged |
|
||||
| Memory behaviour | **unchanged.** Creation, eviction, ranking, pins and lineage are all as in v1.0.0. `memorybank.py` is byte-identical to the tag |
|
||||
| What changed | Diagnostic tooling (`tools/memory_diagnostic.py`, `tools/v11_b1_memory.py`), an opt-in harness mode (`m11_long_run.py --independent-fact`, with every existing behaviour and M04 verdict unchanged when the flag is off) and tests |
|
||||
| New strict xfails | 2 (§D). They document the two diagnosed defects, and the suite stays green. B.2 must remove them deliberately when it fixes the mechanisms |
|
||||
|
||||
## R. Security / local-only
|
||||
|
||||
Checked against the B.1 diff: the `m11_long_run.py` changes plus the four new files.
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| New network client, endpoint, URL or TLS setting | **none.** The only address in the new code is `http://127.0.0.1:9/v1`, a refused loopback port the deterministic scenarios configure so nothing is contacted |
|
||||
| External embedding service or remote vector store | **none.** The deterministic runs use `ConceptEmbedder` in-process. The real-model run uses the configured Ollama embedding model through the application's existing provider |
|
||||
| Endpoint policy (`endpoints.py`, ADR 011) and TLS (`tlstrust.py`) | unchanged; not in the diff |
|
||||
| Network calls in the real-model run | only the configured Ollama host on the trusted LAN, through the application's own provider and probe paths |
|
||||
| Real identifiers in committed files | none. Checked again at staging |
|
||||
| Offline container regression | **not run.** The harness builds and runs a Docker container, and this session's permission policy refused it. B.1 changes no runtime code and nothing in the image (the image carries `backend/app` and `frontend/dist` only), so the container would be byte-identical to the A1/A2 image. That image passed 23/23 on the corrective tree (`V1.1-WP-A1-A2-REPORT.md` §R.3) |
|
||||
|
||||
---
|
||||
|
||||
## S. Final decision
|
||||
|
||||
**What was established.** Every B.1 requirement was carried out except the clean
|
||||
real-model run, which was attempted:
|
||||
- the current mechanism (§B);
|
||||
- a deterministic, isolated harness with four stage outputs and a verdict at recall
|
||||
depth ≥ 100 (§C);
|
||||
- the v1.0.0 baseline (§D);
|
||||
- capacity and eviction (§F), the creation window (§E) and ranking (§G);
|
||||
- the lineage and authority negative controls (§I, §J);
|
||||
- the `recovered_through_memory_independent` long-run verdict, with the M04 verdicts
|
||||
unchanged;
|
||||
- tests and regression (§P), compatibility (§Q) and security (§R).
|
||||
|
||||
**The real-model gap.** The one real-model attempt ran cleanly, but failed isolation
|
||||
(`precondition_failed:absent_from_summary`). The narrator and summariser restated the
|
||||
fact, so a clean memory-only result on a real model was **not obtained**. The owner
|
||||
decided not to make a second attempt, because the same narrator behaviour would very
|
||||
likely repeat. The attempt is reported in full, and its stage diagnosis is used as
|
||||
mechanism evidence only (§K).
|
||||
|
||||
**B.1 DIAGNOSTIC: COMPLETE**
|
||||
|
||||
**FIRST FAILING STAGE: RANKING.** In a real 100-turn campaign the early fact's memory
|
||||
was created (it carries the fact) and retained (active, 33 of 80), but ranked 10th of
|
||||
19 (similarity 0.866) against `top_k` 4. Later memories with near-identical wording
|
||||
outranked it, under a query made of recent narration (§K, §L). Two further failures
|
||||
are proven deterministically, identically on v1.0.0:
|
||||
- **retention**, past `memory_bank_capacity`: least-recently-used eviction removes an
|
||||
early memory first;
|
||||
- **creation**, for a fact early in a block over 2,000 tokens: the summariser's
|
||||
last-2,000-token excerpt drops it.
|
||||
|
||||
**WP-B.2 RECOMMENDED CHANGE: retrieval ranking first.** Build the retrieval query so
|
||||
the newest player input is not diluted by three turns of recent narration, and add one
|
||||
inspectable lexical term for the query's rare words to cosine ranking. The score must
|
||||
be recorded in `memories.used`. Acceptance: a real-model re-run shows `ranked: yes`
|
||||
for the planting-era memory.
|
||||
|
||||
Then, in separate test-first steps:
|
||||
1. make eviction coverage-aware instead of purely least-recently-used, so the only
|
||||
memory of an early range is not discarded first (flips
|
||||
`test_acceptance_an_early_fact_is_recalled_from_memory_past_capacity`);
|
||||
2. choose the summariser excerpt so a fact early in a long block is kept (flips
|
||||
`test_acceptance_a_fact_early_in_a_long_block_is_remembered`).
|
||||
|
||||
Everything in §O stays unchanged. B.2 has not been started.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,475 @@
|
||||
# v1.1 WP-C — Browser Release Coverage
|
||||
|
||||
**Status:** COMPLETE, staged for owner review. The final run passed 91 checks with 0 failed and 0 skipped. The decision is in §R.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD at start | `0c1ba836babe1447ad3693b4d95189a325b7b3b6` — *v1.1 WP-B.2: independent long-term memory retention*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||||
| Its ancestry | `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `ac465ed` (plan v4.1), `432f041` (v1.0.0) |
|
||||
| Working tree at start | clean; nothing staged |
|
||||
| WP-D, WP-E | not started |
|
||||
|
||||
---
|
||||
|
||||
## B. Existing browser harness
|
||||
|
||||
`backend/tools/m11_browser.py` drives Firefox through `backend/tools/m11_webdriver.py`, a
|
||||
dependency-free W3C WebDriver client. It builds its fixture campaign through the API,
|
||||
plays two narrator turns through the streaming endpoint, and then asserts on the
|
||||
rendered DOM. The M11 closeout run (`3652dc6`, run 3) passed **38/38**, 0 failed,
|
||||
0 skipped, on snap Firefox 155.0.1 with geckodriver 0.37.1 and `qwen2.5:3b-instruct`
|
||||
over trusted-LAN HTTPS.
|
||||
|
||||
What it did not do, and WP-C closes: drive Retry and takes, Save Point create and
|
||||
restore, state correction, narration length or failed generation through the UI,
|
||||
and prove an export leaves the browser as a file.
|
||||
|
||||
Before WP-C it also slept, for a fixed time, before several assertions: after Undo
|
||||
and Redo, after planting hostile narration, after choosing a knowledge file, around
|
||||
the delete dialog, and around opening a panel. §J records that as a harness defect.
|
||||
|
||||
---
|
||||
|
||||
## C. Browser and download environment
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Firefox | **155.0.1, the snap** (`/snap/bin/firefox`). No second Firefox was installed |
|
||||
| geckodriver | 0.37.1 (the snap) |
|
||||
| Headless | yes |
|
||||
| Frontend | the production build (`vite build`) served by FastAPI on loopback; no Vite dev server |
|
||||
| Certificates | `acceptInsecureCerts: false`, unchanged |
|
||||
|
||||
**The download profile.** `m11_webdriver.firefox_download_prefs` is passed as
|
||||
`moz:firefoxOptions.prefs`:
|
||||
- `browser.download.folderList` 2, `browser.download.dir` `<--out>/downloads`,
|
||||
`browser.download.useDownloadDir` true;
|
||||
- `browser.download.start_downloads_in_tmp_dir` false;
|
||||
- `browser.download.always_ask_before_handling_new_types` false;
|
||||
- `browser.helperApps.neverAsk.saveToDisk` `application/json,application/octet-stream`;
|
||||
- the download panel suppressed.
|
||||
|
||||
**The folder.** `--out/downloads`, which must be under `$HOME`
|
||||
(`require_under_home`). The harness deletes it at the start of a run and creates it
|
||||
fresh. Evidence lives under `$HOME/v11-evidence/wp-c/`.
|
||||
|
||||
**Does the snap Firefox download?** Yes. Measured first, with a probe
|
||||
(`$HOME/v11-evidence/wp-c/probe/`): a loopback page runs the product's own download
|
||||
pattern (a JSON blob, an `<a download>` click, an immediate `revokeObjectURL`).
|
||||
- Download folder under `~/v11-evidence`: 33-byte file written.
|
||||
- Download folder under `~/Downloads`: 33-byte file written.
|
||||
|
||||
The first probe wrote nothing, and that was a probe defect (J1), not the snap. The
|
||||
non-snap fallback the plan allows was therefore not needed.
|
||||
|
||||
**When a download counts as finished** (`m11_webdriver.wait_for_download`, tested
|
||||
without a browser in `test_v11_c_browser_helpers.py`, 7 tests). All of these at once:
|
||||
- a name absent from the listing taken before the click;
|
||||
- no `*.part` file in the folder;
|
||||
- more than zero bytes;
|
||||
- the same size across 3 consecutive polls.
|
||||
|
||||
A zero-byte, partial, pre-existing or still-growing file never counts, and neither
|
||||
does the "Campaign exported." toast.
|
||||
|
||||
---
|
||||
|
||||
## D. Retry scenario
|
||||
|
||||
Real narration: **yes**. All checks use the reader-facing controls on the play page.
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| Type in "What you do next", press **Send** | a new narration renders and the page is idle; its exact text is recorded | PASS |
|
||||
| — | **Retry** is offered (enabled) on the newest narration | PASS |
|
||||
| Press **Retry** | the take indicator on the newest narration reads **2/2** | PASS |
|
||||
| — | the second take's text differs from the first (otherwise 1/2 could not be told from 2/2) | PASS |
|
||||
| Press **‹** (Previous take) | the indicator reads **1/2**, and the narration is identical to the first recorded text | PASS |
|
||||
| — | the second take's text is not shown anywhere in the transcript | PASS |
|
||||
| Press **›** (Next take) | **2/2**, showing the second take | PASS |
|
||||
| Reload the page | the indicator still reads **2/2** on the live take, showing the second take | PASS |
|
||||
| Press **‹** after the reload | **1/2** still shows the first text, unchanged | PASS |
|
||||
|
||||
All text comparisons are of the rendered `.turn-text`. No database was read.
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## E. Save Point scenario
|
||||
|
||||
Real narration: **yes** (two turns after the Save Point).
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| Press **Save Point**, type a name in "Save this moment", submit | a row with that exact name appears | PASS |
|
||||
| — | the row's moment is the moment being read ("Moment 7") | PASS |
|
||||
| **Send** two turns | two new narrations render; position "Moment 11" | PASS |
|
||||
| In Save Points, press **Restore**, then confirm **Restore** | the position reads "Moment 7 · later story ahead" | PASS |
|
||||
| — | the transcript ends at the Save Point: its last narration is the one read there, and neither later narration is shown | PASS |
|
||||
| — | the position says later story is ahead | PASS |
|
||||
| — | **Redo** is enabled | PASS |
|
||||
| Press **Redo** until it is disabled (each press waits for the position to change) | the last two narrations are the two later turns, text-identical, at "Moment 11" | PASS |
|
||||
| Reload the page | the position is still "Moment 11" | PASS |
|
||||
| — | the named Save Point is still listed | PASS |
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## F. State-correction scenario
|
||||
|
||||
**Owner decision (2026-09-15).** The State panel cannot produce a *partly* refused
|
||||
correction:
|
||||
- "Save correction" sends exactly one `add_fact`, and "That's wrong" sends exactly
|
||||
one `invalidate_fact`;
|
||||
- so a correction is applied whole or refused whole (HTTP 400);
|
||||
- the route's partial application (`refused` alongside applied changes) is
|
||||
reachable only through the API.
|
||||
|
||||
WP-C drives what the reader can reach: an accepted correction that persists, and a
|
||||
refused correction whose refusal and reason are visible, with the refused change
|
||||
not applied. It records the partial refusal as unreachable from the reader UI. No
|
||||
product change was made for it.
|
||||
|
||||
Real narration: **yes** for the refusal (it needs history to step back over); the
|
||||
accepted correction needs none and also passed in the no-narrator smoke runs.
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| In **State**, press **Correct something**, type a fact, press **Save correction** | the form closes and the State panel shows the fact | PASS |
|
||||
| Reload, open **State** | the fact is still shown | PASS |
|
||||
| In a second tab press **Undo** (the correction belongs to the moment it was made at); in the first tab, whose panel still shows the fact, press **That's wrong** on it | a failure notice carrying the correction's refusal ("can't be applied") appears | PASS |
|
||||
| — | it is labelled **"That correction was not applied"**, with no "Try that turn again" and no claim that typed text was kept (§K2) | PASS |
|
||||
| Open **Show technical details** | the reason is visible: "That correction can't be applied — no fact 'f3' to invalidate." | PASS |
|
||||
| Second tab **Redo**, close it; reload the first tab at the corrected moment | the fact still stands, with its **That's wrong** control: the refused withdrawal was not applied | PASS |
|
||||
|
||||
**Partial refusal.** As the owner decided, a *partly* refused correction is not
|
||||
reachable from the reader UI, and was not driven. The route's partial application
|
||||
remains covered by the backend suite.
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## G. Narration-length scenario
|
||||
|
||||
Real narration: **yes** (one turn per band). The model's actual length is not
|
||||
asserted, because nothing in the product contract requires it. What is asserted is
|
||||
that the chosen band reached the turn's own prompt.
|
||||
|
||||
The expected sentence comes from the product's own `builder.length_hint` at the
|
||||
run's 400-token output cap:
|
||||
- **brief:** "must not exceed 180 words, and it should not stop short of about 70";
|
||||
- **long:** "must not exceed 236 words, and it should not stop short of about 118".
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| **Settings** panel: choose **Brief** in "Narration length", press **Save changes** | the button reads "Saved" | PASS |
|
||||
| **Send** a turn; choose **Long**, **Save changes**; **Send** a second turn | both turns render | PASS (both) |
|
||||
| On the brief turn press **Inspect context** | the prompt sections of "The exact text the narrator was sent" contain brief's range, and do **not** contain that turn's own reply (so this is the turn's record, not the dry run of the next) | PASS |
|
||||
| On the long turn press **Inspect context** | the same, with long's range | PASS |
|
||||
| Reload, open **Settings** | "Narration length" still reads **long** | PASS |
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## H. Failed-generation scenario
|
||||
|
||||
**Owner decision (2026-09-15).** The literal sequence (save an unserved model, then
|
||||
submit a turn) cannot be driven. The model check (`modelStatus.jsx`) marks a
|
||||
configured model absent from the endpoint's list as `missing-model`, and
|
||||
`blocksPlay` disables Send, Continue and Retry up front. That is M8's intended
|
||||
behaviour, not a defect. So WP-C drives both of these:
|
||||
1. **The unserved model** saved through Settings: the reader is told and cannot
|
||||
send, and the story is unchanged.
|
||||
2. **A submitted failure:** a model the endpoint lists but that cannot narrate,
|
||||
`nomic-embed-text:latest`. Send stays enabled, and the turn fails in the open.
|
||||
|
||||
Then recovery with the reference model. The report records that path 2 uses a
|
||||
listed, non-narrating model rather than "a name the server does not serve".
|
||||
|
||||
Real narration: **yes**.
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| **Settings**: "Type a model name instead", type an unserved name, **Save** | the page says "Saved" | PASS |
|
||||
| Open the campaign | the header model status is `missing-model` and the setup notice is shown | PASS |
|
||||
| — | **Send** and **Continue** are disabled | PASS |
|
||||
| — | the story is unchanged | PASS |
|
||||
| **Settings**: choose `nomic-embed-text:latest` from the installed-model picker, **Save** | "Saved" | PASS |
|
||||
| Open the campaign (model status `ready`); type a turn; **Send** | a failure notice is shown: "Generation failed", with the server's reason under the details, `"nomic-embed-text:latest" does not support chat` (HTTP 400) | PASS |
|
||||
| — | no narration was added | PASS |
|
||||
| — | the typed text is still in the input box | PASS |
|
||||
| — | the earlier story is text-identical | PASS |
|
||||
| Open **State** | the rendered state is identical to before the failure | PASS |
|
||||
| **Settings**: choose `qwen2.5:3b-instruct`, **Save** | "Saved" | PASS |
|
||||
| Open the campaign; type a turn; **Send** | exactly one new narration | PASS |
|
||||
| — | the earlier story is intact | PASS |
|
||||
| Reload | the successful turn is still the last narration | PASS |
|
||||
|
||||
The failed turn's player moment stays in the transcript, as A05 intends
|
||||
(`player_moment_kept_in_transcript` in the report).
|
||||
|
||||
**Deviation, owner-approved.** Path 2 uses a model the endpoint *lists* but that
|
||||
cannot narrate, not "a name the server does not serve", because an unserved name
|
||||
is caught before a turn can be submitted. Endpoint policy was not bypassed: both
|
||||
paths use the same trusted-LAN HTTPS endpoint, and only the model name changed.
|
||||
|
||||
Final run: `$HOME/v11-evidence/wp-c/final/`, `browser-report.json` (every check with its detail),
|
||||
`server.log`, `geckodriver.log`, and the downloaded files.
|
||||
|
||||
## I. Export-download scenario
|
||||
|
||||
Real narration: not needed for the download itself. In the final run the exported
|
||||
campaign holds 19 moments of real narration, two takes and a Save Point.
|
||||
|
||||
**C6a — the campaign library**
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| On the library page, press **Export** on the "Release Regression" card | a new file is written to `downloads/`; it is finished (no `.part`, stable size) | PASS |
|
||||
| — | it is not empty: **136,739 bytes** | PASS |
|
||||
| Parse the file | `format` is **`ai-dnd-adventure-v3`** | PASS |
|
||||
| Start a **fresh application** (new database); press **Import campaign**; give its file input the downloaded path | the browser lands on the imported campaign's play page | PASS |
|
||||
| Compare what the reader sees of the import with what the reader saw of the original (library card, position, Save Points panel) | same title ("Release Regression") | PASS |
|
||||
| — | same number of moments (19) | PASS |
|
||||
| — | same position ("Moment 18") | PASS |
|
||||
| — | same Save Points, which also match the file | PASS |
|
||||
|
||||
The file's `headDepth` is 17, which the play page shows as "Moment 18", and its
|
||||
`actions` count is 19.
|
||||
|
||||
**C6b — campaign settings**
|
||||
|
||||
| Browser action | Observable assertion | Result |
|
||||
| --- | --- | --- |
|
||||
| In the campaign's **Settings** panel, press **Export campaign** | a second, separate file is written and finished | PASS |
|
||||
| — | not empty: 136,739 bytes | PASS |
|
||||
| Parse the file | `format` is `ai-dnd-adventure-v3` | PASS |
|
||||
|
||||
Both files: `$HOME/v11-evidence/wp-c/final/``downloads/library-export.json` and `downloads/settings-export.json`.
|
||||
The imported application's log is `import-server.log`.
|
||||
|
||||
---
|
||||
|
||||
## J. Harness defects found
|
||||
|
||||
| # | Defect | How found | Fix |
|
||||
| --- | --- | --- | --- |
|
||||
| J1 | The download probe's page wrote `URL.createObjectURL` inside an inline `onclick`, where `URL` is `document.URL`, a string. No blob was made, so it looked exactly like "the snap cannot download" | `gecko.log`: `TypeError: URL.createObjectURL is not a function` | The probe's script uses `window.URL`, the way the product's module does. Both download folders then worked |
|
||||
| J2 | C3's first refusal withdrew a fact in a second tab and withdrew it again in the stale tab. Withdrawing keeps the fact, marked `invalidated` (C04's audit record), so the second withdrawal was valid and accepted. Nothing was refused, and "the refused change was not applied" passed without meaning anything | Smoke run: two C3 failures; `server.log` shows 201 for every correction; `narrative/apply.py` | The second tab steps the story back past the correction with Undo, so the stale tab withdraws a fact the story at that position does not have. The validator then refuses it: `no fact … to invalidate` |
|
||||
| J3 | Clicks on a control just under the play page's fixed composer were intercepted ("Show technical details" in a failure notice, a turn's "Inspect context") | Dev run 1: `element click intercepted` | `Browser.click` scrolls the element to the centre of the view, then uses the real WebDriver click |
|
||||
| J4 | The Settings model field is a text box until the endpoint's model list arrives, then a picker. Choosing before the check finished raced that swap | Dev run 1: `no such element: input#model` | Wait for the header's model status to leave `checking` first |
|
||||
| J5 | C3's "a refused correction is shown" waited for *any* failure notice, so a notice about something else would have passed | Dev run 1: it passed on a notice titled "Generation failed" (§K2) | It now requires the notice to carry this correction's refusal ("can't be applied"), and asserts how it is labelled |
|
||||
| J7 | C4 required the band's sentence in the inspector and the turn's own reply to be absent, to tell a turn's record from the dry run. But the inspector also renders what came back (`raw_output`, "What came back, before the state block was removed") inside the same section, so the absence could never hold | Dev run 2: both C4 inspector checks failed. The stored records show the brief turn's `length_hint` section carrying "must not exceed 180 words … about 70", the long turn's carrying "…236 … 118", and every turn's reply present in `raw_output` | The sentence and the absence are read from the prompt sections only, excluding the "What came back" block |
|
||||
| J6 | M11's checks slept before assertions: after Undo and Redo, after planting hostile narration (1 s), after choosing a knowledge file (0.5 s), around the delete dialog (0.8 s and 0.6 s), and around opening a panel (0.8 s). A sleep is not evidence of what it waited for | Reading the harness against the brief's rule | Each is now a wait on the condition the check needs: the position changed or returned, the planted text rendered, Import enabled, a dialog present or gone, a panel-specific element present. §38's absence check now first waits for the knowledge library to render |
|
||||
|
||||
**Checked and not a defect.** In dev run 2, the original take at depth 6 (action 9)
|
||||
had no `length_hint` in its stored record, while its Retry (action 10) did. The
|
||||
original take's record holds only the per-attempt fields
|
||||
(`attempts.ATTEMPT_KEYS`: world state, narrative state, raw output, usage,
|
||||
accounting), with no sections and no settings. That is how a take that is no
|
||||
longer live is stored, not a prompt built without the range. Every live turn's
|
||||
prompt carried its band.
|
||||
|
||||
Also added: a panel opens only if it is not already open, because a tab toggles
|
||||
its panel closed, and it is recognised by an element only that panel renders, not
|
||||
by its title, which the tab itself already shows.
|
||||
|
||||
---
|
||||
|
||||
## K. Product defects found
|
||||
|
||||
### K1 — "Correct" on an Important Facts row is always refused (not fixed)
|
||||
|
||||
The State panel offers **Correct** on every row with a key. On an Important Facts
|
||||
row the key is the fact's id, and on the scene-summary row it is `"summary"`.
|
||||
`saveCorrection` sends that key as `add_fact.subject`, and the validator checks
|
||||
`subject` as an entity reference. So every such correction is refused.
|
||||
|
||||
**Reproduced deterministically** against the real application (scratch `TestClient`,
|
||||
no browser, no model):
|
||||
|
||||
| Correction the panel sends | Result |
|
||||
| --- | --- |
|
||||
| "Correct" on the Characters row (`subject='mara'`) | **201**, applied |
|
||||
| "Correct" on the Important Facts row (`subject='f1'`) | **400** "That correction can't be applied — add_fact names subject='f1', which does not exist." |
|
||||
|
||||
**Not fixed in WP-C.** It does not prevent any WP-C behaviour: "Correct something"
|
||||
and entity-row corrections work, and C3 uses them. A fix (offer Correct only
|
||||
against entities, or send facts without a subject) is a small UX choice, left to
|
||||
the owner as a v1.1 follow-up.
|
||||
|
||||
### K2 — a refused correction was presented as a failed turn (fixed)
|
||||
|
||||
**Found in the browser** (dev run 1, §J5). When the story refused a correction, the
|
||||
failure notice said:
|
||||
- the title **"Generation failed"**;
|
||||
- the hint "Nothing was added to your story. You can try that turn again.";
|
||||
- the button **Try that turn again**;
|
||||
- the line "what you typed is still in the box below".
|
||||
|
||||
None of that is true of a State-panel correction. `classifyError` had no rule for
|
||||
the server's refusal ("That correction can't be applied — …"), so it fell through
|
||||
to the generation default. That falsified exactly what C3 checks: that a refusal
|
||||
is shown to the reader as a refusal.
|
||||
|
||||
**Fix** (frontend only, no backend change):
|
||||
- `errors.js`: one rule, checked first, for "correction can't be applied". It gives
|
||||
`kind: state`, the title "That correction was not applied", the hint "Nothing in
|
||||
the story or its state was changed. The reason is in the technical details.",
|
||||
`retryable: false` and `keptInput: false`.
|
||||
- `FailureNotice.jsx`: the "what you typed" line is shown unless a failure says
|
||||
`keptInput: false`. Every other kind still shows it, so M8's A05 contract is
|
||||
unchanged.
|
||||
|
||||
**Regression** (`failurePaths.test.jsx`, 3 tests):
|
||||
- a refused correction classifies as a state refusal, not retryable, with the
|
||||
reason kept;
|
||||
- its notice shows the reason and no "Try that turn again" or typed-input claim;
|
||||
- a failed turn still claims the typed words were kept.
|
||||
|
||||
In the browser, C3 now asserts the label, the absence of "Try that turn again" and
|
||||
the absence of the typed-input line.
|
||||
|
||||
Frontend after the fix: **168/168** tests, lint exit 0 (15 pre-existing warnings, 0
|
||||
errors, none in changed files), production build passes.
|
||||
|
||||
---
|
||||
|
||||
## L. Existing 38-check regression
|
||||
|
||||
**38/38 passed, 0 failed, 0 skipped** in the same run, tagged `M11`. The names are
|
||||
identical to the M11 closeout run, so there is a one-to-one mapping and no check was
|
||||
split, merged or dropped:
|
||||
- B01 a turn is accepted (×2);
|
||||
- A/UX: the tab title (×3);
|
||||
- B position indicator (×3), D01 Undo, D04 Redo (×2);
|
||||
- H06 (×3), H07, G09, H04;
|
||||
- G01 knowledge import (×2);
|
||||
- A11y dialog focus (×4);
|
||||
- §38 narrator-only text absent;
|
||||
- F05 context inspector;
|
||||
- H10 (×2), H11/CSP (×2);
|
||||
- A11y names, focus, tabindex, hover, contrast (×4), input focus.
|
||||
|
||||
What changed in them is how they wait (J6), not what they assert.
|
||||
|
||||
## M. New WP-C checks
|
||||
|
||||
**53/53 passed, 0 failed, 0 skipped**, tagged `WP-C`:
|
||||
|
||||
| Scenario | Checks | Result |
|
||||
| --- | --- | --- |
|
||||
| C1 Retry | 8 | 8 PASS |
|
||||
| C2 Save Point | 9 | 9 PASS |
|
||||
| C3 State correction | 6 | 6 PASS |
|
||||
| C4 Narration length | 5 | 5 PASS |
|
||||
| C5 Failed generation | 14 | 14 PASS |
|
||||
| C6a Library export and import | 8 | 8 PASS |
|
||||
| C6b Settings export | 3 | 3 PASS |
|
||||
|
||||
```text
|
||||
existing M11 checks: 38/38
|
||||
WP-C new checks: 53/53
|
||||
failed: 0
|
||||
skipped: 0
|
||||
```
|
||||
|
||||
**Runs that are not the final evidence**, kept under `$HOME/v11-evidence/wp-c/`:
|
||||
|
||||
| Run | Where | Result | Why it is not evidence |
|
||||
| --- | --- | --- | --- |
|
||||
| `probe/` | loopback page | first: no download (J1); second: both folders written | environment probe |
|
||||
| `smoke-1` | no narrator | 44 passed, 2 failed (J2), 6 skipped | partial |
|
||||
| `smoke-2` | no narrator | 43 passed, 0 failed, 7 skipped | partial |
|
||||
| `dev-gpu-1` | GPU host, plain HTTP | 71 passed, 3 failed (J3, J4; J5 found) | development, and not HTTPS |
|
||||
| `dev-gpu-2` | GPU host, plain HTTP | 89 passed, 2 failed (J7) | development, and not HTTPS |
|
||||
| `dev-gpu-3` | GPU host, `--only length` | 7 passed | development, and a subset |
|
||||
|
||||
## N. Production build / frontend verification
|
||||
|
||||
| | Result |
|
||||
| --- | --- |
|
||||
| Frontend suite (`npm test`) | **168/168**, 14 files. It was 165 before; the 3 new tests are K2's regressions |
|
||||
| Lint (`npm run lint`, oxlint) | **exit 0: 0 errors**, 15 warnings. All are the pre-existing `only-export-components` kind, and none is in a file WP-C changed |
|
||||
| Production build (`npm run build`) | passes; `dist/index.html` sha256 `62b6ea5eb4ce02a09f23cc1d48c335c2ada36b208b23c237338b39bf63a26cc5`, built 2026-09-15 15:11 from the final WP-C tree |
|
||||
| What the browser ran | that build, served by FastAPI (`uvicorn app.main:app` on `127.0.0.1`); no Vite server |
|
||||
| Backend product code | **unchanged**: nothing under `backend/app` is in the diff. The full backend suite was not rerun. The changed harness is tested by `test_v11_c_browser_helpers.py` (7 passed) |
|
||||
| Offline regression | **23 passed, 0 failed** (`tools/m11_offline.py`, fresh `--no-cache` image, `--network none`, on the final WP-C tree), rerun because the frontend bundle changed (K2). Evidence: `$HOME/v11-evidence/wp-c/offline/` |
|
||||
|
||||
## O. Trusted-LAN / security
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Narrator | `qwen2.5:3b-instruct` on the CPU reference host, over **trusted-LAN HTTPS**. The certificate is from the private CA in this machine's trust store, and verified; there is no bypass |
|
||||
| Window and A1 | every narrator turn (8): window **verified at 4,096**, accounting **`fits`**. None `exceeded` or `truncation_suspected` |
|
||||
| Protocol echoes (A2) | none of the protocol shapes the harness looks for (a state fence, a hard-limit or reminder bracket, `Events: [`) appeared in any stored narration in this run. This is not a v1.1 protocol-leak result: the release gate owns that, and the mid-reply echo from WP-B.1 remains a separate residual |
|
||||
| Storyteller | loopback only, both applications (the original and the fresh import) |
|
||||
| Browser | `acceptInsecureCerts: false`; the CSP checks (H11) passed |
|
||||
| Endpoint policy | unchanged; C5 changed only the model name, never the endpoint |
|
||||
| Downloads | written only under `$HOME` (enforced); the harness makes no network request of its own beyond loopback and the configured endpoint |
|
||||
| New dependencies | none: the harness still uses only `urllib`, no Selenium or Playwright |
|
||||
| Real identifiers in committed files | none (scanned at staging) |
|
||||
|
||||
## P. Compatibility
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Database schema, migrations | none |
|
||||
| Bundle format | none: the downloads are `ai-dnd-adventure-v3` and import unchanged |
|
||||
| History, Save Points, state semantics, memory, knowledge | none. WP-C drove them and changed nothing in them |
|
||||
| Backend | no application code changed |
|
||||
| Frontend | one behaviour change (K2): the refusal of a State-panel correction is labelled as a refusal rather than as a failed turn. Every other failure's classification, retry offer and typed-input claim is unchanged (M8's tests pass) |
|
||||
| WP-A, WP-B | untouched |
|
||||
|
||||
## Q. Residual risks
|
||||
|
||||
| # | Risk |
|
||||
| --- | --- |
|
||||
| 1 | **K1**: "Correct" on an Important Facts or scene-summary row is always refused. Reproduced, not fixed: a small UX choice for the owner |
|
||||
| 2 | **Partial refusal is API-only.** The reader UI cannot produce a partly refused correction (owner decision). The display for one exists but is unreachable from the panel's own controls |
|
||||
| 3 | **C5 path 2 depends on the reference host listing an embedding model.** On a host without one, the submitted-failure path has no model to use |
|
||||
| 4 | **Model nondeterminism in C1.** "The second take is a different narration" would fail if the model returned identical text for a retry. It did not in any run |
|
||||
| 5 | **The harness runs on this machine's snap Firefox.** A different Firefox or a Chromium would need the download preferences re-checked |
|
||||
| 6 | **The heuristic protocol-shape scan is not A2 evidence.** It is recorded only |
|
||||
| 7 | The WP-B real-model memory limitation and the doubled-full-stop scene text are unchanged, and are not WP-C's |
|
||||
|
||||
## R. Final decision
|
||||
|
||||
```text
|
||||
RETRY: PASS
|
||||
SAVE POINT: PASS
|
||||
STATE CORRECTION: PASS
|
||||
NARRATION LENGTH: PASS
|
||||
FAILED GENERATION: PASS
|
||||
EXPORT DOWNLOAD — LIBRARY: PASS
|
||||
EXPORT DOWNLOAD — SETTINGS: PASS
|
||||
|
||||
EXISTING BROWSER REGRESSION: PASS
|
||||
WP-C NEW BROWSER COVERAGE: PASS
|
||||
|
||||
WP-C OVERALL:
|
||||
PASS
|
||||
```
|
||||
|
||||
**PASS**, on these grounds:
|
||||
- the final run ended with **failed: 0, skipped: 0**;
|
||||
- both export controls produced a real, finished, non-empty `ai-dnd-adventure-v3`
|
||||
file on disk;
|
||||
- the library file imported into a fresh application with the same title, moment
|
||||
count, position and Save Points.
|
||||
|
||||
Two criteria were met in the form the owner approved (2026-09-15):
|
||||
- **State correction:** an accepted correction, and a refused correction with its
|
||||
reason. Partial refusal is recorded as unreachable from the reader UI.
|
||||
- **Failed generation:** the up-front block for an unserved model, plus a submitted
|
||||
failure with a listed model that cannot narrate.
|
||||
|
||||
**Product change:** K2, a mislabelled refusal, fixed narrowly with regression tests.
|
||||
**Product defect left open:** K1.
|
||||
|
||||
Nothing is committed, pushed or tagged. WP-D and WP-E have not started.
|
||||
@@ -0,0 +1,526 @@
|
||||
# v1.1 WP-D — Recovery Honesty
|
||||
|
||||
**Status:** COMPLETE — **PASS**. The decision, and what was deliberately not
|
||||
claimed, is in §P.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD at start | `59b5ebc2d85bdf53d38b6bcf347c496dd3432be1` — *v1.1 WP-C: browser release coverage*, signed by the owner (good signature, RSA key `02C9BF7D…`) |
|
||||
| Its ancestry | `0c1ba83` (WP-B.2), `beb17ad` (WP-B.1), `d63804f` (WP-A1/A2), `432f041` (v1.0.0) |
|
||||
| Working tree at start | clean; nothing staged |
|
||||
| WP-E | not started when WP-D was implemented |
|
||||
|
||||
---
|
||||
|
||||
## B. Existing backup behaviour
|
||||
|
||||
`backend/app/backup.py`, as v1 shipped it. The six questions the brief asks:
|
||||
|
||||
| # | Question | Answer (before WP-D) |
|
||||
| --- | --- | --- |
|
||||
| 1 | How is the backup file created? | SQLite's **online backup API** (`sqlite3.Connection.backup`, `pages=-1`), from the live database opened read-only through a `mode=ro` URI. The destination is a temporary file `….db.partial` **in the destination directory**, so the rename below is atomic |
|
||||
| 2 | Where does validation happen? | `_verify()`, on the finished copy, opened as its own read-only connection — not on the source, and not through the connection that wrote it |
|
||||
| 3 | When does the final filename appear? | Only after verification: `os.replace(working, target)`. A failed or interrupted run never leaves a file wearing a backup's name |
|
||||
| 4 | How are failures cleaned up? | `_discard()` removes the partial file, and `BackupError` is raised with what went wrong. The source is untouched |
|
||||
| 5 | Could an existing good backup be overwritten? | **No.** `_unused_name()` stamps each backup with the time and adds a counter if that name (or its `.partial`) exists |
|
||||
| 6 | Was the check on the source or the copy? | The copy |
|
||||
|
||||
So the mechanism was already right. The one gap was the **strength** of the check:
|
||||
`PRAGMA quick_check`, which reads every page and every record but skips the
|
||||
cross-check between a table and its indexes.
|
||||
|
||||
---
|
||||
|
||||
## C. `integrity_check` implementation
|
||||
|
||||
One function changed:
|
||||
|
||||
```python
|
||||
# app/backup.py, _verify()
|
||||
rows = connection.execute("PRAGMA integrity_check").fetchall() # was quick_check
|
||||
```
|
||||
|
||||
- It still runs on the **finished copy**, opened as its own read-only
|
||||
connection, before the rename.
|
||||
- A failure still raises `BackupError` naming what was wrong, still discards the
|
||||
partial file, and still leaves the source and every earlier backup untouched.
|
||||
- The returned `integrity` field, the filename, the directory, the reported
|
||||
fields (`filename`, `bytes`, `pages`, `seconds`, `integrity`) and the API are
|
||||
unchanged.
|
||||
- The module docstring and `_verify`'s docstring now record why the trade M9
|
||||
made (speed over the index cross-check) was not needed at these sizes, with
|
||||
the measurements in §E.
|
||||
|
||||
**No scheduled backups were added.** There is still no restore endpoint, no
|
||||
retention policy and no timer: a backup happens when the reader asks for one.
|
||||
|
||||
---
|
||||
|
||||
## D. Corruption fixture
|
||||
|
||||
`tests/test_v11_d_recovery.py: build_corrupt_copy()`.
|
||||
|
||||
A small database with `t(id, k, filler)` and an index `i_t_k ON t(k)`, 400 rows
|
||||
with keys `k000000`…`k000399`. `dbstat` names the index's own leaf pages, and one
|
||||
digit inside one indexed key on the first of them is changed (`k000144` →
|
||||
`k000944`). Every page stays structurally sound and every record still parses:
|
||||
what is broken is only the agreement between the index and its table.
|
||||
|
||||
**Proved before it is used as evidence**
|
||||
(`test_the_fixture_is_the_difference_between_the_two_checks`):
|
||||
|
||||
| Pragma | Result |
|
||||
| --- | --- |
|
||||
| `PRAGMA quick_check` | **`ok`** |
|
||||
| `PRAGMA integrity_check` | **`row 145 missing from index i_t_k`** |
|
||||
|
||||
That is the difference the package rests on: not a database both checks reject,
|
||||
but one the old check called healthy.
|
||||
|
||||
---
|
||||
|
||||
## E. Backup performance measurements
|
||||
|
||||
Each pragma was run in its **own fresh process**, alternating, three times, because
|
||||
a first measurement warms the page cache and a naive ordering makes whichever
|
||||
check runs second look faster. (It did: an early single-pass measurement showed
|
||||
`integrity_check` at 1.9 ms against `quick_check` at 65.5 ms, purely from cache
|
||||
warmth.)
|
||||
|
||||
**The real campaign database** — the M11 100-turn evidence campaign,
|
||||
`m04-final/campaign.db`, 2,367,488 bytes (2.26 MB):
|
||||
|
||||
| | Run 1 | Run 2 | Run 3 |
|
||||
| --- | --- | --- | --- |
|
||||
| `quick_check` | 5.0 ms | 3.4 ms | 7.6 ms |
|
||||
| `integrity_check` | 3.5 ms | 5.7 ms | 5.3 ms |
|
||||
|
||||
At this size the two are indistinguishable. A whole backup through
|
||||
`backup.create()` — copy, full check and rename — took **70.8 ms** and **27.4 ms**
|
||||
on two runs (578 pages).
|
||||
|
||||
**An index-heavy synthetic database of 105.2 MB** (105,160,704 bytes; 25,674
|
||||
pages of 4,096 bytes), shaped to give the cross-check real work: `actions` with
|
||||
152,000 rows and `memories` with 15,200, under three indexes
|
||||
(`i_actions_adv_depth`, `i_actions_kind`, `i_memories_adv`):
|
||||
|
||||
| | Run 1 | Run 2 | Run 3 |
|
||||
| --- | --- | --- | --- |
|
||||
| `quick_check` | 172.6 ms | 92.4 ms | 97.1 ms |
|
||||
| `integrity_check` | 227.8 ms | 196.0 ms | 227.4 ms |
|
||||
|
||||
A whole backup of it through `backup.create()`: **0.58 s** for 25,674 pages,
|
||||
`integrity = ok`.
|
||||
|
||||
**A campaign-shaped database of 117.4 MB** (117,403,648 bytes; 2,200 actions),
|
||||
built through the application's own models so the product path can run against
|
||||
it (§E.1):
|
||||
|
||||
| | Run 1 | Run 2 | Run 3 |
|
||||
| --- | --- | --- | --- |
|
||||
| `quick_check` | 89.5 ms | 74.2 ms | 64.4 ms |
|
||||
| `integrity_check` | 90.0 ms | 82.9 ms | 75.9 ms |
|
||||
|
||||
**Reading.** The cost of the cross-check tracks **rows and index entries, not
|
||||
bytes**. On the campaign schema — 2,200 fat rows — the full check costs about
|
||||
7 ms more than the quick one at 117 MB. On the index-heavy synthetic — 167,200
|
||||
rows under three indexes at a similar size — it costs about 96 ms more. Both sit
|
||||
inside a backup of about half a second. There is no performance requirement in
|
||||
this project and WP-D does not invent one; the measurements are here because the
|
||||
M9 trade was made on a speed argument, and at these sizes that argument does not
|
||||
hold.
|
||||
|
||||
### E.1 Through the endpoint the button calls — implementation evidence only
|
||||
|
||||
**This does not satisfy acceptance criterion 3, and an earlier draft of this
|
||||
report wrongly said it did.** The criterion asks for the backup to *complete
|
||||
through the UI*; a reader does not call an endpoint. What follows is evidence
|
||||
about the implementation — the procedure, its result and its cost — and it is
|
||||
kept for that reason. The acceptance evidence is in **§E.2**.
|
||||
|
||||
`POST /api/backups` is the only call the **Back up now** button makes, and it
|
||||
takes no parameters, so it was driven directly against a **copy** of each
|
||||
database — the evidence database is evidence and is not written to:
|
||||
|
||||
| Database | Result |
|
||||
| --- | --- |
|
||||
| Evidence campaign, 2,367,488 bytes | **201** in 0.117 s wall — `pages: 578`, `seconds: 0.038`, `integrity: ok`; `GET /api/backups` then lists 1 backup |
|
||||
| Campaign-shaped, 117,403,648 bytes | **201** in 0.536 s wall — `pages: 28,663`, `seconds: 0.452`, `integrity: ok` |
|
||||
|
||||
**Why a second large database exists.** The index-heavy synthetic has no
|
||||
application schema, and the server runs its migrations at startup, so the app
|
||||
will not start against it: driving the endpoint there fails in a migration that
|
||||
renumbers `actions`, before any backup is attempted. It remains valid evidence
|
||||
for the pragma comparison, which is a property of SQLite and not of this schema,
|
||||
but criterion 3 needs a database the product can actually open — hence the
|
||||
campaign-shaped one, which §E.2 then drives through the browser.
|
||||
|
||||
Artefacts stay under `$HOME` (`v11-evidence/wp-d/`), not in the repository.
|
||||
|
||||
---
|
||||
|
||||
### E.2 Through the UI — the acceptance evidence for criterion 3
|
||||
|
||||
The reader-facing **Back up now** control, clicked in a real Firefox, on the
|
||||
production build served by FastAPI: the WP-C path, with no Vite dev server, no
|
||||
inference, and no sleep used as an assertion. Every wait is on a condition the
|
||||
page or the filesystem can show.
|
||||
|
||||
`backend/tools/wpd_backup_ui.py`, run as
|
||||
`python -m tools.wpd_backup_ui --out $HOME/v11-evidence/wp-d/ui --case both`.
|
||||
**34 checks, 34 passed, 0 failed** —
|
||||
`$HOME/v11-evidence/wp-d/ui/backup-ui-report.json`.
|
||||
|
||||
The route a reader takes is the route the tool takes: open the application, click
|
||||
**Settings** in the navigation, open *Back up everything on this machine* (a
|
||||
`<details>` that loads what is on disk when it opens), then press the button.
|
||||
|
||||
| | **Case 1 — real campaign database** | **Case 2 — ≥100 MB application database** |
|
||||
| --- | --- | --- |
|
||||
| Source | the M11 evidence campaign, copied | the campaign-shaped database of §E, copied |
|
||||
| **Database size** | **2,367,488 bytes** (2.3 MB) | **117,403,648 bytes** (112.0 MB) |
|
||||
| Schema | the application's own, `user_version` 94 | the application's own, `user_version` 94 |
|
||||
| **Browser action** | click **Back up now** (twice — see below) | click **Back up now** |
|
||||
| **Observable UI result** | toast: *"Backup written: adventure-storyteller-20260915-225000.db (2.3 MB)."*, no error toast, the control returns from *Backing up…*, and the file appears in the panel's list | toast: *"Backup written: adventure-storyteller-20260915-225007.db (112.0 MB)."*, no error toast, control returns, file listed |
|
||||
| **Backup path** | `…/ui/real/data/backups/adventure-storyteller-20260915-225000.db` | `…/ui/large/data/backups/adventure-storyteller-20260915-225007.db` |
|
||||
| **Backup file size** | **2,367,488 bytes** | **117,403,648 bytes** |
|
||||
| **`integrity_check` result** | **`ok`** | **`ok`** |
|
||||
| **Integrity-check elapsed** | **5.8 ms** | **78.5 ms** |
|
||||
| **Total elapsed backup** | **0.108 s** (click → the UI says it is done) | **0.719 s** |
|
||||
|
||||
"Total elapsed" is measured from the click to the rendered result, so it is what
|
||||
the reader waits, not what the server reports. The `integrity_check` above is an
|
||||
independent second opinion, run here on the artefact the UI produced — the
|
||||
application had already verified the copy before keeping it.
|
||||
|
||||
**The size the UI reports is the size on disk.** "112.0 MB" in the toast is
|
||||
117,403,648 bytes shown in the page's own units; the file matches the source
|
||||
database byte for byte.
|
||||
|
||||
**Existing-backup protection, established through the UI itself** (Case 1's
|
||||
second press, rather than a backup planted by a library call):
|
||||
|
||||
| Claim | Result |
|
||||
| --- | --- |
|
||||
| A second press writes a *different* file — `…-225000-2.db`, the counter `_unused_name` adds when two backups land in the same second | PASS |
|
||||
| The first backup still exists afterwards | PASS |
|
||||
| And is byte-for-byte what it was — 2,367,488 bytes, `sha256:7d45555566…` | PASS |
|
||||
| And still passes `integrity_check` | PASS |
|
||||
|
||||
No corruption was fabricated through the browser; rejection behaviour is already
|
||||
proved deterministically in §D.
|
||||
|
||||
**Supplemental endpoint evidence.** The servers' own logs corroborate that the
|
||||
button drove each backup: Case 1 logged two `POST /api/backups → 201 Created`,
|
||||
each followed by `GET /api/backups → 200 OK` as the panel reloaded its list;
|
||||
Case 2 logged one of each. Screenshots and logs are beside the report JSON.
|
||||
|
||||
**A harness defect this run found, in my own check and not in the product.** The
|
||||
first pass reported six failures. The toast renders a decorative mark before its
|
||||
message — `<span class="toast-mark" aria-hidden="true">❖</span><span>…</span>` —
|
||||
so the button's `textContent` begins with `❖`, and my assertion matched
|
||||
`textContent.startswith("Backup written:")`. In that same run the UI had shown a
|
||||
non-error toast naming the file, the file was on disk, its name was in the
|
||||
panel's list, and the copy verified: the product was right and the assertion was
|
||||
reading the mark. The check now reads the message span. **No product code was
|
||||
changed** — the expected change for this work was none, and none was needed.
|
||||
|
||||
---
|
||||
|
||||
## F. Existing export/import limit behaviour
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Import ceiling | `limits.MAX_IMPORT_BODY_BYTES` = 20 MB, **unchanged by WP-D** |
|
||||
| Where it is enforced | `BodySizeLimitMiddleware`, on the declared `Content-Length`, before the body is read |
|
||||
| Refusal | HTTP **413**, "Request too large (limit 20 MB)." — it already named the limit |
|
||||
| Export before WP-D | `GET /adventures/{id}/export` returned the bundle. Nothing compared its size with the ceiling, so a campaign could be exported and then refused by its own importer |
|
||||
|
||||
---
|
||||
|
||||
## G. Oversized-export warning implementation
|
||||
|
||||
**Where the answer comes from.** `limits.import_limit_label()` and
|
||||
`limits.oversized_export_warning(size)` derive the sentence from
|
||||
`MAX_IMPORT_BODY_BYTES`. A test changes the constant and asserts the wording
|
||||
follows, so no number is written twice.
|
||||
|
||||
> This export is larger than this version's 20 MB import limit (21,230,000 bytes).
|
||||
> The file was exported successfully, but this version cannot import it.
|
||||
|
||||
**Where it travels: headers, not the body.** The export response *is* the bundle —
|
||||
the browser saves exactly those bytes as the file — so a warning inside it would
|
||||
become part of a portable story file and of every checksum taken over one. The
|
||||
route returns the same body with:
|
||||
|
||||
```text
|
||||
X-Export-Bytes the serialised size
|
||||
X-Import-Limit-Bytes MAX_IMPORT_BODY_BYTES
|
||||
X-Importable-By-This-Version true / false
|
||||
X-Export-Warning only when false
|
||||
```
|
||||
|
||||
**Which size is measured.** The compact serialisation (`separators=(",", ":")`,
|
||||
`ensure_ascii=False`), which is what this response sends *and* what the browser
|
||||
POSTs back on import — the bytes `BodySizeLimitMiddleware` weighs. The
|
||||
pretty-printed file the reader downloads is larger and is not what import reads.
|
||||
The route serialises once and returns those bytes, so `X-Export-Bytes` is the
|
||||
length of the body actually sent.
|
||||
|
||||
**The bundle is unchanged.** No key was added to it (§I), and the format stays
|
||||
`ai-dnd-adventure-v3`.
|
||||
|
||||
---
|
||||
|
||||
## H. Oversized export evidence
|
||||
|
||||
`test_an_oversized_export_is_still_delivered_and_says_it_cannot_come_back`
|
||||
writes real rows into a campaign until its bundle genuinely exceeds the ceiling —
|
||||
nothing is mocked, and the export serialises all of it.
|
||||
|
||||
| Claim | Result |
|
||||
| --- | --- |
|
||||
| The response body is larger than 20 MB | PASS |
|
||||
| It parses, and is `ai-dnd-adventure-v3` with its actions | PASS |
|
||||
| `X-Importable-By-This-Version: false` | PASS |
|
||||
| `X-Export-Warning` names the limit ("20 MB"), says the export succeeded, and says this version cannot import it | PASS |
|
||||
| `X-Export-Bytes` equals the body length | PASS |
|
||||
| The same bundle is refused by import, with the limit named (§J) | PASS |
|
||||
| A normal campaign's export carries **no** warning, and `X-Importable-By-This-Version: true` | PASS |
|
||||
| The warning follows the constant (changed to 50 MB in a test: the sentence says 50 MB, and no longer says 20 MB) | PASS |
|
||||
|
||||
---
|
||||
|
||||
## I. Normal-export compatibility
|
||||
|
||||
The same campaign database (`m04-final/campaign.db`) exported through a worktree
|
||||
at `59b5ebc` (before WP-D) and through the WP-D tree:
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Before | 2,699,076 bytes |
|
||||
| After | 2,699,076 bytes |
|
||||
| Comparison | **identical** — the parsed documents compare equal, key for key |
|
||||
|
||||
No timestamp allowance was needed: this campaign's bundle carries no field that
|
||||
moves between exports. The format is `ai-dnd-adventure-v3` in both.
|
||||
|
||||
`test_the_bundle_itself_never_carries_the_warning` additionally asserts that no
|
||||
key of an oversized bundle mentions the warning, the limit or importability.
|
||||
|
||||
---
|
||||
|
||||
## J. Import-refusal behaviour
|
||||
|
||||
Unchanged in behaviour, and already naming the limit:
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Status | **413**, from the middleware, on `Content-Length`, before the body is parsed |
|
||||
| Message | "Request too large (limit 20 MB)." |
|
||||
| WP-D test | `test_that_same_bundle_is_refused_by_import_naming_the_limit` posts the *actual oversized export* and asserts 413 and that the detail contains `limits.import_limit_label()` |
|
||||
| Not relaxed | the same boundary, the same status, the same middleware. `test_a_bundle_under_the_limit_still_imports` keeps the ordinary path honest (201) |
|
||||
|
||||
No wording correction was needed.
|
||||
|
||||
---
|
||||
|
||||
## K. Frontend behaviour
|
||||
|
||||
`api.exportAdventure` now returns the bundle **and** what the server said about
|
||||
it (`exportBytes`, `importLimitBytes`, `importable`, `warning`), read from the
|
||||
headers.
|
||||
|
||||
**The no-header case.** `importable` is `resp.headers.get(…) !== 'false'`, so
|
||||
only the literal string `false` is read as a refusal: an older server that sends
|
||||
no headers — or a header that arrives malformed — yields `importable: true` and
|
||||
`warning: null`, and the page says nothing it was not told. The two size fields
|
||||
fall back to `null` unless they parse as a positive number.
|
||||
|
||||
This is asserted directly: three tests stub `fetch` with real response headers
|
||||
and check what `api.exportAdventure` makes of them — a warning with the sizes it
|
||||
named, an ordinary export marked importable, and an older server sending no
|
||||
headers at all. They exist because the page-level tests below mock
|
||||
`api.exportAdventure` itself and so cannot see a header name, which meant a typo
|
||||
on the frontend side would have left the whole suite green (§O.7).
|
||||
|
||||
Both reader-facing entry points keep delivering the file first and then report:
|
||||
|
||||
| Entry point | Under the limit | Over the limit |
|
||||
| --- | --- | --- |
|
||||
| Campaign library, **Export** | file written; "Campaign exported." | file written; the warning, as an error-styled toast |
|
||||
| Campaign settings, **Export campaign** | file written; "Campaign exported." | file written; the warning |
|
||||
|
||||
No new modal, no redesign: the existing toast carries it.
|
||||
|
||||
**Tests** (`pages/exportHonesty.test.jsx`, **7**): three read real response
|
||||
headers through a stubbed `fetch` (above), and four drive both entry points —
|
||||
the file is delivered in every case; the warning is shown when the server sends
|
||||
one, naming 20 MB and saying the export succeeded; the ordinary confirmation is
|
||||
shown when it does not.
|
||||
|
||||
---
|
||||
|
||||
## L. Offline regression
|
||||
|
||||
`tools/m11_offline.py`, against the production build, in the offline container:
|
||||
|
||||
**23 checks, 23 passed, 0 failed** —
|
||||
`$HOME/v11-evidence/wp-d/offline/offline-report.json`.
|
||||
|
||||
The two this package could have broken are in it and passed:
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| a campaign exports offline | ok |
|
||||
| and imports offline, with its state | ok |
|
||||
| no secret is present in the export | ok |
|
||||
| a local file imports offline | ok |
|
||||
|
||||
The export route now sets headers and serialises compactly; the offline
|
||||
container still exports a campaign and imports it back with its state, so
|
||||
neither the round trip nor the secret-scrubbing changed.
|
||||
|
||||
## M. Full regression
|
||||
|
||||
| Suite | Result |
|
||||
| --- | --- |
|
||||
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,712 passed, 17 skipped, 0 failed, 0 xfailed** (936.6 s) |
|
||||
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
|
||||
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
|
||||
| **Production build** (`npm run build`) | succeeded |
|
||||
| **Offline regression** | 23 passed, 0 failed (§L) |
|
||||
|
||||
**The backend count reconciles exactly.** WP-B's closing tree was 1,693; WP-C
|
||||
added 7 (`test_v11_c_browser_helpers.py`) and did not rerun the suite because no
|
||||
application code changed; WP-D adds the 12 in `test_v11_d_recovery.py`.
|
||||
1,693 + 7 + 12 = **1,712**.
|
||||
|
||||
**The 17 skips are named, not assumed.** Run with `-rs`, every one is an
|
||||
environment-gated real-model test, and none is new:
|
||||
|
||||
| File | Skipped | Gate |
|
||||
| --- | --- | --- |
|
||||
| `test_knowledge_real_model.py` | 7 | `AIDND_TEST_ENDPOINT` (and `AIDND_TEST_EMBED_MODEL`) |
|
||||
| `test_context_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
|
||||
| `test_narrative_realistic.py` | 3 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
|
||||
| `test_m11_real_window.py` | 3 | two the same; one `AIDND_TEST_WIDE_MODEL` |
|
||||
| `test_provider_wiring.py` | 1 | `AIDND_TEST_ENDPOINT` and `AIDND_TEST_MODEL` |
|
||||
|
||||
That is the same 17 B.1 and B.2 recorded. WP-D used no model and added no skip.
|
||||
|
||||
**The plan's named regression requirements**, run individually rather than
|
||||
assumed to be inside the total:
|
||||
|
||||
| Requirement | Result |
|
||||
| --- | --- |
|
||||
| I01-I07, L01-L04 | **25 passed**, 1,704 deselected |
|
||||
| The backup case in `test_m11_migration.py` | **1 passed** |
|
||||
| `backup.test.jsx` | **7 passed** |
|
||||
| The offline container's export and import | passed (§L) |
|
||||
|
||||
**Frontend arithmetic.** WP-C's baseline was 168. WP-D adds 4 page-level export
|
||||
tests and 3 header-parsing tests: 168 + 7 = **175**. The 15 lint warnings are
|
||||
the same pre-existing `only-export-components` and unused-import kind recorded
|
||||
at WP-C, and **none is in a file WP-D changed**.
|
||||
|
||||
## N. Compatibility
|
||||
|
||||
| Area | Effect |
|
||||
| --- | --- |
|
||||
| Database schema | unchanged; no migration |
|
||||
| Bundle format | `ai-dnd-adventure-v3`, unchanged |
|
||||
| Normal bundle contents | byte-identical (§I) |
|
||||
| Existing backups | still valid files with the same names; only the check that admits a new one is stricter |
|
||||
| v1.0.0 databases | open unchanged (no schema or data path changed) |
|
||||
| Import limit | unchanged at 20 MB |
|
||||
| Scheduled backups | none, as before |
|
||||
| History, branches, Save Points, state, memory, knowledge | untouched |
|
||||
| Endpoint policy, local-only operation | untouched |
|
||||
|
||||
## O. Residual risks
|
||||
|
||||
1. **Closed: the backup is now driven through the UI.** This risk previously
|
||||
read "the browser click was not driven for the backup", and the owner
|
||||
correctly refused criterion 3 on endpoint evidence. §E.2 drives the real
|
||||
**Back up now** control in Firefox on both databases — 34 checks, 0 failed —
|
||||
including the ≥100 MB case, which does mean the harness starts a server
|
||||
against a 117 MB database and takes 0.719 s to do it. What remains is
|
||||
ordinary coverage scope: this runs as its own tool rather than inside the
|
||||
release harness, so it is not part of the 101-check run in the WP-E report.
|
||||
2. **The 20 MB ceiling is unchanged.** A campaign past it still cannot be
|
||||
imported by this version. WP-D makes that audible at export; raising the
|
||||
limit or streaming import stays v1.2, as the plan assigns it.
|
||||
3. **The warning is about *this* version.** It says what this build's importer
|
||||
will accept. A future build with a higher ceiling could import a file this
|
||||
one warned about, and the wording ("this version") is chosen so that stays
|
||||
true rather than becoming a lie.
|
||||
4. **`integrity_check` is not a guarantee of recoverability.** It proves the
|
||||
copy's pages, records and indexes agree. A database that was already
|
||||
logically wrong when it was copied is copied faithfully and passes. The
|
||||
package makes the check honest, not omniscient.
|
||||
5. **No restore path.** There is still no restore button and no scheduled
|
||||
backup — both explicitly out of scope. A reader who needs a backup restores
|
||||
it by replacing the file themselves, as before.
|
||||
6. **The header channel depends on the browser reaching headers.** Both export
|
||||
call sites read them through `fetch`, so a proxy that stripped `X-` headers
|
||||
would silently return to v1 behaviour: the file still arrives, the warning
|
||||
does not. Local-only operation makes that unlikely, and the failure is the
|
||||
old behaviour rather than a wrong claim.
|
||||
7. **Header parsing was untested — now closed.** Writing this section exposed
|
||||
it: the page-level tests mock `api.exportAdventure` and the backend tests
|
||||
assert what the server sends, so nothing read an actual header name, and a
|
||||
frontend-side typo would have left every test green and the reader silently
|
||||
uninformed. Three tests that stub `fetch` with real headers now cover it
|
||||
(§K). What remains is the ordinary version of this risk: the two sides agree
|
||||
by matching string literals in two files, and only a browser-level test would
|
||||
catch a mismatch introduced in both at once.
|
||||
|
||||
## P. Final decision
|
||||
|
||||
**Against the plan's acceptance criteria** (§ WP-D, *Recovery honesty*):
|
||||
|
||||
| # | Criterion | Verdict |
|
||||
| --- | --- | --- |
|
||||
| 1 | A healthy database's backup reports `integrity_check` ok and is kept | **PASS** — `test_a_healthy_backup_passes_the_full_check_and_is_kept`, and both endpoint runs returned `integrity: ok` (§E.1) |
|
||||
| 2 | A copy `integrity_check` rejects and `quick_check` does not is rejected and not kept; existing backups still never overwritten | **PASS** — the fixture is proved to be exactly that difference (§D), and three tests cover rejection, the untouched earlier backup, and the unchanged file semantics |
|
||||
| 3 | The time is recorded on the evidence database and on a synthetic ≥100 MB, and the backup completes through the UI on both | **PASS** — pragma timings in §E on three databases, and the backup completes **through the UI** on both in §E.2: 0.108 s with `integrity_check` `ok` in 5.8 ms on the 2.3 MB campaign, 0.719 s with `ok` in 78.5 ms on the 117 MB application database. An earlier draft claimed this criterion on endpoint evidence (§E.1); that claim was wrong and is corrected |
|
||||
| 4 | An over-20 MB export succeeds, delivers the file, and warns naming the limit, in the API response and a component test; under the limit, no warning | **PASS** — §H, and 7 frontend tests (§K) |
|
||||
| 5 | Importing that bundle is refused with a message naming the limit | **PASS** — §J, asserted on the actual oversized export |
|
||||
| 6 | A normal export is byte-identical before and after, apart from timestamps | **PASS** — fully identical, no timestamp allowance needed (§I) |
|
||||
|
||||
**What I corrected rather than reported around.** Two claims of my own failed
|
||||
checking and were fixed, not softened: §E originally rested criterion 3 on
|
||||
`backup.create()` calls, and driving the real endpoint showed the index-heavy
|
||||
synthetic cannot serve that criterion at all (the app's startup migrations abort
|
||||
against a schema-less database) — hence the campaign-shaped database in §E.1.
|
||||
And §K asserted the no-header fallback from the code alone, which exposed that
|
||||
**no test read any header name**; three tests now do (§O.7).
|
||||
|
||||
**Not done, and not claimed:** the import ceiling is unchanged at 20 MB, by the
|
||||
plan's own assignment of the raise to v1.2. The UI backup runs as its own tool
|
||||
rather than inside the release harness (§O.1).
|
||||
|
||||
**No product code changed for the UI verification.** The run exposed one defect,
|
||||
and it was in my own assertion, not in the application (§E.2). Backend
|
||||
**1,723 passed / 17 skipped / 0 failed** and offline **23/23** therefore stand
|
||||
as prior evidence, unchanged and not re-run; the targeted backup and browser
|
||||
checks were re-run instead.
|
||||
|
||||
```text
|
||||
BACKUP INTEGRITY: PASS
|
||||
BACKUP THROUGH UI — REAL DB: PASS
|
||||
BACKUP THROUGH UI — >=100 MB DB: PASS
|
||||
EXPORT HONESTY: PASS
|
||||
|
||||
WP-D OVERALL:
|
||||
PASS
|
||||
```
|
||||
|
||||
All WP-D changes are **staged and uncommitted**. No commit, no push, no tag.
|
||||
WP-E has not begun.
|
||||
@@ -0,0 +1,380 @@
|
||||
# v1.1 WP-E — Control-Boundary Contrast
|
||||
|
||||
**Status:** COMPLETE — **PASS**, with owner approval of the screenshots
|
||||
outstanding. The decision, and what was deliberately not claimed, is in §O.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Branch | `v1.1-development` |
|
||||
| HEAD | `59b5ebc` — *v1.1 WP-C: browser release coverage*, signed by the owner |
|
||||
| Working tree at start | WP-D staged (10 files), nothing committed |
|
||||
| Criterion | WCAG 2.1 **1.4.11 Non-text Contrast**, 3:1, for control boundaries; **1.4.3** 4.5:1 for body text, unchanged |
|
||||
|
||||
---
|
||||
|
||||
## B. What v1.0.0 actually did
|
||||
|
||||
`tools/contrast_audit.py` measured control boundaries, printed that two of them
|
||||
were below 3:1, and **exited 0**. Its own comment argued the position:
|
||||
|
||||
> in this design a control is identified by its *label*, which is measured above
|
||||
> and passes, not by its edge. So a boundary below 3:1 is reported with its
|
||||
> number and does not fail the run.
|
||||
|
||||
So the audit was a report, not a gate: no palette change could ever fail it on a
|
||||
boundary. The two numbers it printed were **1.33:1** (`--border` on
|
||||
`--bg-panel`) and **1.75:1** (`--border-bright`), against a floor of 3.0.
|
||||
|
||||
**WP-E overturns that argument.** 1.4.11 covers the visual information needed to
|
||||
identify a component *and its boundary*; a reader who cannot see where a text box
|
||||
ends cannot see that there is a text box to type into, label or no label. The
|
||||
tokens were raised rather than the criterion re-argued.
|
||||
|
||||
---
|
||||
|
||||
## C. Token inventory
|
||||
|
||||
| Token | v1.0.0 | v1.1 | Why |
|
||||
| --- | --- | --- | --- |
|
||||
| `--border` | `#2b2b3d` | **`#676792`** | every control's resting edge |
|
||||
| `--border-bright` | `#3d3d55` | **`#7a7aaa`** | hover edges, the composer's resting edge, the open panel tab |
|
||||
| `--bg-panel`, `--bg-input`, `--bg`, `--text`, `--text-dim`, `--accent*`, `--danger`, `--warning`, `--player`, `--chart-*` | — | **unchanged** | WP-E is a boundary package; no text or accent colour moved |
|
||||
|
||||
**The floor is taken against `--bg-input`, not `--bg-panel`.** Inputs and buttons
|
||||
are drawn on `--bg-input` (`styles/forms.css`), which is lighter than
|
||||
`--bg-panel` and therefore the harder case. The audit had been checking only
|
||||
`--bg-panel`, so a token could have passed the audit while the real control
|
||||
failed. Measured on the new values:
|
||||
|
||||
| | vs `--bg-input` | vs `--bg-panel` | vs `--bg` |
|
||||
| --- | --- | --- | --- |
|
||||
| `--border` | **3.21:1** | 3.44:1 | 3.70:1 |
|
||||
| `--border-bright` | **4.24:1** | 4.55:1 | 4.88:1 |
|
||||
|
||||
Two properties were preserved deliberately: the rest→hover step is the same size
|
||||
as before (1.318 → 1.320), so hover still reads as a change rather than a jump;
|
||||
and `--border-bright` stays *below* body text against the same panel (2.98:1
|
||||
between them), so no edge outshines the words inside it.
|
||||
|
||||
---
|
||||
|
||||
## D. Component inventory
|
||||
|
||||
Where these tokens are actually drawn, from the stylesheets:
|
||||
|
||||
| Control | Rule | Rest | Hover / active | Focus |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Story composer | `.input-bar` (story.css) | `--border-bright` on `--bg-panel` | — | `--accent-dim` + `--accent-glow` ring |
|
||||
| Story controls | `.story-controls button` | `--border` on `--bg-panel` | `--border-bright` | (M11 focus check) |
|
||||
| Fields and buttons | `forms.css` | `--border` on `--bg-input` | `--accent-dim` | `--accent-dim` + ring |
|
||||
| Panel tabs | `.panel-tabs button` | **`transparent`** | `--border` on hover, `--border-bright` when active | — |
|
||||
| Top navigation | `.topnav` | `--border` bottom edge on **`--bg-panel-glass`** | — | — |
|
||||
| Scrollbar thumb | `base.css` | `--border-bright` as a *fill* on `--bg` | `--accent-dim` | — |
|
||||
|
||||
Two of these cannot be answered by token arithmetic at all, and both are
|
||||
measured in the browser instead (§G): the nav sits on a translucent panel, and
|
||||
the panel tab's edge is `transparent` until the panel is open.
|
||||
|
||||
---
|
||||
|
||||
## E. The audit is now a gate
|
||||
|
||||
`tools/contrast_audit.py`:
|
||||
|
||||
1. **Boundary pairs fail.** `text` and `boundary` rows are both pass/fail; the
|
||||
advisory branch is gone. The run returns 1 if either kind falls short.
|
||||
2. **Eight boundary pairs replace two.** Each border is checked against every
|
||||
background it is drawn on — `--bg-input`, `--bg-panel` and `--bg` — plus the
|
||||
focused edge (`--accent-dim`) on both panel and field backgrounds.
|
||||
3. **The verdict is taken on the number that is printed** (rounded to two
|
||||
decimals), so a pair shown as `3.00:1` is not failed for arithmetic the
|
||||
reader cannot see.
|
||||
4. The comment block that argued the old position is replaced by one recording
|
||||
what changed and why, including why `--bg-panel-glass` is not in the list.
|
||||
|
||||
## F. Gate tests
|
||||
|
||||
`backend/tests/test_v11_e_contrast.py` — **11 passed**. The threshold is
|
||||
exercised from both sides, on real token files:
|
||||
|
||||
| Test | Result |
|
||||
| --- | --- |
|
||||
| A boundary at **2.99:1** against `--bg-input` fails the run (exit 1) | PASS |
|
||||
| A boundary at **3.00:1** passes (exit 0) | PASS |
|
||||
| The **v1.0.0 value** `#2b2b3d` fails, at 1.24:1 against `--bg-input` | PASS |
|
||||
| A dimmed `--text-dim` still fails as a *text* pair | PASS |
|
||||
| A renamed token is a failure, not a silent skip | PASS |
|
||||
| The shipped palette passes both criteria | PASS |
|
||||
| Every boundary pair is measured against the background it is drawn on | PASS |
|
||||
| The M11 text baselines are unchanged: 14.57 / 13.57 / 5.48 / 5.88 | PASS |
|
||||
| The hover edge stays brighter than the resting edge | PASS |
|
||||
| No boundary becomes as loud as body text | PASS |
|
||||
| The WCAG ratio formula is anchored on known values (21:1, 1:1, symmetry) | PASS |
|
||||
|
||||
The 2.99 and 3.00 values are worth noting: **both clear 3:1 against
|
||||
`--bg-panel`** (3.21 and 3.22). They decide the gate only because the floor is
|
||||
now taken against the background the control is really on — so these two tests
|
||||
also prove §C's change is doing work.
|
||||
|
||||
---
|
||||
|
||||
## G. Browser measurement
|
||||
|
||||
`tools/m11_browser.py` gains a third suite, **WP-E**, counted separately from
|
||||
M11's 38 and WP-C's 53. It measures the *rendered* edge — `borderColor` from
|
||||
`getComputedStyle` — against what is actually behind it, with every translucent
|
||||
layer composited bottom-up.
|
||||
|
||||
A boundary is measured against **both** adjacent colours (the control's own fill
|
||||
inside it, the background outside it) and passes on the better of the two: an
|
||||
edge that matches its fill but contrasts with the page is still a visible
|
||||
outline. What 1.4.11 asks is that the component's extent be perceivable.
|
||||
|
||||
Two harness capabilities were added for this (`tools/m11_webdriver.py`):
|
||||
|
||||
- **`hover()`** moves a real pointer through the WebDriver Actions API.
|
||||
Dispatching a `mouseover` event from JavaScript does *not* trigger CSS
|
||||
`:hover`, so a synthetic event would have re-measured the resting edge and
|
||||
reported it as the hover edge.
|
||||
- **`screenshot()`** writes the viewport as a PNG, for the before/after evidence.
|
||||
|
||||
**A defect this found in my own first measurement.** The first run reported the
|
||||
hover edge as `rgb(114, 114, 160)` and the focused edge as `rgb(144, 120, 81)` —
|
||||
neither of which is any token. Both controls carry `transition: border-color
|
||||
0.15s`, so the measurement was taken mid-animation, on a colour no state
|
||||
actually has. `_settled()` now polls until the computed edge colour is the same
|
||||
on two consecutive reads before measuring (polled, not slept, per this harness's
|
||||
own rule). After the fix the same edges read exactly `rgb(122, 122, 170)`
|
||||
(`--border-bright`) and `rgb(150, 119, 58)` (`--accent-dim`).
|
||||
|
||||
---
|
||||
|
||||
## H. Before and after, measured in the browser
|
||||
|
||||
Both passes were taken the same way — `--only boundaries --no-narrator`, the
|
||||
production build — with only `tokens.css` differing. Evidence under
|
||||
`$HOME/v11-evidence/wp-e/before/` and `.../after/`.
|
||||
|
||||
| Control (state) | Before | After | Floor |
|
||||
| --- | --- | --- | --- |
|
||||
| Story composer — resting edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
|
||||
| Story control — resting edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
|
||||
| Open panel tab — resting edge | **1.75:1** FAIL | **4.55:1** pass | 3.0 |
|
||||
| Story control — hover edge | **1.88:1** FAIL | **4.88:1** pass | 3.0 |
|
||||
| Top navigation — translucent edge | **1.43:1** FAIL | **3.70:1** pass | 3.0 |
|
||||
| Story composer — focused edge | 4.70:1 pass | 4.70:1 pass | 3.0 |
|
||||
|
||||
**Suite result: before 5 passed / 5 failed; after 10 passed / 0 failed / 0
|
||||
skipped.**
|
||||
|
||||
Two things this table says that a summary would blur:
|
||||
|
||||
- **Focus was never the defect.** The focused edge (`--accent-dim`) already
|
||||
cleared 3:1 in v1.0.0 at 4.70:1, and WP-E did not change it. What failed was
|
||||
rest and hover — the states a reader spends all their time in.
|
||||
- **The translucent edge is real evidence.** The nav's background composited to
|
||||
`rgb(17, 17, 29)` — `--bg-panel-glass` (rgba 19,19,32 @ 0.82) over
|
||||
`rgb(10, 10, 15)` — not a fallback. That is the case token arithmetic cannot
|
||||
reach, and it moved from 1.43:1 to 3.70:1.
|
||||
|
||||
## I. Screenshots
|
||||
|
||||
| File | |
|
||||
| --- | --- |
|
||||
| `before/control-boundaries.png` | 131,175 bytes, 1366×682 |
|
||||
| `before/control-boundaries-nav.png` | 63,376 bytes, 1366×682 |
|
||||
| `after/control-boundaries.png` | 132,430 bytes, 1366×682 |
|
||||
| `after/control-boundaries-nav.png` | 63,541 bytes, 1366×682 |
|
||||
|
||||
All four are PNG, 1366×682, taken on the production build through the same
|
||||
harness path, differing only in `tokens.css`. The play-page pair shows the
|
||||
composer, the story controls and the open panel tab; the nav pair shows the
|
||||
translucent top edge on the library route.
|
||||
|
||||
## J. A finding this package created and fixed
|
||||
|
||||
Raising `--border-bright` broke something that had nothing to do with control
|
||||
boundaries. `.slice-7` in the context inspector's token breakdown was painted
|
||||
with `var(--border-bright)`, so it followed the token to `#7a7aaa` — an OKLab ΔE
|
||||
of **0.035** from `.slice-6` (`#7c86b8`), making two neighbouring chart slices
|
||||
effectively the same colour. The other slices sit **0.100–0.119** from their
|
||||
nearest neighbour.
|
||||
|
||||
`.slice-7` is now pinned to `#3d3d55`, the literal value it already rendered, so
|
||||
its appearance is unchanged from v1.0.0 and its separation (ΔE **0.251**) is the
|
||||
widest in the set. A chart fill and a control edge have different jobs and should
|
||||
not share a token.
|
||||
|
||||
Reassigning it to a fresh hue was considered and rejected on evidence: inside the
|
||||
palette's own chroma (0.045–0.120) and lightness (0.586–0.804) bands, the only
|
||||
hues clearing the set's 0.100 separation floor are pinks near 14°, which is
|
||||
`--danger`'s territory. Painting an ordinary prompt section in the colour this
|
||||
application reserves for failure would trade an accessibility fix for a semantic
|
||||
lie.
|
||||
|
||||
*(A first attempt, `#5d7f9e`, was rejected by the same measurement at ΔE 0.061 —
|
||||
below every real slice. It is recorded here because it was written into the file
|
||||
before it was measured.)*
|
||||
|
||||
## K. Text contrast regression
|
||||
|
||||
Unchanged, and asserted so in §F: **14.57:1** body text on the page, **13.57:1**
|
||||
in a panel, **5.48:1** secondary text in a panel, **5.88:1** on the page. Every
|
||||
text pair still clears 1.4.3, and no text token was touched.
|
||||
|
||||
## L. Browser release regression
|
||||
|
||||
The full harness, all three suites, against a real narrator on the production
|
||||
build. Evidence: `$HOME/v11-evidence/wp-e/release/browser-report.json`.
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Kind | **`release regression`** — not `partial`, not `development (only …)` |
|
||||
| Narrator | `qwen2.5:3b-instruct`, the reference model, over plain HTTP on the LAN GPU host |
|
||||
| Browser | Firefox 155.0.1, geckodriver 0.37.1 |
|
||||
| Served | FastAPI on loopback, the **built** SPA (`dist` 2026-09-15T21:08:52) |
|
||||
| Duration | 71 s |
|
||||
| Turns played | **8**, of which `turns_not_clean` **0** and `protocol_shapes_in_narration` **0** |
|
||||
|
||||
**Counted per suite, as the brief requires:**
|
||||
|
||||
| Suite | Passed | Failed | Skipped |
|
||||
| --- | --- | --- | --- |
|
||||
| **M11** (the v1 release regression) | **38** | **0** | **0** |
|
||||
| **WP-C** (browser release coverage) | **53** | **0** | **0** |
|
||||
| **WP-E** (control boundaries) | **10** | **0** | **0** |
|
||||
| **Total** | **101** | **0** | **0** |
|
||||
|
||||
M11's 38 and WP-C's 53 are unchanged in count and in name: WP-E added a suite
|
||||
beside them rather than altering either. The `kind` field is quoted above
|
||||
because a fast run invites the question — 71 s for 101 checks including 8
|
||||
narrated turns is the GPU host being quick with a 3B model, and the 8 recorded
|
||||
turns with no unclean accounting are what rule out narration having been
|
||||
skipped.
|
||||
|
||||
The WP-E rows in this run are the same ten as the standalone capture in §H,
|
||||
re-measured with a narrator present and a full story on the page.
|
||||
|
||||
## M. Full regression
|
||||
|
||||
| Suite | Result |
|
||||
| --- | --- |
|
||||
| **Full backend suite** (`pytest -q`, no `AIDND_TEST_*` set) | **1,723 passed, 17 skipped, 0 failed, 0 xfailed** (1,054.7 s) |
|
||||
| **Frontend suite** (`npm test`) | **175 passed**, 15 files, 0 failed |
|
||||
| **Lint** (`npm run lint`, oxlint) | **exit 0**, 0 errors, 15 warnings |
|
||||
| **Production build** (`npm run build`) | succeeded |
|
||||
| **Browser harness** (M11 + WP-C + WP-E) | **101 passed, 0 failed, 0 skipped** (§L) |
|
||||
| **Contrast audit** (`python tools/contrast_audit.py`) | **exit 0** — every text pair and every boundary pair passes |
|
||||
|
||||
**The backend count reconciles exactly.** WP-D's tree was 1,712; WP-E adds the
|
||||
11 in `test_v11_e_contrast.py`. 1,712 + 11 = **1,723**. The 17 skips are the same
|
||||
environment-gated real-model tests recorded since B.1 — WP-E used no model and
|
||||
added no skip.
|
||||
|
||||
**The plan's regression requirements for WP-E** were the frontend suite and
|
||||
lint, and the harness's accessibility checks. All three pass: 175 and exit 0
|
||||
above, and M11's `A11y` rows — accessible names, visible keyboard focus, no
|
||||
positive tabindex, nothing revealed only on hover, and the four rendered text
|
||||
contrasts — are inside the 38/38 in §L.
|
||||
|
||||
**Lint detail.** The 15 warnings are the same pre-existing
|
||||
`only-export-components` and unused-import kind recorded at WP-C and WP-D, and
|
||||
**none is in a file WP-E changed**. `tokens.css` and `context.css` are not
|
||||
flagged.
|
||||
|
||||
## N. Residual risks
|
||||
|
||||
1. **The screenshots are unapproved.** The measurements say every boundary now
|
||||
clears 3:1; whether the result *looks* right in this design is a judgment
|
||||
the numbers cannot make. Recorded as PENDING in §O, not assumed.
|
||||
2. **Five controls are measured in the browser; the rest inherit.** The audit
|
||||
checks token pairs, and the harness measures the composer, a story control,
|
||||
the open panel tab, the nav edge and the focused composer. Every other
|
||||
bordered surface — modals, cards, the knowledge and context panels — draws
|
||||
the same two tokens, so it moves with them, but none is individually
|
||||
measured. A component that overrides a border with a literal colour would
|
||||
not be caught by either check.
|
||||
3. **`--bg-panel-glass` is outside the token audit by nature.** It is rgba over
|
||||
a gradient, so no token pair can express it; its edge is covered only by the
|
||||
browser measurement, which runs in the harness rather than in CI.
|
||||
4. **Disabled controls are deliberately not measured.** `.story-controls
|
||||
button:disabled` carries `opacity: 0.35`, so a disabled control's rendered
|
||||
edge is dimmer than any value here. WCAG 1.4.11 exempts inactive components,
|
||||
and the harness selects `:not(:disabled)` on purpose — stated so that the
|
||||
exclusion is visible rather than looking like an oversight.
|
||||
5. **The chart set was re-checked only where WP-E disturbed it.** `.slice-7`'s
|
||||
separation and colour-blind distance were measured against the other seven
|
||||
(§J); the set as a whole was not re-audited, which is outside this package.
|
||||
6. **`a11y.test.jsx` was not extended**, though the plan listed it as likely
|
||||
affected. It asserts structure and names, not colours, and adding colour
|
||||
assertions in jsdom would test the stylesheet's text rather than a rendered
|
||||
result. Boundary contrast is asserted instead where it can be measured: the
|
||||
gate tests (§F) and the browser (§G).
|
||||
|
||||
**One risk that turned out not to exist.** §G's rule — a boundary passes on the
|
||||
better of its two adjacent colours — was written to avoid failing an edge that
|
||||
contrasts with the page but matches its own fill. In the event it never did any
|
||||
work: every measured boundary clears 3:1 against **both** neighbours (composer
|
||||
4.55/4.88, story control 3.44/3.70, panel tab 4.24/4.55, hover 4.55/4.88, focus
|
||||
4.38/4.70, nav 3.51/3.70). The stricter reading would have produced the same
|
||||
verdict on every row.
|
||||
|
||||
## O. Final decision
|
||||
|
||||
**Against the plan's acceptance criteria** (§ WP-E, *Control-boundary contrast*):
|
||||
|
||||
| # | Criterion | Verdict |
|
||||
| --- | --- | --- |
|
||||
| 1 | `tools.contrast_audit` exits non-zero on any boundary pair below 3:1, and 0 on the package tree | **PASS** — 2.99:1 exits 1 and 3.00:1 exits 0 (§F); the v1.0.0 value exits 1; the package tree exits 0 |
|
||||
| 2 | Every text pair still clears 4.5:1, and the rendered text contrasts do not fall below v1's 14.57 / 5.48 / 13.57 / 5.88 | **PASS** — asserted as exact baselines in §F and re-measured on the rendered page in §L's M11 rows. No text token changed |
|
||||
| 3 | The rendered boundary of the story input and of a primary control, at rest and on hover, is at least 3:1; the focus indicator is still visible | **PASS** — composer 4.88:1, story control 3.70:1 at rest and 4.88:1 on hover (§H), and M11's visible-focus check passes in §L |
|
||||
| 4 | The owner approves the before-and-after screenshots, and the report records the approval | **PENDING** — not a criterion I can satisfy. The four PNGs are in §I |
|
||||
|
||||
**Where I exceeded the criterion, said plainly.** The plan asks for 3:1 "against
|
||||
their panel" and names two controls. This package measures against
|
||||
**`--bg-input`** as well — the lighter background inputs and buttons are really
|
||||
drawn on, and the one that decides the gate — and adds the open panel tab, the
|
||||
translucent navigation edge and the focused state. The stricter floor is the
|
||||
reason the two threshold tests in §F are decided by `--bg-input` rather than
|
||||
`--bg-panel`, where both would have passed.
|
||||
|
||||
**What I got wrong and corrected.** Three of my own claims failed checking and
|
||||
were fixed rather than softened: the first hover and focus measurements were
|
||||
taken mid-transition and reported colours no state has (§G); two contrast
|
||||
figures were written into `context.css` before being measured, and were wrong
|
||||
(§J); and my first gate check reported the current tokens as failing because my
|
||||
harness crashed on a shallow path, not because of any contrast.
|
||||
|
||||
**Not done, and not claimed:** owner approval (criterion 4), individual
|
||||
measurement of every bordered component (§N.2), and any palette work beyond the
|
||||
two boundary tokens and the one chart fill that borrowed from them.
|
||||
|
||||
```text
|
||||
WP-E CONTRAST GATE: PASS
|
||||
WP-E BOUNDARY MEASUREMENT: PASS
|
||||
|
||||
WP-E OVERALL:
|
||||
PASS, pending owner approval of the screenshots (criterion 4)
|
||||
```
|
||||
|
||||
All WP-E changes are **staged and uncommitted**. No commit, no push, no tag.
|
||||
v1.1 release validation has not begun.
|
||||
|
||||
```text
|
||||
OWNER SCREENSHOT APPROVAL: APPROVED
|
||||
```
|
||||
|
||||
The before/after screenshots in §I are the evidence for a change a reader judges
|
||||
by looking at it. The measurements say every boundary now clears 3:1; whether the
|
||||
result looks right in this design was the owner's call.
|
||||
|
||||
**Approved by the owner on 2026-09-16**, in the v1.1 release-validation brief,
|
||||
after reviewing the before/after pair in `$HOME/v11-evidence/wp-e/`. This line
|
||||
was `PENDING` in the signed commit `87a4032` because the report predated that
|
||||
review; it is updated here as part of the release closeout, with its source and
|
||||
date recorded rather than the approval being assumed. No visual code changed
|
||||
during release validation, so the approval stands (release report §Q).
|
||||
Reference in New Issue
Block a user