From 44edece67e7f65bacf78010ef56f98c3c8864073 Mon Sep 17 00:00:00 2001 From: JesseMarkowitz Date: Mon, 7 Sep 2026 01:55:45 -0400 Subject: [PATCH] M9: a campaign you can actually get back MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A campaign could already be exported and imported. What could not survive the trip was everything that explains it: the state events behind the authoritative document, the prompt each turn was actually given, the passages it was shown, the summaries that carry long-story continuity, and which take belonged to which turn. An imported campaign could be read and could no longer say why it was what it was — and a manual correction, the one state change no narration explains, was indistinguishable from something the story had established. The bundle is now `ai-dnd-adventure-v3`, and the version is the design rather than a side effect. Everything added here could have been another optional key, the way persona, Save Points, narrative state and imported knowledge each were. That mechanism stops working at exactly this addition: a v2 file with no prompt provenance is ambiguous between "written before M9" and "written by M9 from a campaign that has none", and those are different facts about a campaign. A version number is how a recovery file states what it was capable of recording. v1 and v2 still import, and every seam from pre-active-head onward is tested for the rule that an older file is never reinterpreted under a newer assumption. Two categories became three. "Chosen travels, derived is recomputed" was enough until stored prompts had to be decided: they are derived, and they must travel anyway. The test that separates evidence from cache is not "could this be recomputed" but "would a recomputation answer the same question" — a rebuilt search index answers the same question, a rebuilt prompt says what the turn would be told *now*, which is the opposite of what the inspector is for. Also here: a real SQLite backup, through the online backup API rather than a file copy, taken while the application is running and verified before it is kept; story cards settled as compatibility-only legacy data and taken out of the narrator's prompt, because they were the untracked path around knowledge authority that IMPORTED-KNOWLEDGE-DESIGN §73 already forbade; and no schema change at all, proved against a database M8's own code wrote. Three defects, found by running the milestone's own tests rather than by reading them. Deleting a campaign leaked its FTS index rows, and SQLite then handed the freed ids to the next source imported into any campaign, which failed with an integrity error that Reindex could not repair — both ends are closed, and a database already carrying the damage now repairs itself. An imported node with no state snapshot was being stamped with the campaign's head state, so an Undo to turn 2 showed what the story knew at turn 20. And the snapshot relink did not persist at all, because it mutated a dict in place on a column SQLAlchemy tracks by assignment: it looked correct in memory and wrote the wrong ids to disk. Carrying per-turn prompts looked like it would halve the length of campaign that can be restored. Measured — and after compressing them inside the file — everything M9 added costs 12% of it: the import ceiling moves from about 318 turns to about 279, against a 100-turn certification target. The dominant cost is not M9's at all. The per-position narrative state document is 74% of a bundle, and v2 already carried it. Backend 1,102 passed / 14 skipped / 0 failed. Frontend 145 passed. Lint, production build and Docker build clean. Verified across two server processes with two data directories, and in a real browser against a real narrator. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B --- .gitignore | 5 + DEVELOPMENT.md | 102 ++ README.md | 56 +- backend/app/backup.py | 278 +++ backend/app/bundle.py | 945 +++++++++- backend/app/context/builder.py | 67 +- backend/app/knowledge/fts.py | 44 +- backend/app/knowledge/importer.py | 22 + backend/app/knowledge/inject.py | 15 +- backend/app/main.py | 6 +- backend/app/routers/adventures/bundle_io.py | 59 +- backend/app/routers/adventures/crud.py | 8 + backend/app/routers/backups.py | 75 + backend/app/schemas.py | 19 + backend/app/starter.py | 4 +- backend/tests/m9_fixture.py | 406 +++++ backend/tests/test_bundle_v2.py | 25 +- backend/tests/test_imported_knowledge.py | 45 +- backend/tests/test_m9_backup.py | 451 +++++ backend/tests/test_m9_clean_import.py | 460 +++++ backend/tests/test_m9_corrupt_bundles.py | 747 ++++++++ backend/tests/test_m9_legacy_bundles.py | 365 ++++ backend/tests/test_m9_portability.py | 1186 ++++++++++++ backend/tests/test_retry_variants.py | 18 +- backend/tests/test_starter_adventure.py | 14 +- backend/tests/test_story_tree_baseline.py | 11 +- backend/tools/m9_migration_proof.py | 331 ++++ backend/tools/m9_portability_report.py | 354 ++++ backend/tools/m9_scale_report.py | 264 +++ frontend/src/api.js | 8 + frontend/src/components.jsx | 44 +- frontend/src/filePicker.test.jsx | 110 ++ frontend/src/pages/Settings.jsx | 101 + frontend/src/pages/backup.test.jsx | 116 ++ frontend/src/styles/library.css | 15 + planning/BROWSER-UX-SPEC.md | 29 + planning/BUILD-MILESTONES.md | 194 +- planning/DATA-MODEL.md | 112 ++ planning/IMPORTED-KNOWLEDGE-DESIGN.md | 41 + planning/README.md | 84 +- planning/TECHNICAL-DESIGN.md | 83 + planning/TEST-CAMPAIGN-FIXTURE.md | 114 ++ planning/V1-ACCEPTANCE-TESTS.md | 251 ++- planning/VERSION.md | 99 +- .../M8-IMPLEMENTATION-REPORT.md | 0 planning/reports/M9-IMPLEMENTATION-REPORT.md | 1622 +++++++++++++++++ 46 files changed, 9227 insertions(+), 178 deletions(-) create mode 100644 backend/app/backup.py create mode 100644 backend/app/routers/backups.py create mode 100644 backend/tests/m9_fixture.py create mode 100644 backend/tests/test_m9_backup.py create mode 100644 backend/tests/test_m9_clean_import.py create mode 100644 backend/tests/test_m9_corrupt_bundles.py create mode 100644 backend/tests/test_m9_legacy_bundles.py create mode 100644 backend/tests/test_m9_portability.py create mode 100644 backend/tools/m9_migration_proof.py create mode 100644 backend/tools/m9_portability_report.py create mode 100644 backend/tools/m9_scale_report.py create mode 100644 frontend/src/filePicker.test.jsx create mode 100644 frontend/src/pages/backup.test.jsx rename planning/{reports => archive/milestone-reports}/M8-IMPLEMENTATION-REPORT.md (100%) create mode 100644 planning/reports/M9-IMPLEMENTATION-REPORT.md diff --git a/.gitignore b/.gitignore index 770af61..59b2c05 100644 --- a/.gitignore +++ b/.gitignore @@ -9,6 +9,11 @@ __pycache__/ # Database *.db +# M9: verified database backups land beside the database. `*.db` already covers +# the files; this names the directory so its purpose is obvious in a listing and +# so nothing else that ends up there is committed by accident. +backend/backups/ +data/backups/ # Node node_modules/ diff --git a/DEVELOPMENT.md b/DEVELOPMENT.md index af08856..2cba4f5 100644 --- a/DEVELOPMENT.md +++ b/DEVELOPMENT.md @@ -328,6 +328,100 @@ subsystem comes back as a route, if an API key becomes settable again, if the model timeout stops being configurable or becomes unbounded, or if a supported start path stops binding loopback. +## Backing up, and getting a campaign back + +There are two recovery tools and they answer different questions. Using the +wrong one is the most common way to be surprised later, so they are described +together. + +| | Campaign export | Database backup | +| --- | --- | --- | +| Covers | one campaign | every campaign, and your settings | +| Shape | a JSON file you can read | a copy of the SQLite database | +| Moves between machines | **yes** — this is the supported way | no; it is this machine's database | +| Taken from | Export, on a campaign | Settings → *Back up everything on this machine* | +| Restored by | Import campaign, on the library screen | replacing the database file, below | + +### Exporting and importing a campaign + +Export is on each campaign in the library, and in the campaign's own Settings +panel. It writes one `.json` file holding the whole campaign: the story and its +entire retained tree, the branch you are on and **the exact position you are +reading at** — including one you undid back to — every alternate take, your Save +Points, the authoritative state and its per-position snapshots, the state +history that explains it, your imported knowledge with its classifications, the +summaries and memories, and the prompt each turn was actually given. + +Import is on the library screen and takes that file back, into this or any other +installation. Nothing about the file refers to the machine that wrote it: the +imported files come back from their content, not from a path, and no setting of +yours is changed by importing somebody's campaign. + +Two things it deliberately does **not** carry: your inference endpoint and model +settings, which describe your machine rather than the campaign, and the +rebuildable search indexes, which are rebuilt from the imported content before +the import returns. + +**A campaign imports whether or not the model that wrote it is installed here.** +Recovering a campaign and being able to play it on are separate questions; the +first never depends on the second. + +### Backing up the whole database + +Settings → Advanced → *Back up everything on this machine*. It writes a verified +copy into a `backups/` directory beside the database itself, and tells you where. + +It is a real backup rather than a file copy. It uses SQLite's online backup API, +so it is safe to take **while you are playing** — a `cp` of a live database can +read one page before a transaction and another after it, producing a file that +opens, reports a schema, and is quietly missing rows. The copy is checked with +`PRAGMA quick_check` before it is kept, an existing backup is never overwritten, +and a failure leaves nothing behind. + +You can also take one from the command line, or from `cron`: + +```bash +curl -s -X POST http://127.0.0.1:8000/api/backups | python3 -m json.tool +``` + +### Restoring a whole database + +There is deliberately no restore button, because restoring means replacing the +file the running application has open — which is how you lose both copies at +once. It is a three-step procedure and each step needs the application stopped: + +```bash +# 1. Stop the application. Nothing below is safe while it is running. +# (Ctrl-C the server, or `docker compose down`.) + +# 2. Keep what is there now, whatever state it is in. You may want it back. +mv backend/data.db backend/data.db.before-restore + +# 3. Put the backup in its place, and start the application again. +cp backend/backups/adventure-storyteller-20260907-043000.db backend/data.db +``` + +Check the file before you trust it, and check it again after starting: + +```bash +sqlite3 backend/backups/adventure-storyteller-20260907-043000.db 'PRAGMA quick_check;' +# -> ok +``` + +The database path is `backend/data.db` by default, and whatever `AIDND_DB_PATH` +names otherwise — in Docker that is the mounted volume. + +There is one file to move and no others: this build leaves SQLite in its default +rollback-journal mode, so there are no `-wal` or `-shm` companions beside the +database (`PRAGMA journal_mode` reports `delete`). A build that switched to WAL +would have to move those too, and leaving them behind would pair a new database +with an old write-ahead log. + +**Prefer the campaign export for anything smaller than "everything".** Restoring +a whole database rolls every campaign back to the moment the backup was taken, +including the ones you did not mean to touch. To recover one campaign, export it +and import it. + ## The context window your Ollama actually enforces **Check this before a long campaign.** The application budgets a prompt up to @@ -380,6 +474,14 @@ If you would rather not raise it at all, set **How much story to send** in Settings to the number `/api/ps` reports, and the prompt will be assembled to fit. +**This matters most on the machine you import to.** A campaign carries its +history, not the window the machine that wrote it had, and a long imported +campaign fills a prompt on its very first turn — so a deployment that has applied +neither the derived model above nor a matching budget meets its ceiling +immediately rather than gradually. Importing succeeds either way; it is the first +turn afterwards that truncates. Check `/api/ps` on the destination before playing +on an imported campaign, not after. + ## What was made offline-safe, and how to check Two runtime downloads were removed in Milestone M1. Both were invisible on a diff --git a/README.md b/README.md index a67ee1a..a3d4771 100644 --- a/README.md +++ b/README.md @@ -65,11 +65,16 @@ that isn't the live one starts a new branch. change is recorded with what it was before and which turn caused it, so the Story State panel can show what changed and why. You can correct it by hand, and your correction outranks the story. -- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world - info) are triggered by keywords in recent story text, then assembled under a token budget - (`backend/app/context/builder.py`). Story cards are the inherited authored-lore primitive and - are kept; they are **not** the knowledge library below, which is a first-class subsystem with - its own classification, provenance, chunking and index. +- **A context engine you can account for.** Memory, the author's note, the campaign's own + rules, the authoritative state, the summary that applies here, and the retrieved imported + passages are assembled under one token budget, in an order chosen so that a section which + changes does not re-price the cached prefix above it (`backend/app/context/builder.py`). + + **Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as + legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched + card used to arrive in front of it as a world fact with no class, no visibility, no source and + nothing to switch it off, competing with your imported Canon for the same budget; the knowledge + library below replaces it, and does all of that explicitly. - **Total prompt transparency.** Every turn stores the exact prompt sent to the model. **Inspect context** on any narrator turn opens a readable account of what it was given — what it remembered, what it read, what it believes, and what each part cost — with the @@ -122,13 +127,33 @@ that isn't the live one starts a new branch. no story, and deleting a branch a Save Point is kept on is refused until you remove the Save Point yourself, so nothing takes a named moment away behind your back. -- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree: - every branch, every take, the fork points, which branches the story has left behind, the Save - Points and the position it is being read at — all of them chosen rather than computed, which is - the rule for what a bundle carries. A campaign opens where its head says, never at a Save Point - merely because it has one. A campaign exported after two Undos imports still undone, with its - retained future intact, instead of silently reopening at its newest turn. Files that predate - the head position, and files saved in the old single-line format, still import. +- **Import and export, as a recovery contract.** A campaign exports as one JSON file, + `ai-dnd-adventure-v3`, and imports into a clean install on another machine. It carries the whole + tree — every branch, every take, the fork points, which branches the story has left behind, the + Save Points and the position it is being read at — and, since it is meant to be *recovery* + rather than a copy of the text, everything that explains that story: the authoritative state + and the typed events behind it, **the exact prompt each turn was given and the passages it was + shown**, the summaries with the coordinates that decide whether they still apply, and your + imported files with their classifications. A restored campaign can still answer "why does the + state say this?" and "what was the narrator actually told?" — after the source file has been + deleted and the canon edited since. + + A campaign opens where its head says, never at a Save Point merely because it has one. Exported + after two Undos, it imports still undone, with its retained future intact. Search indexes are + not carried: they are rebuilt from the content, before the import returns. Nothing about your + machine travels — no endpoint, no model, no path — so importing somebody's campaign never + reconfigures your inference, and a campaign imports whether or not you have the model that + wrote it. Older files still import: the flat single-line format, files that predate the head + position, and files that predate everything above. AI Dungeon-compatible scenario format is + still read and written for scenarios and story cards. +- **A verified backup of everything, taken while you play.** Settings → *Back up everything on + this machine* writes a copy of the whole database through SQLite's online backup API — not a + file copy, which of a live database can read one page before a transaction and another after it + and produce a file that opens and is quietly missing rows. It is checked with `PRAGMA + quick_check` before it is kept, and an existing backup is never overwritten + (`backend/app/backup.py`). Restoring one is a documented stop-move-start procedure in + `DEVELOPMENT.md`, deliberately not a button: replacing the file the running application has + open is how you lose both copies. - **Single user, no accounts.** There is no sign-up, no login, no session and no API key anywhere in the product. The storyteller API binds to loopback and is unauthenticated by design, because the only person who can reach it is the person running it. A new install @@ -234,8 +259,8 @@ leave it there. player input → assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions] + [plot essentials] + [story summary] + [retrieved memories] - + [triggered story cards] + [retrieved imported knowledge, - framed by class and bounded by its own budget] + + [retrieved imported knowledge, framed by class and + bounded by its own budget] + [history along this branch, token-budgeted] + [author's note] + [player action] → snapshot context (Insights) @@ -262,7 +287,8 @@ frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI ├─ worldstate/ the inherited RPG stat engine — legacy, no longer authoritative ├─ memorybank.py auto-summarization + embedding retrieval ├─ knowledge/ the imported library: import, chunk, FTS5, embed, rank, inject - ├─ bundle.py the export/import formats, v2 (tree) and a v1 reader + ├─ bundle.py the export/import formats: v3, and readers for v2 and v1 + ├─ backup.py a verified whole-database copy, via SQLite's backup API ├─ providers/ OpenAI-compatible adapter, streaming └─ data.db SQLite (path overridable via AIDND_DB_PATH) ``` diff --git a/backend/app/backup.py b/backend/app/backup.py new file mode 100644 index 0000000..13edb3f --- /dev/null +++ b/backend/app/backup.py @@ -0,0 +1,278 @@ +"""M9: a consistent copy of the whole database, taken while the app is running. + +This is **not** the campaign bundle, and the two are not alternatives. They are +different recovery tools and M9 keeps them apart deliberately: + + campaign bundle one campaign, logical, portable between installations, + importable into a clean data directory on another + machine, readable by a human and by a later build + database backup every campaign, every setting, physical, this machine, + restored by putting the file back + +The bundle is the primary cross-install recovery path and is what the acceptance +tests measure. This exists for the other question: the reader has one database +holding everything they have ever played, and wants a copy of it before they +upgrade, move a disk, or try something they might regret. + +## Why not `cp data.db backup.db` + +Because a copy taken with the application running is a copy of a moving target. +SQLite writes a database in pages, and a plain file copy can read page 5 before +a transaction and page 900 after it — the result is a file that opens, reports a +schema, and is silently missing or duplicating rows. In WAL mode it is worse: the +committed data may be in a `-wal` file the copy never touched. Nothing warns +anyone. The corruption is found later, by which time the original may be gone. + +So this uses SQLite's own **online backup API** (`sqlite3.Connection.backup`), +which is the supported mechanism for exactly this: it copies page by page while +holding the right locks, restarts if a write moves the source underneath it, and +produces a file that is a transactionally consistent snapshot of some committed +point. The application keeps running throughout; no session is closed and no +turn is blocked. + +## What the procedure guarantees + +1. The source database is opened **read-only** and is never written to. A backup + that could damage what it is backing up would be worse than no backup. +2. The copy is written to a temporary file beside the destination and renamed + into place only after it has been verified, so an interrupted or failed run + never leaves a half-written file wearing a backup's name. `os.replace` is + atomic on the same filesystem, which is why the temporary sits in the + destination's own directory rather than in `/tmp`. +3. `PRAGMA quick_check` runs against the finished copy, opened as its own + database, before it is renamed. A backup nobody verified is a belief. +4. An existing file is never overwritten. Each run writes a new name stamped + with the time, so yesterday's backup survives today's mistake — which is most + of what a backup is for. +5. Failure is reported and leaves nothing behind but the log line. + +## What it does not do + +There is no restore endpoint. Restoring a whole database means replacing the +file the running application has open, and doing that from inside that +application is a way to lose both copies. The procedure is in `DEVELOPMENT.md`: +stop the app, move the file into place, start it. Campaign-level recovery — the +common case, and the one that crosses machines — is the bundle. + +No path comes from a caller. The destination directory is derived from the +database the application is already using and the filename is generated here, so +there is no request that can direct a write anywhere else (H08). +""" + +from __future__ import annotations + +import logging +import os +import sqlite3 +from dataclasses import dataclass +from datetime import datetime +from pathlib import Path + +from .database import DB_PATH + +log = logging.getLogger(__name__) + +#: Where backups go: a directory beside the database itself. Beside, rather than +#: inside a configurable location, because the one thing this must not do is +#: write somewhere a request can name. +DIRECTORY_NAME = "backups" + +#: The stem every backup file carries, so a directory listing sorts by date and +#: says what these files are without being opened. +PREFIX = "adventure-storyteller" + + +class BackupError(RuntimeError): + """A backup did not complete. The source database is untouched.""" + + +@dataclass(frozen=True) +class Backup: + """One finished, verified backup file.""" + + path: Path + bytes: int + pages: int + seconds: float + integrity: str + + def as_dict(self) -> dict: + return { + # The name alone, not the path. The full path is a fact about this + # machine's filesystem, and the reader is told the directory once by + # the endpoint that lists them. + "filename": self.path.name, + "bytes": self.bytes, + "pages": self.pages, + "seconds": round(self.seconds, 3), + "integrity": self.integrity, + } + + +def directory(db_path: Path | None = None) -> Path: + """The backup directory for a database, created if it does not exist.""" + root = (db_path or DB_PATH).parent / DIRECTORY_NAME + root.mkdir(parents=True, exist_ok=True) + return root + + +def create(db_path: Path | None = None, *, now: datetime | None = None) -> Backup: + """Takes one verified backup of the live database, and returns it. + + Raises `BackupError` on any failure, having removed whatever it had written. + The source database is opened read-only and is never modified, so a failure + here costs the backup and nothing else. + """ + source_path = db_path or DB_PATH + if not source_path.exists(): + raise BackupError(f"There is no database at {source_path}.") + stamp = (now or datetime.now()).strftime("%Y%m%d-%H%M%S") + target = _unused_name(directory(source_path), stamp) + # The temporary sits in the destination directory so the rename below is a + # rename rather than a copy across filesystems, which would not be atomic. + working = target.with_name(target.name + ".partial") + started = datetime.now() + try: + pages = _copy(source_path, working) + integrity = _verify(working) + except BackupError: + _discard(working) + raise + except Exception as exc: # noqa: BLE001 - reported, never raised raw + _discard(working) + log.exception("Backup of %s failed", source_path) + raise BackupError(f"{type(exc).__name__}: {exc}") from exc + size = working.stat().st_size + # Only now does the file get the name a reader would trust. + os.replace(working, target) + return Backup( + path=target, + bytes=size, + pages=pages, + seconds=(datetime.now() - started).total_seconds(), + integrity=integrity, + ) + + +def _copy(source_path: Path, working: Path) -> int: + """Runs SQLite's online backup from `source_path` into a new file. + + The source is opened through a URI with `mode=ro`, so this connection cannot + write to it even by accident. The destination is a fresh database that this + function creates; `backup()` overwrites whatever is in it, and the caller has + guaranteed the name is unused. + + Returns the number of pages copied, which is the one honest measure of how + much was actually written — the file size counts pages the source had + already allocated. + """ + source = sqlite3.connect(f"file:{source_path}?mode=ro", uri=True) + try: + destination = sqlite3.connect(working) + try: + copied = 0 + + def progress(_status, remaining, total): + nonlocal copied + copied = total - remaining + + # `pages=-1` copies the whole database in one step while holding the + # source's read lock, which is the right trade for a local + # single-user database: it is the fastest option, it cannot restart + # partway, and the lock it holds does not block readers. + source.backup(destination, pages=-1, progress=progress) + return copied + finally: + destination.close() + finally: + source.close() + + +def _verify(working: Path) -> str: + """Runs `PRAGMA quick_check` against the finished copy. + + Opened as its own connection, so what is checked is the file on disk rather + than any page cache the copy left behind. `quick_check` rather than + `integrity_check` because it does the structural work — every page reachable, + every record readable — without the full index cross-check, which on a large + database is minutes rather than moments. A backup nobody verified is a + belief; a backup verified slowly enough that nobody takes one is worse. + """ + connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True) + try: + rows = connection.execute("PRAGMA quick_check").fetchall() + finally: + connection.close() + result = ", ".join(str(row[0]) for row in rows) if rows else "no result" + if result != "ok": + raise BackupError( + f"The backup was written but did not verify: {result}. It has been " + f"discarded; the original database is untouched." + ) + return result + + +def _unused_name(root: Path, stamp: str) -> Path: + """A name in `root` that nothing is using. + + An existing backup is never overwritten. Two backups taken inside one second + are the only way to collide, and the counter settles that rather than one of + them silently replacing the other. + """ + candidate = root / f"{PREFIX}-{stamp}.db" + counter = 2 + while candidate.exists() or candidate.with_name(candidate.name + ".partial").exists(): + candidate = root / f"{PREFIX}-{stamp}-{counter}.db" + counter += 1 + return candidate + + +def _discard(working: Path) -> None: + """Removes a partial file, ignoring a file that is already gone.""" + try: + working.unlink() + except OSError: + pass + + +def existing(db_path: Path | None = None) -> list[dict]: + """Every backup in the directory, newest first. + + Names and sizes only. Reading one to report what is inside it would mean + opening a database on every page load for a screen that is a list. + + `taken_at` is read out of the **filename**, which is the stamp `create` + wrote when it took the backup, and falls back to the file's modification + time only for a name that does not parse. The two usually agree, and where + they disagree the name is the one telling the truth: copying a backup to + another disk, restoring it from an archive, or touching it all move the + mtime, and a list that then reordered itself would report when the file was + last handled rather than when the backup was taken. + """ + root = directory(db_path) + rows = [] + for path in root.glob(f"{PREFIX}-*.db"): + try: + stat = path.stat() + except OSError: + continue + rows.append({ + "filename": path.name, + "bytes": stat.st_size, + "taken_at": ( + _stamp_in(path.name) or datetime.fromtimestamp(stat.st_mtime) + ).isoformat(timespec="seconds"), + }) + rows.sort(key=lambda row: (row["taken_at"], row["filename"]), reverse=True) + return rows + + +def _stamp_in(filename: str) -> datetime | None: + """The time in a backup's name, or `None` if it does not carry one.""" + rest = filename[len(PREFIX) + 1:].removesuffix(".db") + # A collision within one second gets a `-2` suffix, which is not the stamp. + stamp = "-".join(rest.split("-")[:2]) + try: + return datetime.strptime(stamp, "%Y%m%d-%H%M%S") + except ValueError: + return None diff --git a/backend/app/bundle.py b/backend/app/bundle.py index e19c8b8..2e13a59 100644 --- a/backend/app/bundle.py +++ b/backend/app/bundle.py @@ -50,23 +50,141 @@ Every hand-editable coordinate is therefore checked before a row is written, in `plan`, rather than repaired afterwards. An import that fails partway leaves an adventure holding half a tree, and a tree missing a branch is a story that stops without reporting anything. + +## Version 3, and why the version was bumped (M9) + +Version 3 adds the evidence that explains a campaign, rather than more of the +campaign itself: + +* `stateEvents` and `stateProposals` — the audit half of the hybrid + (`DATA-MODEL.md` §17). Version 2 carried the per-position snapshots, so an + imported campaign could be *read* at any position but could not say why the + state there was what it was, and a manual correction was indistinguishable + from something the story established. +* a per-node `contextSnapshot` — the exact prompt a turn was given, the + passages it was shown and their text, and the model and generation settings it + ran under. This is `SPECIFICATION.md` §6.1's "exact prompt/context snapshot or + reproducible equivalent", and it is the one thing here that genuinely cannot + be recreated: an imported source can be deleted, canon can be edited, the + chunker can change, and the evidence of what an old turn was actually told has + to survive all of it (`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50). +* `summaries` — the generated and hand-written rolling summaries, each with the + coordinate that decides whether it is eligible. Version 2 carried only the + `storySummary` mirror, which has no lineage, so a restored campaign resumed + with no usable long-story continuity at all. +* a per-node `parent` — which attempts belong to the same turn (SP9). Version 2 + left every imported node parentless, and the coordinate fallback merges the + attempts under two different takes of the turn above into one pager. +* per-memory `authority`, and per-source `parserVersion`/`chunkingVersion`. + +### Why this is a version and not five more optional keys + +Everything above could have been added to version 2 the way `persona`, +`checkpoints`, `narrativeState` and `knowledge` were: read with `.get`, absent +meaning the campaign had none. That mechanism is real and this module has used +it four times. + +It stops working here, for the reason the `headDepth` rule above exists. A +version 2 file with no `contextSnapshot` anywhere is **ambiguous**: it may have +been written before M9, when no bundle could carry one, or by M9 from a campaign +whose turns predate the column. Those are different facts and a reader has to be +able to tell them apart — the same distinction I07 draws when it says a file +written before the head was carried opens at the tip *because tip was the only +position that format could represent*. A version number is how a recovery file +states what it was capable of recording. So the writer bumps, and the reader +keeps every older version. + +### What each version can be trusted to say + + v1 a linear story, its turns, and its retries as a repeating group + v2 + the tree, the live flags, the after-snapshots, the chosen head, + Save Points, the narrative state document, imported knowledge + v3 + state events and proposals, historical prompt/context provenance, + lineage-anchored summaries, take parentage, memory authority + +An older file is never reinterpreted under a newer rule. A v2 file has no state +events because v2 could not carry them, not because the campaign had none, and +the import says so rather than inventing an audit trail. + +## Three kinds of data, and the rule for each + +M9 states the rule the module has been applying, because the additions above are +the first ones where the second and third categories differ: + +1. **Chosen.** What a person decided: the story, the head, the takes, the Save + Points, the classifications, the canon. It travels. +2. **Historical evidence.** What happened, and what the application was told at + the time. It travels *even though* something resembling it could be + regenerated, because a regeneration would describe today's campaign rather + than the turn it claims to explain. Historical evidence is not a cache. +3. **Rebuildable derived data.** A deterministic function of what travels: + knowledge passages, the FTS index, embeddings, the branch lineage cache. It + does not travel, and the import rebuilds it. + +The test that separates 2 from 3 is not "could this be recomputed" but "would a +recomputation answer the same question". Rebuilding the FTS index answers the +same question it answered before. Rebuilding an old turn's prompt does not: it +would say what that turn *would be told now*, which is the opposite of what the +inspector is for. + +## Why the snapshots are compressed inside the file + +A per-turn prompt contains the story so far, so storing one per turn costs +O(turns²) in a campaign's length. That is already true of the database — +`compression.py` records the column as 89% of production storage and compresses +it for exactly this reason — and carrying the snapshots makes it true of the +export too. Measured on this build with a 16,384-token budget: 20,797 bytes of +snapshot per turn at turn 20, 55,291 at turn 120, and 68% of a 9.7 MB file. +Against `limits.MAX_IMPORT_BODY_BYTES` a campaign would stop being importable — +stop being *recoverable* — somewhere around its 170th turn. + +So the snapshot travels as `contextSnapshotZ`: the same JSON, zlib-compressed by +`compression.pack` and base64-encoded so the file is still JSON. It is the same +bytes on the way out as on the way in; nothing is dropped, summarised or +reconstructed, and `compression.py`'s existing pack/unpack is the only +implementation. The measured cost falls to about a quarter, which moves the +ceiling out by roughly the square root of that. + +Everything else in the file stays plain, readable JSON — the story, the tree, +the head, the Save Points, the state events, the imported source text. The one +field that is encoded is the machine-assembled prompt archive, which nobody +reads by eye in a text editor and which the context inspector shows properly. + +`contextSnapshot`, holding the plain object, is still **read** on import and +takes precedence when both keys are present. A file somebody hand-edited to make +a prompt legible has to keep importing, and the plain form is what they will +have written. """ +import base64 +import binascii +import zlib from datetime import datetime from fastapi import HTTPException from sqlalchemy import insert, update from sqlalchemy.orm import Session, undefer -from . import attempts, models, schemas +from . import attempts, compression, memorybank, models, schemas, summaries from .context import cursors, lineage from .knowledge import chunking as knowledge_chunking from .knowledge import classes as knowledge_classes from .knowledge import importer as knowledge_importer from .narrative import model as narrative_model -FORMAT = "ai-dnd-adventure-v2" +#: What `export` writes. The family name is inherited from the production base +#: and is deliberately unchanged: renaming it would break every reader for no +#: gain, and `PROVENANCE.md` is where the fork is recorded. +FORMAT = "ai-dnd-adventure-v3" + +#: Every version the importer still accepts, newest first. A backup that stops +#: importing is not a backup, so this list only ever grows. +TREE_FORMAT = "ai-dnd-adventure-v2" LEGACY_FORMAT = "ai-dnd-adventure-v1" +READABLE = (FORMAT, TREE_FORMAT, LEGACY_FORMAT) + +#: The versions that carry the tree — everything except the flat v1 shape. +TREE_FORMATS = (FORMAT, TREE_FORMAT) # `actions.type` is VARCHAR(20), and a raw-dict import bypasses the schemas. TYPE_MAX = 20 @@ -75,17 +193,27 @@ TYPE_MAX = 20 # ---------------------------------------------------------------- exporting def export(db: Session, adventure: models.Adventure) -> dict: - """Returns the whole adventure as a version 2 bundle. + """Returns the whole adventure as a version 3 bundle. The export is not scoped to a path. A backup holds the entire tree, not the - branch its owner is currently reading. Both after-snapshots are undeferred in - the one query, because they are per-node columns nothing else reads in bulk - and requesting them a row at a time would cost one query per turn. + branch its owner is currently reading. Every deferred per-node column is + undeferred in the one query, because they are columns nothing else reads in + bulk and requesting them a row at a time would cost one query per turn — and + since M9 that includes `context_snapshot`, which is the largest of them. - The bundle carries no context snapshots. It has never carried the assembled - prompts, and it still does not. They run about 163 kB per turn, they explain - a generation rather than form part of the story, and the Insights viewer they - feed reads the adventure they came from. + **The bundle now carries the context snapshots (M9).** Every version before + this one left them out, on the reasoning that they explain a generation + rather than form part of the story and that the viewer reading them reads + the adventure they came from. That reasoning holds right up to the moment + the campaign moves to another machine, which is what this milestone is + about: an imported campaign with no snapshots has no historical prompt + provenance at all, so `SPECIFICATION.md` §6.1's per-turn record is local to + the machine that played it. See the module docstring on the three kinds of + data — this is evidence, not a cache. + + They are stored once per turn on the live attempt (`attempts.py`), so a + campaign with many retries pays for the prompt once and for the few hundred + bytes of per-attempt slices per take. The measured cost is in the M9 report. """ branches = ( db.query(models.Branch) @@ -103,12 +231,15 @@ def export(db: Session, adventure: models.Adventure) -> dict: .options( undefer(models.Action.state_after), undefer(models.Action.world_state_after), + undefer(models.Action.narrative_state_after), + undefer(models.Action.context_snapshot), ) .order_by( models.Action.branch_id, models.Action.depth, models.Action.id, ) .all() ) + knowledge_sources = list(adventure.knowledge_sources) return { "format": FORMAT, "title": adventure.title, @@ -181,7 +312,30 @@ def export(db: Session, adventure: models.Adventure) -> dict: # # A bundle written before M7 has no key here and imports with an empty # library, which is what such a campaign had. - "knowledge": [_exported_source(k) for k in adventure.knowledge_sources], + "knowledge": [_exported_source(k) for k in knowledge_sources], + # M6/M9. The lineage-anchored summaries, which version 2 did not carry. + # A summary is derived from the story, so the earlier rule left it out — + # but it is derived by a *model*, at a cost, from a stretch of story that + # may since have been abandoned, and regenerating one on the importing + # machine would produce different prose describing a different reading. + # It is provenance-bearing derived data, and the coordinate on it is + # what decides eligibility (`CONTEXT-AND-MEMORY.md`; E03), so the row + # travels and its vectors do not. + "summaries": [_exported_summary(s, local) for s in adventure.summaries], + # M5/M9. The audit half of the hybrid. `DATA-MODEL.md` §17 keeps the + # events for audit and the snapshots for restore, and version 2 carried + # only the snapshots — so a moved campaign could be read at any position + # and could no longer say what changed there, who asserted it, or what + # the value was before. A manual correction was the worst case: it is + # the one state change no narration explains, so with the events gone + # nothing distinguished it from something the story established. + # + # Both tables travel, because the inspector reads both: the event says + # what was accepted, and the proposal beside it says what the model + # actually asked for and what was refused. Exporting only the accepted + # half would keep the answers and lose every question. + "stateProposals": _exported_proposals(db, adventure, local), + "stateEvents": _exported_events(db, adventure, local), "actions": [_exported_node(a, local) for a in nodes], } @@ -189,6 +343,13 @@ def export(db: Session, adventure: models.Adventure) -> dict: def _exported_source(source: models.KnowledgeSource) -> dict: """One knowledge source, as it goes into the file.""" return { + # M9. The id this source had in the campaign that wrote the file. It is + # not an identity the import reuses — the row gets a new id like + # everything else — but the historical retrieval records inside the + # context snapshots name it, and without it a restored inspector could + # not tell which of five imported sources an old turn was shown. The + # import maps it forward; see `_relink_snapshot`. + "sourceId": source.id, "title": source.title, "originalFilename": source.original_filename, "classification": source.classification, @@ -198,11 +359,107 @@ def _exported_source(source: models.KnowledgeSource) -> dict: "contentHash": source.content_hash, "mediaType": source.media_type, "notes": source.notes, + # M9. What produced the passages this source last had. The import + # rebuilds the passages with *this* build's chunker and records its own + # versions, so these two say what the source was chunked by when it was + # exported — which is how a later reader can tell that a historical + # retrieval record describes passages a newer parser would not produce. + "parserVersion": source.parser_version, + "chunkingVersion": source.chunking_version, "importedAt": source.imported_at.isoformat() if source.imported_at else None, "content": source.content, } +def _exported_summary(summary: models.Summary, local: dict[int, int]) -> dict: + """One summary, with the coordinate that decides whether it is eligible.""" + return { + "text": summary.text, + "branch": _local(summary.branch_id, local), + "depth": summary.depth, + "sourceStart": summary.source_start, + "sourceEnd": summary.source_end, + "trigger": summary.trigger, + "model": summary.model_name, + "createdAt": summary.created_at.isoformat() if summary.created_at else None, + } + + +def _exported_proposals( + db: Session, adventure: models.Adventure, local: dict[int, int], +) -> list[dict]: + """Every state proposal, in the order they were made. + + `id` is the row's identity in the source campaign, exported for the reason + `_exported_node` exports one: the events below name it, and the import needs + something to translate. It is a *file-local* identity — the imported row + gets a fresh id — and the only guarantee it carries is that references + inside one file agree with each other. + + `detail` is deferred and compressed, and it is the largest thing here — a + busy turn's rejections — so it is undeferred once for the whole list rather + than touched per row. + """ + rows = ( + db.query(models.StateProposal) + .filter(models.StateProposal.adventure_id == adventure.id) + .options(undefer(models.StateProposal.detail)) + .order_by(models.StateProposal.id) + .all() + ) + return [ + { + "id": row.id, + "action": row.action_id, + "branch": local.get(row.branch_id) if row.branch_id is not None else None, + "depth": row.depth, + "model": row.model_name, + "source": row.source, + "status": row.status, + "rawOutput": row.raw_output, + "detail": row.detail, + "createdAt": row.created_at.isoformat() if row.created_at else None, + } + for row in rows + ] + + +def _exported_events( + db: Session, adventure: models.Adventure, local: dict[int, int], +) -> list[dict]: + """Every accepted state event, in the order they were applied. + + `proposal` and `action` name rows by the identity those rows export, and the + import translates both. An event whose proposal or action is missing from + the file keeps its coordinate and loses the pointer, which is what a + `SET NULL` and a `CASCADE` respectively already mean in the schema. + """ + rows = ( + db.query(models.StateEvent) + .filter(models.StateEvent.adventure_id == adventure.id) + .order_by(models.StateEvent.id) + .all() + ) + return [ + { + "proposal": row.proposal_id, + "action": row.action_id, + "branch": local.get(row.branch_id) if row.branch_id is not None else None, + "depth": row.depth, + "sequence": row.sequence, + "eventType": row.event_type, + "payload": row.payload, + "before": row.before, + # Who asserted this. `manual_correction` is the value that has to + # survive: it is what C04 and the audit view use to say the user + # overruled the story rather than the story establishing it. + "source": row.source, + "createdAt": row.created_at.isoformat() if row.created_at else None, + } + for row in rows + ] + + _ROOT = {"parent": None, "forkDepth": None} @@ -246,6 +503,23 @@ def _exported_node(action: models.Action, local: dict[int, int]) -> dict: "text": action.text, "createdAt": action.created_at.isoformat() if action.created_at else None, } + # M9. Which turn this node is an attempt at (SP9). A coordinate cannot + # answer it: attempts under two different takes of the turn above share a + # branch and a depth, and an attempt forked onto its own line leaves the + # coordinate its siblings still hold. Version 2 exported neither, so every + # imported node landed parentless and `attempts.group` fell back to the + # coordinate — which merges those two groups into one pager reading 5/5 + # where the reader should see 2/2 and 3/3. + # + # The value is the parent's own database id, translated on the way in. A + # root node has none, and so does a row written before SP9. + if action.parent_id is not None: + node["parentId"] = action.parent_id + # M9. The node's own id, which is what `parentId` above and the audit + # records refer to. Exported rather than derived from list position because + # a hand-edited file that reorders `actions` would otherwise silently + # re-parent every take. + node["id"] = action.id if action.reasoning: node["reasoning"] = action.reasoning # `{}` and an absent key mean different things. `{}` means the node left an @@ -267,14 +541,64 @@ def _exported_node(action: models.Action, local: dict[int, int]) -> dict: node["stateChanges"] = action.state_changes if action.world_delta: node["worldDelta"] = action.world_delta + # M9. What this turn was actually given: the assembled prompt section by + # section, the passages retrieved and their rendered text, which summary was + # eligible, and the model and generation settings the call ran under. + # + # This is the one field here that is neither a decision somebody made nor a + # deterministic function of the rows beside it. It is evidence, and the + # module docstring says why that is a third category: regenerating a + # historical prompt would produce what the turn *would* be told now, from + # today's canon, today's sources and today's state, which is precisely the + # question the inspector does not ask. + # + # Stored once per turn on the live attempt, so a superseded take carries + # only `attempts.ATTEMPT_KEYS` — its own reply, its own state proposal and + # its own token accounting — and a retried turn is not a multiplier on the + # largest thing in the file. + # + # Compressed, for the size reason in the module docstring. The plain form is + # still accepted on the way in. + if action.context_snapshot is not None: + node["contextSnapshotZ"] = _packed(action.context_snapshot) return node +def _packed(snapshot: dict) -> str: + """One context snapshot as base64-encoded zlib, for the file.""" + return base64.b64encode(compression.pack(snapshot)).decode("ascii") + + +def _unpacked(value) -> dict | None: + """Reads a `contextSnapshotZ` field back, or `None` if it cannot be read. + + A snapshot that will not decode is dropped rather than raising, which is the + rule `compression.CompressedJSON.process_result_value` already applies to + the same data in the database: the evidence for one turn is worth less than + the campaign, and a turn with no snapshot is a state the inspector has + always been able to render. It is never *partially* decoded — the whole + value is one compressed document, so it arrives intact or not at all. + """ + if not isinstance(value, str): + return None + try: + unpacked = compression.unpack(base64.b64decode(value, validate=True)) + except (binascii.Error, zlib.error, UnicodeDecodeError, ValueError): + return None + return unpacked if isinstance(unpacked, dict) else None + + def _exported_memory(memory: models.Memory, local: dict[int, int]) -> dict: return { "text": memory.text, "pinned": memory.pinned, "forgotten": memory.forgotten, "sourceStart": memory.source_start, "sourceEnd": memory.source_end, "useCount": memory.use_count, + # M9. How much weight the narrator may give this (`CONTEXT-AND-MEMORY.md` + # §14, F07). `accepted_story` is something the story established; + # `heuristic` is a reading of it. Version 2 carried neither, so every + # imported memory landed on the column default — which silently promoted + # a heuristic memory to accepted story, the one direction F07 forbids. + "authority": memory.authority, # The node this memory is attached to. A hand-written memory summarizes # no node, so it has a branch and no depth, and it keeps that shape # here. @@ -343,13 +667,30 @@ def _exported_anchor( # ---------------------------------------------------------------- importing def check_format(bundle: dict) -> str: - """Returns the bundle's version, or raises a 400.""" + """Returns the bundle's version, or raises a 400. + + Every version this build has ever written is accepted, newest first, because + a backup that stops importing is not a backup. A version this build has never + heard of is refused rather than read as the newest one it knows: a file from + a later build may use a key this one would misread, and guessing is how an + import silently loses half a campaign. + """ + if not isinstance(bundle, dict): + raise HTTPException(400, "This file does not contain an adventure export.") fmt = bundle.get("format") - if fmt in (FORMAT, LEGACY_FORMAT): + if fmt in READABLE: return fmt + if isinstance(fmt, str) and fmt.startswith("ai-dnd-adventure-v"): + raise HTTPException( + 400, + f"This file was written in format {fmt}, which this version of the " + f"application does not understand. The newest format it can read is " + f"{FORMAT}.", + ) raise HTTPException( 400, - f"Not an adventure export file (expected format {FORMAT} or {LEGACY_FORMAT}).", + "Not an adventure export file (expected format " + + ", ".join(READABLE) + ").", ) @@ -361,48 +702,81 @@ def plan(bundle: dict, version: str) -> dict: about the shape of a tree, because the alternative is an import that fails partway and leaves an adventure holding a story with a gap in it. - Both versions produce the same shape, so `write` never learns that there are - two formats. A version 1 bundle is a linear story, which is a tree with one - branch, and its `variants` array is a sibling group written the old way. + Every version produces the same shape, so `write` never learns that there + are three formats. A version 1 bundle is a linear story, which is a tree + with one branch, and its `variants` array is a sibling group written the old + way; a version 2 bundle is a version 3 bundle with the evidence sections + absent, and absent means the format could not carry them. """ - branches = ( - _planned_branches(bundle) if version == FORMAT else [dict(_ROOT)] - ) + tree = version in TREE_FORMATS + branches = _planned_branches(bundle) if tree else [dict(_ROOT)] nodes = ( - _planned_nodes(bundle, len(branches)) if version == FORMAT + _planned_nodes(bundle, len(branches)) if tree else _planned_v1_nodes(bundle) ) head = _as_index(bundle.get("headBranch"), len(branches), default=0) + # M9. Which node in the file each exported node id refers to. Built here so + # that `parentId` and every audit reference can be checked against it before + # a row is written, and so a reference the file cannot satisfy is a refusal + # rather than a dangling pointer discovered later. + by_id = _planned_node_ids(nodes) return { "branches": branches, "nodes": nodes, "memories": _planned_memories(bundle, len(branches)), - # M4. Empty for a version 1 bundle and for any version 2 bundle written + # M4. Empty for a version 1 bundle and for any tree bundle written # before Save Points existed, which is the same answer: no one had named # a position in those campaigns. "checkpoints": ( - _planned_checkpoints(bundle, len(branches), nodes) - if version == FORMAT else [] + _planned_checkpoints(bundle, len(branches), nodes) if tree else [] ), "head": head, # None means the file does not say, which is every version 1 bundle and # every version 2 bundle written before M3. `_point_the_head` derives it # then, which is what those files were written expecting. "headDepth": _planned_head_depth(bundle, branches, nodes, head), - # Version 2 records where the derived work reached. Version 1 counted - # it, and a count cannot become a node until the nodes exist. See - # `settle`. - "anchors": _planned_anchors(bundle, len(branches)) if version == FORMAT else None, - "positions": None if version == FORMAT else { + # Versions 2 and 3 record where the derived work reached. Version 1 + # counted it, and a count cannot become a node until the nodes exist. + # See `settle`. + "anchors": _planned_anchors(bundle, len(branches)) if tree else None, + "positions": None if tree else { "memory": _as_int(bundle.get("memoryCursor"), 0), "summary": _as_int(bundle.get("summaryCursor"), 0), }, # M7. Checked here with everything else, before a row is written, so a # hand-edited library fails the import rather than half-landing in it. "knowledge": _planned_knowledge(bundle), + # M9. Three sections version 3 introduced. Each is empty for an older + # file, and that is a statement about the format rather than about the + # campaign: a version 2 bundle has no state events because version 2 + # could not carry any, and the import does not invent an audit trail to + # fill the gap. + "summaries": _planned_summaries(bundle, len(branches)) if tree else [], + "proposals": _planned_proposals(bundle, len(branches), by_id), + "events": _planned_events(bundle, len(branches), by_id), } +def _planned_node_ids(nodes: list[dict]) -> dict[int, int]: + """Maps each exported node id to its position in the planned node list. + + A file that names the same node id twice is refused. The ids are what the + take parentage and the whole audit trail hang off, so two rows claiming one + identity would silently attach half the evidence to the wrong turn. + """ + by_id: dict[int, int] = {} + for position, node in enumerate(nodes): + node_id = node.get("id") + if node_id is None: + continue + if node_id in by_id: + raise HTTPException( + 400, f"Two actions in this file both call themselves {node_id}." + ) + by_id[node_id] = position + return by_id + + def _planned_knowledge(bundle: dict) -> list[dict]: """The knowledge sources in a bundle, checked and normalized. @@ -472,10 +846,141 @@ def _planned_knowledge(bundle: dict) -> list[dict]: "content": content, "content_hash": actual, "hash_mismatch": isinstance(stated, str) and bool(stated) and stated != actual, + # M9. The id this source had where the file was written, so the + # retrieval records inside the historical context snapshots can be + # pointed at the row that arrives here. Not an identity: two files + # imported side by side will both have named their sources 1, 2, 3. + "source_id": _int_or_none(entry, "sourceId"), }) return planned +def _planned_summaries(bundle: dict, branches: int) -> list[dict]: + """The summaries in a bundle, with the coordinates that make them eligible. + + A summary with no usable coordinate is dropped rather than refused, and + dropped rather than repaired. The reasoning is `_planned_checkpoints`': + losing one summary costs a paragraph of derived prose, and refusing the + whole file over it would lose the campaign to save the paragraph. Repairing + it would be worse than either — a summary placed at a guessed coordinate is + a summary that may become eligible on a line it does not describe, which is + the E03 leak arriving by a new route. + """ + raw = bundle.get("summaries") + out: list[dict] = [] + for entry in raw if isinstance(raw, list) else []: + if not isinstance(entry, dict) or not str(entry.get("text") or "").strip(): + continue + branch = _as_index(entry.get("branch"), branches, default=None) + depth = entry.get("depth") + if branch is None or not _is_int(depth): + continue + out.append({ + "text": str(entry["text"]), + "branch": branch, + "depth": depth, + "sourceStart": _int_or_none(entry, "sourceStart"), + "sourceEnd": _int_or_none(entry, "sourceEnd"), + "trigger": str(entry.get("trigger") or "interval")[:20], + "model": str(entry.get("model") or "")[:200], + "createdAt": _as_time(entry.get("createdAt")), + }) + return out + + +def _planned_proposals( + bundle: dict, branches: int, by_id: dict[int, int] +) -> list[dict]: + """The state proposals in a bundle, checked against the tree beside them. + + A proposal names the node whose narration produced it. A proposal naming a + node this file does not contain is **refused**, not dropped: unlike a Save + Point, an audit record that quietly did not arrive leaves a campaign whose + state cannot be explained, and it is the explanation the reader would go + looking for precisely when something looks wrong. A file disagreeing with + itself about which turn asserted something is a file whose audit trail is + not worth restoring at all. + + A proposal with **no** action is a different thing and is accepted: that is + a manual correction, which has a coordinate and no narration behind it. + """ + raw = bundle.get("stateProposals") + if raw is None: + return [] + if not isinstance(raw, list): + raise HTTPException(400, "The state history in this file is not a list.") + out: list[dict] = [] + for i, entry in enumerate(raw): + if not isinstance(entry, dict): + raise HTTPException(400, f"State proposal {i + 1} is not an object.") + out.append({ + "id": _int_or_none(entry, "id"), + "action": _planned_action_ref(entry.get("action"), by_id, + f"State proposal {i + 1}"), + "branch": _as_index(entry.get("branch"), branches, default=None), + "depth": _int_or_none(entry, "depth"), + "model": str(entry.get("model") or "")[:200], + "source": str(entry.get("source") or "accepted_story")[:40], + "status": str(entry.get("status") or "accepted")[:30], + "rawOutput": str(entry.get("rawOutput") or ""), + "detail": _dict_or_none(entry, "detail"), + "createdAt": _as_time(entry.get("createdAt")), + }) + return out + + +def _planned_events( + bundle: dict, branches: int, by_id: dict[int, int] +) -> list[dict]: + """The accepted state events in a bundle, checked the same way. + + `proposal` is resolved against the proposals in the same file at write time + rather than here, because a proposal's identity is its own exported id and + the planner has both lists. An event naming a proposal the file does not + contain keeps its coordinate and loses the pointer, which is what the + schema's `ON DELETE SET NULL` already says happens when a proposal goes + away. An event naming a *node* the file does not contain is refused, for + the reason above. + """ + raw = bundle.get("stateEvents") + if raw is None: + return [] + if not isinstance(raw, list): + raise HTTPException(400, "The state events in this file are not a list.") + out: list[dict] = [] + for i, entry in enumerate(raw): + if not isinstance(entry, dict): + raise HTTPException(400, f"State event {i + 1} is not an object.") + event_type = entry.get("eventType") + if not isinstance(event_type, str) or not event_type.strip(): + raise HTTPException(400, f"State event {i + 1} has no type.") + out.append({ + "proposal": _int_or_none(entry, "proposal"), + "action": _planned_action_ref(entry.get("action"), by_id, + f"State event {i + 1}"), + "branch": _as_index(entry.get("branch"), branches, default=None), + "depth": _int_or_none(entry, "depth"), + "sequence": _as_int(entry.get("sequence"), 0), + "eventType": event_type[:60], + "payload": _dict_or_none(entry, "payload"), + "before": _dict_or_none(entry, "before"), + "source": str(entry.get("source") or "accepted_story")[:40], + "createdAt": _as_time(entry.get("createdAt")), + }) + return out + + +def _planned_action_ref(value, by_id: dict[int, int], subject: str) -> int | None: + """Resolves an audit record's node reference to a position in the plan.""" + if value is None: + return None + if not _is_int(value) or value not in by_id: + raise HTTPException( + 400, f"{subject} names action {value!r}, which is not in this file." + ) + return by_id[value] + + def _derived_tip(branches: list[dict], nodes: list[dict], head: int) -> int: """Returns where the head branch's story ends, which is where a file that does not state a head depth is opened. @@ -638,6 +1143,29 @@ def _planned_nodes(bundle: dict, branches: int) -> list[dict]: "worldStateAfter": _as_dict(entry.get("worldStateAfter")), "worldDelta": _as_dict(entry.get("worldDelta")), "createdAt": _as_time(entry.get("createdAt")), + # M9. The node's identity within this file, and the turn it is an + # attempt at. Both are absent from every version 2 bundle, and a + # node with no id simply takes part in nothing that refers to one: + # `attempts.group` falls back to the coordinate for it, exactly as + # it does for a pre-SP9 row, which is the rule such a file was + # written under. + "id": _int_or_none(entry, "id"), + "parentId": _int_or_none(entry, "parentId"), + # M9. Historical evidence, restored verbatim. Nothing here is + # regenerated and nothing is validated beyond its being an object: + # it is a record of what an older build assembled, and imposing + # today's expectations on it would be the retroactive reading the + # inspector exists to rule out. The one thing the import does touch + # is the source ids inside it, which name rows on the machine that + # wrote the file; see `_relink_snapshots`. + # + # The plain key wins over the compressed one when a file carries + # both, because somebody who hand-edited a prompt into readable JSON + # meant the readable one. + "contextSnapshot": ( + _as_dict(entry.get("contextSnapshot")) + or _unpacked(entry.get("contextSnapshotZ")) + ), }) return nodes @@ -680,6 +1208,14 @@ def _planned_v1_nodes(bundle: dict) -> list[dict]: "worldStateAfter": None, "worldDelta": None, "createdAt": _as_time(variant.get("createdAt")), + # A version 1 file names no node and carries no prompt. Both + # keys are present so that `_write_nodes` reads one shape for + # every version; `None` here means the format could not say, + # which for parentage puts the row on `attempts.group`'s + # coordinate fallback — the rule it was written under. + "id": None, + "parentId": None, + "contextSnapshot": None, }) return nodes @@ -702,7 +1238,25 @@ def _planned_memories(bundle: dict, branches: int) -> list[dict]: # memory on a branch nothing can see never reaches a prompt # again. "branch": _as_index(entry.get("branch"), branches, default=0), - "depth": entry.get("depth") if _is_int(entry.get("depth")) else None, + "depth": _int_or_none(entry, "depth"), + # M9. `heuristic` or `accepted_story` (F07). An unrecognised value + # is read as `heuristic`, which is the safe direction: the failure + # this guards against is a guess about the story being presented to + # the narrator as something the story established, and an + # accepted-story memory demoted to a hedged one costs a little + # weight in ranking rather than a false fact. + # + # A file with no key here is every bundle before version 3, and it + # gets the column default. That is not a repair — the value simply + # was not recorded, and re-classifying the text now would be this + # build's reading of a memory an older one wrote. + "authority": ( + entry["authority"] + if entry.get("authority") in (memorybank.ACCEPTED_STORY, + memorybank.HEURISTIC) + else memorybank.HEURISTIC if "authority" in entry + else None + ), }) return out @@ -773,24 +1327,73 @@ def _planned_anchors(bundle: dict, branches: int) -> dict: # ------------------------------------------------------------------ writing -def write(db: Session, adventure: models.Adventure, story: dict) -> None: +def write(db: Session, adventure: models.Adventure, story: dict) -> dict: """Writes a planned tree onto a newly created adventure. - The order is fixed. Branches come first, because a node needs a branch id. - The nodes come next, because the head and the anchors name a node. + The order is fixed, and every step in it needs something the step before it + produced. Branches come first, because a node needs a branch id. The nodes + come next, because the head, the anchors, the audit records and the take + parentage all name a node. The knowledge comes before the snapshots are + relinked, because the relink needs the ids the sources land on. + + Returns what the caller has to report to the reader: the knowledge sources + whose rebuildable index could not be built. The authoritative campaign is + committed either way — see `_write_knowledge` — and the difference between + "the import failed" and "the import succeeded and the search index did not" + is one a reader has to be told, not one to leave in a column. """ ids = _write_branches(db, adventure, story["branches"]) - _write_nodes(db, adventure, story["nodes"], ids) + rows = _write_nodes(db, adventure, story["nodes"], ids) + # The nodes need ids of their own before anything can point at them. One + # flush covers the parentage, the audit records and the Save Points. + db.flush() + _link_take_parents(story["nodes"], rows) _write_memories(db, adventure, story["memories"], ids) _point_the_head(adventure, story, ids) _write_checkpoints(db, adventure, story["checkpoints"], ids) _write_anchors(adventure, story, ids) - _write_knowledge(db, adventure, story.get("knowledge") or []) + _write_summaries(db, adventure, story.get("summaries") or [], ids) + _write_state_history(db, adventure, story, ids, rows) + sources = _write_knowledge(db, adventure, story.get("knowledge") or []) + # Last, because it needs both halves: the nodes carrying the snapshots and + # the knowledge rows the snapshots' retrieval records point at. + _relink_snapshots(rows, sources) + return { + "knowledge_index_failures": [ + {"title": source.title, "detail": source.index_detail} + for source in sources.values() if source.index_state == "failed" + ], + } + + +def _link_take_parents(specs: list[dict], rows: list[models.Action]) -> None: + """Points each imported node at the turn it is an attempt at (M9, SP9). + + The file names the parent by the id it had where the file was written, so + this walks the pair of lists once and translates. A node whose parent is not + in the file is left with none — that is a root node, a pre-SP9 row, or an + older bundle that carried no parentage at all, and for all three + `attempts.group` falls back to the coordinate, which is the rule they were + written under. + + The parentage is deliberately not validated into a refusal. A wrong parent + costs a take pager that groups oddly; a refused import costs the campaign. + """ + by_exported_id = { + spec["id"]: row for spec, row in zip(specs, rows) + if spec.get("id") is not None + } + for spec, row in zip(specs, rows): + parent = by_exported_id.get(spec.get("parentId")) + # A node cannot be its own parent, and `attempts.group` would loop on + # one that was. A file claiming it is treated as claiming nothing. + if parent is not None and parent is not row: + row.parent_id = parent.id def _write_knowledge( db: Session, adventure: models.Adventure, specs: list[dict] -) -> None: +) -> dict[int, models.KnowledgeSource]: """Restores the imported library, and rebuilds the index it needs. The bundle carries the source and not its passages, so this is where they @@ -812,6 +1415,7 @@ def _write_knowledge( failure is visible on the source and in the knowledge status endpoint, and Reindex is the repair. """ + landed: dict[int, models.KnowledgeSource] = {} for spec in specs: source = models.KnowledgeSource( adventure_id=adventure.id, @@ -845,11 +1449,183 @@ def _write_knowledge( ).strip() db.add(source) db.flush() + if spec.get("source_id") is not None: + landed[spec["source_id"]] = source try: knowledge_importer.build_index(db, source) except Exception as exc: # noqa: BLE001 - recorded, not raised source.index_state = "failed" source.index_detail = f"{type(exc).__name__}: {exc}"[:2000] + return landed + + +def _write_summaries( + db: Session, adventure: models.Adventure, specs: list[dict], ids: list[int] +) -> None: + """Restores the rolling summaries onto the coordinates that hold them. + + Nothing about eligibility is decided here, and that is the point. A summary + is eligible when its coordinate lies on the active capped lineage, which + `summaries.current` works out from the tree at read time — so restoring the + coordinate restores the answer, including the answer "no" for a summary + belonging to a line this campaign has left (E03). + + `adventures.story_summary` is the reader-facing mirror and is set from the + bundle's own `storySummary` by `materialize`. `settle_the_mirror` corrects it + afterwards from whichever summary the restored head is actually entitled to, + because a mirror that disagrees with the lineage is exactly the M6 defect + that made summaries rows in the first place. + """ + for spec in specs: + summary = models.Summary( + adventure_id=adventure.id, + text=spec["text"], + branch_id=ids[spec["branch"]], + depth=spec["depth"], + source_start=spec["sourceStart"], + source_end=spec["sourceEnd"], + trigger=spec["trigger"], + model_name=spec["model"], + ) + if spec["createdAt"] is not None: + summary.created_at = spec["createdAt"] + db.add(summary) + + +def _write_state_history( + db: Session, adventure: models.Adventure, story: dict, ids: list[int], + rows: list[models.Action], +) -> None: + """Restores the audit half of the hybrid: proposals, then the events. + + Proposals first, because an event names one. Both are written as they were + recorded and neither is replayed: `DATA-MODEL.md` §17 is explicit that the + events are the audit record and the snapshots are the restore path, so + nothing here recomputes a state document and nothing here checks that the + events add up to the snapshot beside them. They are two accounts of the same + history and the import is restoring both, not reconciling them. + + A rejected proposal is restored exactly like an accepted one. It changed + nothing then and it changes nothing now — it is the record that explains why + the state does not say what the narration seems to say, which is the case + where a reader most needs it. + """ + proposals: dict[int, models.StateProposal] = {} + for spec in story.get("proposals") or []: + proposal = models.StateProposal( + adventure_id=adventure.id, + action_id=_action_id(spec["action"], rows), + branch_id=ids[spec["branch"]] if spec["branch"] is not None else None, + depth=spec["depth"], + model_name=spec["model"], + source=spec["source"], + status=spec["status"], + raw_output=spec["rawOutput"], + detail=spec["detail"], + ) + if spec["createdAt"] is not None: + proposal.created_at = spec["createdAt"] + db.add(proposal) + if spec["id"] is not None: + proposals[spec["id"]] = proposal + # The events name proposals, so the proposals need ids first. + if proposals: + db.flush() + for spec in story.get("events") or []: + proposal = proposals.get(spec["proposal"]) + event = models.StateEvent( + adventure_id=adventure.id, + proposal_id=proposal.id if proposal is not None else None, + action_id=_action_id(spec["action"], rows), + branch_id=ids[spec["branch"]] if spec["branch"] is not None else None, + depth=spec["depth"], + sequence=spec["sequence"], + event_type=spec["eventType"], + payload=spec["payload"], + before=spec["before"], + source=spec["source"], + ) + if spec["createdAt"] is not None: + event.created_at = spec["createdAt"] + db.add(event) + + +def _action_id(position: int | None, rows: list[models.Action]) -> int | None: + """The database id of the node at `position` in the planned list.""" + return rows[position].id if position is not None else None + + +def _relink_snapshots( + rows: list[models.Action], sources: dict[int, models.KnowledgeSource] +) -> None: + """Points a restored snapshot's retrieval records at the sources that arrived. + + The snapshot itself is evidence and is restored verbatim — the prompt, the + passages, the text each one supplied, the settings the call ran under. What + is *not* evidence is `source_id`: it is a pointer at a row on the machine + that wrote the file, and the imported library holds the same sources under + different ids. Left alone it would point the "open this source" control in + the context inspector at whatever happens to hold that id here, which is + either nothing or somebody else's file. + + So the pointer is translated and the evidence is not. Where the file's own + knowledge section contains the source, the record names the row it landed + on; where it does not — a source deleted before the export was taken — the + record is left saying `null`, and the inspector shows the filename as plain + text rather than a link. That is the honest answer: this passage came from a + file that is no longer here, and here is what it said. + + `chunk_id` is deliberately untouched. Passages are rebuilt on import and get + new ids, so no translation exists; the record's `chunk_index` and + `heading_path` still say which passage of the source it was, and the + rendered text is in the record either way. + + ## A new document, never a mutation in place + + `context_snapshot` is a plain `CompressedJSON` column and not a + `MutableDict`, so SQLAlchemy tracks it by assignment and not by content. + Editing the dict the attribute already holds and assigning it back changes + nothing: at flush time the loaded value and the current value are the same + object, the history reports no net change, and no UPDATE is emitted. The + import then looks correct in memory and writes the untranslated ids to disk. + + So a modified snapshot is built as a new document and assigned. The copy is + shallow at every level this touches — a new outer dict, a new `knowledge` + dict, new lists, and a new dict per record — and shares everything it does + not change, including the rendered passage text, which is the large part. + """ + if not sources: + # Nothing to translate against. Every record keeps whatever it says, + # which for a campaign with no library is nothing that resolves anyway. + return + for row in rows: + snapshot = row.context_snapshot + if not isinstance(snapshot, dict): + continue + knowledge = snapshot.get("knowledge") + if not isinstance(knowledge, dict): + continue + rebuilt = {} + touched = False + for key in ("used", "dropped", "suppressed"): + records = knowledge.get(key) + if not isinstance(records, list): + continue + replaced = [] + for record in records: + if not isinstance(record, dict) or "source_id" not in record: + replaced.append(record) + continue + landed = sources.get(record["source_id"]) + replaced.append( + dict(record, source_id=landed.id if landed is not None else None) + ) + touched = True + rebuilt[key] = replaced + if touched: + row.context_snapshot = dict( + snapshot, knowledge=dict(knowledge, **rebuilt) + ) def _write_branches( @@ -898,13 +1674,18 @@ def _write_branches( def _write_nodes( db: Session, adventure: models.Adventure, specs: list[dict], ids: list[int] -) -> None: +) -> list[models.Action]: """Writes the nodes, grouped into the turns they are attempts at. + Returns the rows in the order the specs were given, so that the caller can + match each one back to the file entry that produced it — the take parentage + and every audit record need exactly that correspondence. + One value is decided here rather than read from the file. Exactly one attempt in each group is made live, because a file can name none or several, and a turn with no live node disappears from the story. """ + rows: list[models.Action] = [] groups: dict[tuple[int, int], list[models.Action]] = {} for spec in specs: key = (spec["branch"], spec["depth"]) @@ -919,18 +1700,43 @@ def _write_nodes( state_after=spec["stateAfter"], world_state_after=spec["worldStateAfter"], world_delta=spec["worldDelta"], - narrative_state_after=spec.get("narrativeStateAfter"), + # M9. Written explicitly, including when the file carries none. + # + # Leaving it NULL hands the decision to `tree.stamp_outcome`, which + # runs on every flush and fills a missing snapshot from *the + # campaign's current state*. On an import that is the state at the + # exported head, so every position whose snapshot the file did not + # carry would come back holding the newest position's state — an + # Undo landing on turn 2 showing what the story knew at turn 20, + # which is exactly the failure M5's review finding 3 was about, + # arriving by a different route. + # + # The empty document is the honest value: this position established + # nothing that the file records. It is also what `restore_state` + # falls back to for a NULL and what migration 88 backfilled for + # pre-M5 rows, so a legacy bundle lands exactly where it landed + # before — the fallback wrote `empty()` for those files too, because + # a campaign with no `narrativeState` has no head state to copy. + narrative_state_after=( + spec.get("narrativeStateAfter") + if spec.get("narrativeStateAfter") is not None + else narrative_model.empty() + ), state_changes=spec.get("stateChanges"), + # M9. Restored exactly as it was written. See `_planned_nodes`. + context_snapshot=spec.get("contextSnapshot"), ) if spec["createdAt"] is not None: action.created_at = spec["createdAt"] db.add(action) + rows.append(action) groups.setdefault(key, []).append(action) - for rows in groups.values(): - live = next((row for row in rows if row.live), rows[0]) - for row in rows: + for group in groups.values(): + live = next((row for row in group if row.live), group[0]) + for row in group: row.live = row is live + return rows def _write_memories( @@ -948,6 +1754,8 @@ def _write_memories( branch_id=ids[spec["branch"]], depth=spec["depth"], ) + if spec.get("authority") is not None: + memory.authority = spec["authority"] # A version 1 memory has no depth of its own. `source_end` is the index # of the last action it summarizes, which on one branch is that node's # depth. @@ -1052,7 +1860,7 @@ def settle(adventure: models.Adventure, story: dict) -> None: def materialize( db: Session, payload: dict, story: dict, user_id: int | None -) -> models.Adventure: +) -> tuple[models.Adventure, dict]: """Writes a planned bundle into a new adventure owned by `user_id`. Call `check_format` and `plan` first, and apply any rate or size limits @@ -1060,9 +1868,14 @@ def materialize( import endpoint alone: the guest starter writes a file the server ships, so it has nothing to rate-limit and no untrusted list to cap. - The caller commits. Two flushes happen here, because the anchors and the - legacy counts describe one boundary in two coordinate systems, and aligning - them needs the actions to be queryable. + Returns the adventure and a report of the derived work that did not + complete. **The caller commits**, and everything below happens in that one + transaction: an import that raises anywhere in here leaves no adventure at + all rather than half of one (L01's rule, applied to the import path). + + Two flushes happen here, because the anchors and the legacy counts describe + one boundary in two coordinate systems, and aligning them needs the actions + to be queryable. """ # A raw-dict import bypasses the schemas, so truncate strings bound for # VARCHAR columns. Postgres enforces the widths. See `schemas.py`. @@ -1104,11 +1917,32 @@ def materialize( # story, its tree, its cards and its memories still import intact, which is # what the bundle is for. - write(db, adventure, story) + report = write(db, adventure, story) db.flush() db.expire(adventure, ["actions"]) settle(adventure, story) - return adventure + settle_the_mirror(db, adventure) + return adventure, report + + +def settle_the_mirror(db: Session, adventure: models.Adventure) -> None: + """Points `story_summary` at whichever restored summary the head is entitled to. + + The bundle carries the mirror column as well as the rows, because it always + has and because a v1 or v2 file carries nothing else. But the column has no + lineage of its own, and the value that was correct where the file was written + is only correct here if the head landed in the same place — which it does, + but relying on that would be relying on a coincidence rather than on the + rule. M6's finding M6-F1 was precisely a mirror that had drifted from the + lineage, so the import ends by asking the lineage. + + A campaign whose file carried no summary rows keeps whatever the column + said: an older bundle's mirror is the only summary it has, and blanking it + would lose the one thing that file recorded. + """ + if not adventure.summaries: + return + summaries.refresh_mirror(db, adventure) # ------------------------------------------------------------------ reading # Small coercions. A raw-dict import bypasses the schemas, so every value from a @@ -1142,6 +1976,23 @@ def _as_dict(value) -> dict | None: return value if isinstance(value, dict) else None +def _int_or_none(entry: dict, key: str) -> int | None: + """An optional integer field, or `None` when the file gives something else. + + Distinct from `_as_int`, which substitutes a default: the fields this serves + are ones where absent and unusable mean the same thing — *the file does not + say* — and where inventing a number would be inventing a coordinate. + """ + value = entry.get(key) + return value if _is_int(value) else None + + +def _dict_or_none(entry: dict, key: str) -> dict | None: + """An optional object field, or `None` when the file gives something else.""" + value = entry.get(key) + return value if isinstance(value, dict) else None + + def _as_time(value) -> datetime | None: if not isinstance(value, str): return None diff --git a/backend/app/context/builder.py b/backend/app/context/builder.py index c066f0d..8c85e31 100644 --- a/backend/app/context/builder.py +++ b/backend/app/context/builder.py @@ -29,7 +29,9 @@ from ..knowledge import records as knowledge_records from . import encoding, history AUTHORS_NOTE_DEPTH = 3 # actions from the end of history -CARD_BUDGET_SHARE = 0.4 # max share of non-reserved budget that story cards may take +# `CARD_BUDGET_SHARE = 0.4` was here, and is gone with the injection it bounded +# (M9). It is named rather than deleted silently because two other places +# reasoned about their own share against it. NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene") SEPARATOR = "\n\n" @@ -508,30 +510,47 @@ def build_context( adventure, available_after_knowledge, count_tokens, exclude_action_id ) - # ----- Story cards: triggered by recent story text (the window history could fill) ----- - trigger_window = truncate_to_last_tokens( - SEPARATOR.join(a.text for a in actions), available_after_knowledge - ) - triggered = match_cards(adventure.story_cards, trigger_window) - - card_budget = int(available_after_knowledge * CARD_BUDGET_SHARE) - card_records = [] - lore_lines: list[str] = [] + # ----- Story cards: legacy, and no longer part of the narrator's prompt (M9) + # + # Until M9 a keyword-triggered story card was injected here as + # `World Lore: `, taking up to 40% of what was left after the + # imported knowledge had been placed. + # + # `IMPORTED-KNOWLEDGE-DESIGN.md` §73 settles that Story Cards are not the + # production imported-knowledge store and, in as many words, that they "must + # not become an alternate untracked path around the new knowledge + # authority/provenance rules". That is exactly what this was. A card entry + # arrived in front of the narrator as a world fact with: + # + # * no class — nothing said whether it was Canon, Reference or Inspiration, + # so nothing framed how far the narrator could rely on it; + # * no visibility — no narrator-only distinction at all; + # * no source, no hash, no lifecycle, nothing to disable it with; + # * no browser surface, since M8 removed the editor — so a reader could + # neither see it nor switch it off; + # * and no row in the context inspector, which renders `knowledge` and + # never rendered `cards`. + # + # It also competed with imported Canon for one budget, which is the + # arrangement M7 spent a milestone separating. + # + # M9's decision, recorded in the milestone report: story cards are + # **compatibility-only legacy data**. Nothing is deleted. The rows stay, the + # `/api/story-cards` endpoints stay, the bundle carries them out and back so + # a round trip destroys nothing, and `memorybank.cast_brief` still reads them + # as the summariser's character roster — a roster names who is on stage so a + # memory says "Aldric" rather than "he", it never reaches the narrator, and + # every memory written from it is authority-classified by the application + # afterwards. What stops is the one path that asserted campaign facts to the + # narrator without any of the controls §73 requires. + # + # `cards` stays in the report and is now always empty for a new turn. + # Removing the key would break the historical snapshots that have one, which + # M9 has just made portable: an old turn's evidence says story cards were + # included, and it must go on saying so. + card_records: list[dict] = [] + lore_section = None used = 0 - for match in triggered: - line = f"World Lore: {match['entry'].strip()}" - tokens = count_tokens(line) - included = used + tokens <= card_budget - if included: - lore_lines.append(line) - used += tokens - card_records.append( - {"id": match["id"], "name": match["name"], "keyword": match["keyword"], - "included": included} - ) - lore_section = ( - Section("world_lore", "\n".join(lore_lines)) if lore_lines else None - ) # ----- Story history: newest first until the remaining budget is spent ----- history_budget = available_after_knowledge - used diff --git a/backend/app/knowledge/fts.py b/backend/app/knowledge/fts.py index 479e762..21c77ca 100644 --- a/backend/app/knowledge/fts.py +++ b/backend/app/knowledge/fts.py @@ -137,13 +137,53 @@ def index_line(heading_path: str, text_: str) -> str: def add(db: Session, chunk_id: int, heading_path: str, text_: str) -> None: - """Indexes one passage. The caller supplies the chunk's id as the rowid.""" + """Indexes one passage. The caller supplies the chunk's id as the rowid. + + `OR REPLACE`, and the reason is a defect M9 found rather than a defensive + habit. The rowid is a chunk's primary key, so a row already sitting at it is + by definition stale: the chunk that owned it does not exist, or is being + rewritten by the reindex that called this. Either way the new passage is the + truth and the old row is not. + + Without it, an orphaned index row makes an ordinary import fail. SQLite + reuses primary keys once the highest row is gone, so the next campaign to + import a source is handed rowid 1 again, collides with an orphan, and gets a + 500 from `INSERT` — and `clear_index` cannot clear the orphan, because it + finds index rows *through* the chunks, and there are none. That made Reindex, + which is the documented repair, unable to repair this. `REPLACE` closes it + from both ends: a leaked row is overwritten the moment the id comes round + again, so an existing database repairs itself rather than needing a + migration, and Reindex is the repair it is described as. + + The leak itself is closed separately, in `importer.clear_campaign_index`. + """ db.execute( - sql(f"INSERT INTO {TABLE} (rowid, text) VALUES (:id, :text)"), + sql(f"INSERT OR REPLACE INTO {TABLE} (rowid, text) VALUES (:id, :text)"), {"id": chunk_id, "text": index_line(heading_path, text_)}, ) +def remove_adventure(db: Session, adventure_id: int) -> int: + """Drops every index row belonging to one campaign. Returns how many. + + Scoped through the chunks, which is the only place the campaign is + recorded — the index deliberately holds no copy of it + (see "The table" above). So this has to run **before** the chunk rows go, + which is what `importer.clear_campaign_index` is for. + """ + result = db.execute( + sql( + f""" + DELETE FROM {TABLE} WHERE rowid IN ( + SELECT id FROM knowledge_chunks WHERE adventure_id = :adventure_id + ) + """ + ), + {"adventure_id": adventure_id}, + ) + return result.rowcount or 0 + + def remove_chunks(db: Session, chunk_ids: list[int]) -> None: """Drops passages from the index by id. diff --git a/backend/app/knowledge/importer.py b/backend/app/knowledge/importer.py index 76a9410..53ed6c2 100644 --- a/backend/app/knowledge/importer.py +++ b/backend/app/knowledge/importer.py @@ -362,6 +362,28 @@ def clear_index(db: Session, source: models.KnowledgeSource) -> None: db.expire(source, ["chunks"]) +def clear_campaign_index(db: Session, adventure: models.Adventure) -> int: + """Removes a whole campaign's lexical index rows. Returns how many. + + Called before a campaign is deleted, and it has to be: the FTS index is a + virtual table, so no foreign key reaches it and no `ON DELETE CASCADE` + covers it. Deleting a campaign cascades `knowledge_sources` to + `knowledge_chunks` and stops there, leaving one index row per passage + belonging to a chunk that no longer exists. + + Found in M9. The leak is not cosmetic. SQLite hands out the lowest free + primary key, so once the highest chunk is gone the *next* source imported + into *any* campaign is given a chunk id that an orphan already occupies, and + the import fails with an integrity error — a 500 on an ordinary upload, in a + campaign that has nothing to do with the deleted one. `fts.add` now repairs + such a collision when it meets one; this stops it happening. + + Vectors and passages need no equivalent, because both are real tables whose + foreign keys cascade. + """ + return fts.remove_adventure(db, adventure.id) + + def delete_source(db: Session, source: models.KnowledgeSource) -> None: """Removes a source and everything derived from it. diff --git a/backend/app/knowledge/inject.py b/backend/app/knowledge/inject.py index 76bed59..353051c 100644 --- a/backend/app/knowledge/inject.py +++ b/backend/app/knowledge/inject.py @@ -50,10 +50,17 @@ from .records import Candidate, Result #: Share of the non-protected budget that retrieved knowledge may spend. #: -#: Story cards already take up to 40% (`CARD_BUDGET_SHARE`), and the history is -#: what is left. A third is enough for several passages at the chunker's -#: typical size and leaves the majority of the window to the story itself, -#: which is the thing the reader came for. +#: A third is enough for several passages at the chunker's typical size and +#: leaves the majority of the window to the story itself, which is the thing the +#: reader came for. +#: +#: This share was chosen when story cards could take up to 40% of the same +#: budget and the history took what was left. M9 removed that injection +#: (`IMPORTED-KNOWLEDGE-DESIGN.md` §73), so the history now gets that 40% back. +#: The number here is deliberately unchanged: a third of the budget was chosen +#: as the right amount of *imported material* to put in front of the narrator, +#: not as a leftover, and raising it because room appeared would be changing +#: retrieval behaviour under cover of a portability milestone. KNOWLEDGE_SHARE = 0.33 #: What each class may take of the knowledge budget. Canon may take all of it; diff --git a/backend/app/main.py b/backend/app/main.py index 4f372db..032a085 100644 --- a/backend/app/main.py +++ b/backend/app/main.py @@ -10,7 +10,9 @@ from starlette.exceptions import HTTPException as StarletteHTTPException from .database import engine from .limits import BodySizeLimitMiddleware from .migrations import bootstrap -from .routers import adventures, chat, debug, scenarios, settings, story_cards +from .routers import ( + adventures, backups, chat, debug, scenarios, settings, story_cards, +) from .seed import seed_public_scenarios bootstrap(engine) @@ -112,6 +114,8 @@ app.include_router(scenarios.router) app.include_router(adventures.router) app.include_router(story_cards.router) app.include_router(settings.router) +# M9: a verified copy of the whole database, taken while the app is running. +app.include_router(backups.router) app.include_router(chat.router) app.include_router(debug.router) diff --git a/backend/app/routers/adventures/bundle_io.py b/backend/app/routers/adventures/bundle_io.py index b94dd36..966b63d 100644 --- a/backend/app/routers/adventures/bundle_io.py +++ b/backend/app/routers/adventures/bundle_io.py @@ -1,7 +1,31 @@ """Exporting an adventure to a bundle, and importing one back. `app/bundle.py` owns the format and the version handling. These two endpoints -only check ownership and hand the work over. +only check ownership, apply the caps, and hand the work over. + +## Why the import is one transaction and two phases + +`bundle.plan` reads the whole file and returns a checked, normalised tree +without opening a session, touching a row or creating an adventure. Everything a +hand-edited file can get wrong about its own shape — a node on a branch that is +not listed, a fork from a branch listed after it, a head past the story, an +audit record naming a turn that is not there — is a 400 from a function with no +side effects. + +Only then does `bundle.materialize` write, and it writes inside the single +transaction this endpoint commits at the end. So there are exactly two outcomes +a caller can see, and M9 requires them to be distinguishable: + + the authoritative import failed 4xx, and no campaign exists + the authoritative import succeeded 201, and the campaign is complete + +A third state — the campaign landed and a *rebuildable* index did not — is not a +failure of the import and does not roll it back. Passages, the lexical index and +vectors are all a deterministic function of content the file carries, so losing +them costs a rebuild rather than data. It is reported on the response as a +warning, it is visible per source in the Knowledge panel, and Reindex is the +repair. Refusing a whole campaign because a search index would not build would +trade the valuable thing for the cheap one. """ from fastapi import Body, Depends, Request @@ -18,15 +42,15 @@ def export_adventure( db: Session = Depends(get_db), adv: models.Adventure = Depends(current_adventure), ): - """Returns a full backup: plot components, story cards, scripts, state, and tree. + """Returns a full backup: the story, the tree, the state, and the evidence. - `app/bundle.py` owns the format, in both of its versions. A backup outlives - the schema, so no call site decides anything about its shape. + `app/bundle.py` owns the format, in all three of its versions. A backup + outlives the schema, so no call site decides anything about its shape. """ return bundle.export(db, adv) -@router.post("/import", response_model=schemas.AdventureOut, status_code=201) +@router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201) def import_adventure( request: Request, payload: dict = Body(...), @@ -58,18 +82,29 @@ def import_adventure( branches=story["branches"], ) - adventure = bundle.materialize(db, payload, story, user.id) - - db.commit() + try: + adventure, report = bundle.materialize(db, payload, story, user.id) + db.commit() + except Exception: + # Explicit, rather than left to the session closing. The planner has + # already refused everything it can see, so anything raising here is a + # write that surprised us — the case where leaving a partial campaign + # behind would be worst, and the case a test can only assert on if the + # rollback is a statement rather than a side effect of teardown. + db.rollback() + raise db.refresh(adventure) # A campaign exported while undone imports undone (M3), so the history # controls have to be right on the response that opens it — otherwise the # first thing the reader sees about a story with a retained future is a # greyed-out Redo. - out = schemas.AdventureOut.model_validate(adventure) + out = schemas.ImportedAdventureOut.model_validate(adventure) out.can_undo = head.can_undo(db, adventure) out.can_redo = head.can_redo(db, adventure) - # This is not a funnel step. A returning player imports a bundle, so it - # says nothing about how far a first-time visitor got. It is counted anyway, - # because it is the clearest evidence that anyone uses the export format. + out.import_warnings = [ + f"The search index for “{failure['title']}” could not be rebuilt " + f"({failure['detail']}). The file itself imported intact — use Reindex " + f"in the Knowledge panel to try again." + for failure in report["knowledge_index_failures"] + ] return out diff --git a/backend/app/routers/adventures/crud.py b/backend/app/routers/adventures/crud.py index 3e88958..71dada2 100644 --- a/backend/app/routers/adventures/crud.py +++ b/backend/app/routers/adventures/crud.py @@ -14,6 +14,8 @@ from ... import ( worldstate, ) from ...database import get_db +from ...knowledge import embeddings as knowledge_embeddings +from ...knowledge import importer as knowledge_importer from .deps import CurrentUser, current_adventure, router from .paging import action_window, annotate_takes @@ -348,8 +350,14 @@ def delete_adventure( db: Session = Depends(get_db), adventure: models.Adventure = Depends(current_adventure), ): + # M9. The lexical index first, while the chunks that locate it still exist. + # It is a virtual table, so nothing cascades into it, and an orphaned index + # row makes the *next* import into *any* campaign fail — see + # `knowledge.importer.clear_campaign_index`. + knowledge_importer.clear_campaign_index(db, adventure) db.delete(adventure) db.commit() # No later request reads this adventure's vectors, so drop them now. The # cache would otherwise hold them until the process restarted. memorybank.forget_cached_vectors(adventure_id) + knowledge_embeddings.forget_cached(adventure_id) diff --git a/backend/app/routers/backups.py b/backend/app/routers/backups.py new file mode 100644 index 0000000..596962c --- /dev/null +++ b/backend/app/routers/backups.py @@ -0,0 +1,75 @@ +"""M9: taking a verified copy of the whole database, from the browser. + +Two endpoints and no third. `app/backup.py` owns the procedure and every +guarantee it makes; these only decide who may ask. + +## Why there is no restore endpoint, and no download + +**Restore** means replacing the database file the running process has open. +Doing that from inside that process is how someone loses both copies at once: +the connection pool still holds handles on the old file, the WAL belongs to the +old file, and a half-swapped database is not something a running application can +notice. The supported procedure is in `DEVELOPMENT.md` — stop the application, +move the file into place, start it — and it is a procedure precisely because +each step needs the application not to be running. Campaign-level recovery, the +common case and the only one that crosses machines, is the export bundle. + +**Download** is not offered either. The file is a copy of every campaign on the +machine, and streaming it through the browser would put it in the download +directory, in the browser's own cache, and in whatever the reader does with it +next — for a local single-user application whose whole premise is that the story +does not leave the machine, that is a worse default than a path the reader can +copy. So the response names the directory and the reader takes it from there. + +## Where the file goes + +Nowhere a request can name. The destination is derived from the database the +application is already using, and the filename is generated from the clock. No +part of either comes from the caller, so there is no traversal to attempt (H08), +and the endpoints below accept no body at all. +""" + +import logging + +from fastapi import APIRouter, Depends, HTTPException + +from .. import auth, backup, models + +router = APIRouter(prefix="/api/backups", tags=["backups"]) + +log = logging.getLogger(__name__) + + +@router.get("") +def list_backups(_user: models.User = Depends(auth.get_current_user)): + """The backups already on disk, newest first, and where they are. + + The directory is reported once here rather than on every row, because it is + the same for all of them and it is what the reader needs in order to find + the files at all. + """ + return { + "directory": str(backup.directory()), + "backups": backup.existing(), + } + + +@router.post("", status_code=201) +def create_backup(_user: models.User = Depends(auth.get_current_user)): + """Takes one verified backup, and reports what it wrote. + + Synchronous. A backup of a local single-user database is a page copy that + finishes in well under a second, and a reader who pressed the button is + entitled to be told whether it worked rather than to be told it started. + + A failure is a 500 carrying the reason. There is nothing for the caller to + fix by retrying differently — the request has no parameters — so the useful + thing is the message, and `backup.create` guarantees that the source database + is untouched and no partial file is left behind. + """ + try: + result = backup.create() + except backup.BackupError as exc: + log.error("Backup failed: %s", exc) + raise HTTPException(500, str(exc)) from exc + return {"directory": str(result.path.parent), **result.as_dict()} diff --git a/backend/app/schemas.py b/backend/app/schemas.py index e9f8ea4..dd70cf4 100644 --- a/backend/app/schemas.py +++ b/backend/app/schemas.py @@ -497,6 +497,25 @@ class AdventureOut(ORMModel): can_redo: bool = False +class ImportedAdventureOut(AdventureOut): + """A campaign that has just been restored from a bundle (M9). + + Exactly `AdventureOut` plus what could not be rebuilt. The extra field is on + a subclass rather than on the base, because "which of your search indexes + failed to rebuild" is a fact about one import and not a property of a + campaign — putting it on `AdventureOut` would attach it to every read of + every campaign forever. + + An empty list is the ordinary answer and means the whole campaign, its + evidence and its derived indexes all landed. A non-empty one means the + authoritative import succeeded and a rebuildable index did not, which is a + distinction M9 requires a caller to be able to draw: the campaign is intact, + and Reindex is the repair. + """ + + import_warnings: list[str] = [] + + class ActionPage(BaseModel): """A slice of the story, counted back from the newest action.""" diff --git a/backend/app/starter.py b/backend/app/starter.py index 3b123dd..da47e76 100644 --- a/backend/app/starter.py +++ b/backend/app/starter.py @@ -63,7 +63,9 @@ def give(db: Session, user: models.User) -> models.Adventure | None: # flush whatever part of the adventure the session still held. with db.begin_nested(): story = bundle.plan(payload, bundle.check_format(payload)) - adventure = bundle.materialize(db, payload, story, user.id) + # The starter ships with no imported knowledge, so the derived + # report is always empty here and nothing reads it. + adventure, _ = bundle.materialize(db, payload, story, user.id) _link_scenario(db, adventure, payload) return adventure except Exception: diff --git a/backend/tests/m9_fixture.py b/backend/tests/m9_fixture.py new file mode 100644 index 0000000..209eef0 --- /dev/null +++ b/backend/tests/m9_fixture.py @@ -0,0 +1,406 @@ +"""M9: one campaign that exercises every portable data family at once. + +`TEST-CAMPAIGN-FIXTURE.md` describes the standard Continuity Test campaign, and +the acceptance suites use it. This is a different thing and does not replace it: +the Continuity Test is shaped to read like a story, and this one is shaped to +break a round trip. Every property M9 promises has a source in this campaign that +would be silently lost by a plausible mistake in the exporter or the importer. + + Opening + | + +-- normal turns transcript, state events, snapshots + +-- Retry two takes at one coordinate + +-- knowledge retrieval imported passages in a stored prompt + +-- Save Point S1 a named coordinate on the first line + +-- more turns a future the reader will leave + | + +-- Undo x2 the head steps back + | + +-- divergent continuation a second branch, and a second future + +-- Save Point S2 a named coordinate on the second line + +-- manual state correction an event nothing narrated + +-- Undo x1 the head ends behind the newest row + +The shape is chosen so that no single fact identifies a position. The active head +is not the newest row, not the deepest row, not the last row written, and not on +the branch that holds the most story — an importer that guesses any one of those +lands somewhere else. + +Two campaigns are built, not one. `build` returns the rich campaign; the fixture +also leaves a neighbour beside it, because a bundle that accidentally exported +another campaign's rows would otherwise export nothing and pass. + +The builder speaks HTTP throughout. A fixture that wrote rows directly would +prove the exporter can read what the fixture wrote, which is not the claim. +""" + +from __future__ import annotations + +import asyncio + +from app import memorybank + +from fakes import ScriptedProvider, state_block + +# --------------------------------------------------------------- source files +# Three imported sources, one per class, plus the two lifecycle states that a +# round trip most easily loses: a source someone switched off, and one only the +# narrator may see. + +CANON_MD = """# Westhaven + +## The Old Abbey + +The abbey above Westhaven has stood since the founding. Its crypt is sealed, +and the seal has never been broken. + +## What cannot happen here + +The dead do not return. No rite, relic or bargain in Westhaven has ever +returned anyone from death, and none ever will. +""" + +REFERENCE_MD = """# The Crooked Lantern + +The tavern on Fen Street is timber-framed, low-beamed, and older than the +street it stands on. The hearth is never allowed to go out. + +## The keeper + +Mara keeps the Crooked Lantern. She was born in Westhaven and has never left +it. +""" + +INSPIRATION_MD = """# Weather notes + +Rain on shutters. Lantern light through wet glass. The smell of a hearth +banked for the night. +""" + +SECRET_MD = """# The seal + +The abbey seal was broken once, sixty years ago, and set again by a hand that +is still alive. Nobody in Westhaven knows this. +""" + +DISABLED_MD = """# Discarded draft + +An earlier draft of the Westhaven material, kept for reference and switched off +so it cannot reach the narrator. +""" + +#: The campaign's own rule, so the correction and the canon block have something +#: real to be measured against. +CAMPAIGN_CANON = {"rules": ["The dead do not return."]} + +OPENING = "Aldric sits in the Crooked Lantern with Mara, and the rain starts." + + +# ------------------------------------------------------------------- helpers + +def _play(client, adv_id, text, prose, events=None, kind="do"): + ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"] + response = client.post( + f"/api/adventures/{adv_id}/actions", json={"type": kind, "text": text} + ) + assert response.status_code == 200, response.text[:400] + return response + + +def _fact(predicate, value, fact_id): + return {"type": "add_fact", "predicate": predicate, "value": value, + "fact_id": fact_id} + + +def upload(client, adv_id, name, body, classification, **fields): + """Imports a file the way the browser does: multipart, and no pathname.""" + data = {"classification": classification} + data.update({k: str(v).lower() if isinstance(v, bool) else str(v) + for k, v in fields.items()}) + response = client.post( + f"/api/adventures/{adv_id}/knowledge", + files={"file": (name, body.encode("utf-8"), "text/markdown")}, + data=data, + ) + assert response.status_code == 201, response.text[:400] + return response.json()["id"] + + +def _checkpoint(client, adv_id, name, note=""): + response = client.post( + f"/api/adventures/{adv_id}/checkpoints", json={"name": name, "note": note} + ) + assert response.status_code == 201, response.text[:400] + return response.json() + + +def _undo(client, adv_id, times=1): + for _ in range(times): + response = client.post(f"/api/adventures/{adv_id}/undo") + assert response.status_code == 200, response.text[:400] + + +def settle_derived(adv_id): + """Runs the background memory and summary pass to completion. + + The turn endpoint fires this as a fire-and-forget task, which a test client + does not wait for. Calling it directly is the same code on the same rows — + what is skipped is the scheduling, not the work — and it is what + `test_context_realistic.py` does for the same reason. + """ + asyncio.run(memorybank.run_post_turn(adv_id)) + + +# --------------------------------------------------------------------- build + +def build(client, adv_id) -> dict: + """Plays the fixture campaign onto `adv_id`, and returns what it built. + + The returned dictionary is the assertion source for every round-trip test: + it names the properties that must survive, measured from the campaign as it + stands here rather than restated as constants, so a test compares the copy + against the original instead of against a guess about the original. + """ + # Story memory and the rolling summary on, because a campaign that + # generated neither would let an exporter omit both and still pass. The + # abandoned line below gets long enough to earn its own, which is what E03 + # is about after a round trip. + switched_on = client.patch( + f"/api/adventures/{adv_id}", + json={"auto_summarize": True, "memory_bank_enabled": True}, + ) + assert switched_on.status_code == 200, switched_on.text[:400] + + sources = { + "canon": upload(client, adv_id, "canon.md", CANON_MD, "canon", + always_include=True), + "reference": upload(client, adv_id, "reference.md", REFERENCE_MD, + "reference"), + "inspiration": upload(client, adv_id, "inspiration.md", INSPIRATION_MD, + "inspiration"), + "secret": upload(client, adv_id, "secret.md", SECRET_MD, "canon", + visibility="hidden"), + "disabled": upload(client, adv_id, "draft.md", DISABLED_MD, "reference"), + } + disable = client.patch( + f"/api/adventures/{adv_id}/knowledge/{sources['disabled']}", + json={"enabled": False}, + ) + assert disable.status_code == 200, disable.text[:400] + + # ---- the first line of story ----------------------------------------- + # Turn 1 asks about the abbey, so the canon source is retrieved and the + # stored prompt for this turn holds an imported passage. That turn is the + # one the provenance tests read back after the round trip. + _play(client, adv_id, "ask Mara about the abbey", + "Mara sets down the cloth. The abbey, she says, is sealed.", + [_fact("tally", 10, "tally-10")]) + _play(client, adv_id, "walk up to the abbey", + "The path climbs out of the town and the rain follows.", + [_fact("tally", 20, "tally-20")]) + + # A retry, so one coordinate holds two takes and the earlier one is + # retained but not selected. + ScriptedProvider.replies = [ + "The door is oak, and the seal on it is unbroken.\n" + + state_block([_fact("tally", 30, "tally-30")]) + ] + _play(client, adv_id, "try the crypt door", + "The door will not move.", [_fact("tally", 30, "tally-30")]) + retry = client.post(f"/api/adventures/{adv_id}/retry") + assert retry.status_code == 200, retry.text[:400] + + s1 = _checkpoint(client, adv_id, "At the crypt door", + "Before anything is decided.") + + # The future the reader is about to leave behind. It is played out far + # enough to earn derived data of its own — `memorybank.MEMORY_INTERVAL` is + # six actions — because a summary and a memory belonging to an abandoned + # line are what E03 forbids reaching an active prompt, and a round trip is + # a new way to leak one. + _play(client, adv_id, "force the door", + "The seal gives, and the stair below is dark.", + [_fact("tally", 40, "tally-40")]) + _play(client, adv_id, "go down", + "The crypt is dry, and the air has not moved in years.", + [_fact("tally", 50, "tally-50")]) + _play(client, adv_id, "read the names on the slabs", + "Sixty years of Westhaven dead, and one slab with no name at all.", + [_fact("tally", 60, "tally-60")]) + _play(client, adv_id, "touch the nameless slab", + "The stone is warm, which stone in a crypt is not.", + [_fact("tally", 70, "tally-70")]) + + # Derived data for the line that is about to be abandoned, written while + # the head is still on it. This is the summary and the memory that must + # come back after a round trip and must still be ineligible there. + settle_derived(adv_id) + tip_state = client.get(f"/api/adventures/{adv_id}/state").json() + + # ---- step back, and go somewhere else --------------------------------- + _undo(client, adv_id, 4) + _play(client, adv_id, "turn back and return to the tavern", + "The rain has not let up, and the Lantern's windows are lit.", + [_fact("tally", 41, "tally-41")]) + s2 = _checkpoint(client, adv_id, "Back at the Lantern", "The other way.") + _play(client, adv_id, "ask Mara what she is not saying", + "She looks at the fire for a while before she answers.", + [_fact("tally", 51, "tally-51")]) + _play(client, adv_id, "wait", + "The rain fills the silence, and then she starts talking.", + [_fact("tally", 61, "tally-61")]) + + # A manual correction: an accepted state change with no narration behind + # it, which is the one kind of state event a replay could never recreate. + correction = client.post( + f"/api/adventures/{adv_id}/state/corrections", + json={ + "events": [{ + "type": "add_fact", + "predicate": "keeper_of_the_lantern", + "value": "Mara", + "fact_id": "keeper", + }], + "note": "Established in play before the state system saw it.", + }, + ) + assert correction.status_code == 201, correction.text[:400] + + # Derived data for the line the reader stayed on, so the copy has both an + # eligible and an ineligible summary to tell apart. The generated one landed + # on the abandoned line, which is the E03 case; this one is typed at the + # current head, so it is the eligible case beside it. A round trip has to + # keep them on opposite sides of that line. + settle_derived(adv_id) + + # One more Undo, so the head finishes behind the retained tip of its own + # branch as well as behind the abandoned line's. + _undo(client, adv_id, 1) + + # Typed at the final head, so it is the eligible summary and the generated + # one on the abandoned line is not. A round trip has to keep them on + # opposite sides of that line. + typed = client.patch( + f"/api/adventures/{adv_id}", + json={"story_summary": "Aldric went back to the Lantern instead."}, + ) + assert typed.status_code == 200, typed.text[:400] + + return snapshot_of(client, adv_id, sources=sources, s1=s1, s2=s2, + tip_state=tip_state) + + +def snapshot_in(action: dict) -> dict | None: + """The stored prompt in one bundle entry, decoded. + + The export compresses it (`bundle._packed`), so a test that reached for a + plain dict would conclude the evidence was missing when it is merely + encoded. Both keys are read, plain first, exactly as the importer does. + """ + from app import bundle + + plain = action.get("contextSnapshot") + if isinstance(plain, dict): + return plain + return bundle._unpacked(action.get("contextSnapshotZ")) + + +def with_snapshot(action: dict, snapshot: dict | None) -> dict: + """A bundle entry carrying `snapshot`, written in the plain form. + + Tests that break a snapshot on purpose write the readable key, because the + importer prefers it and because a test that had to compress its own fixture + would be testing the encoding rather than the thing it edited. + """ + edited = {k: v for k, v in action.items() if k != "contextSnapshotZ"} + if snapshot is None: + edited.pop("contextSnapshot", None) + else: + edited["contextSnapshot"] = snapshot + return edited + + +# ------------------------------------------------------------------- reading + +def snapshot_of(client, adv_id, *, sources=None, s1=None, s2=None, + tip_state=None) -> dict: + """Everything about a campaign that a round trip has to reproduce. + + Read through the API, so the comparison is between what a reader can see in + the source campaign and what a reader can see in the copy. Two campaigns + that agree here agree on everything the product promises about a restored + campaign; nothing below is a database id, because ids are expected to + differ. + """ + head = client.get(f"/api/adventures/{adv_id}").json() + branches = client.get(f"/api/adventures/{adv_id}/branches").json() + checkpoints = client.get(f"/api/adventures/{adv_id}/checkpoints").json() + knowledge = client.get(f"/api/adventures/{adv_id}/knowledge").json() + state = client.get(f"/api/adventures/{adv_id}/state").json() + events = client.get(f"/api/adventures/{adv_id}/state/events?limit=500").json() + memories = client.get(f"/api/adventures/{adv_id}/memories").json() + derived = client.get(f"/api/adventures/{adv_id}/derived").json() + return { + "id": adv_id, + "title": head["title"], + "canon_rules": head.get("canon_rules") or [], + "can_undo": head.get("can_undo"), + "can_redo": head.get("can_redo"), + "transcript": [(a["type"], a["text"]) for a in head["actions"]], + # Every branch's own story, which is the whole retained tree as text. + "branch_count": len(branches), + "checkpoints": sorted( + (c["name"], c["note"]) for c in checkpoints + ), + "knowledge": sorted( + (k["title"], k["classification"], k["enabled"], k["visibility"], + k["always_include"], k["content_hash"]) + for k in knowledge + ), + "state": _comparable_state(state), + "state_events": sorted( + (e["event_type"], e["source"], _payload_key(e["payload"])) + for e in events + ), + "memories": sorted(m["text"] for m in memories), + "summaries": sorted( + (s["preview"], s["trigger"], s["eligible"]) + for s in derived.get("summaries", []) + ), + # Carried through from `build`, for the tests that need the original + # ids or the state at a position the head has since left. + "sources": sources, + "s1": s1, + "s2": s2, + "tip_state": _comparable_state(tip_state) if tip_state else None, + } + + +def _comparable_state(state: dict) -> dict: + """The authoritative state, with only what a reader is shown. + + Groups arrive from the API as display sections, which is the right shape to + compare: two campaigns whose State panels read identically hold the same + state, whatever ids sit underneath. + """ + groups = state.get("groups") if isinstance(state, dict) else None + if not isinstance(groups, list): + return {} + return { + str(group.get("title")): sorted( + ", ".join(f"{k}={group_row[k]}" for k in sorted(group_row)) + for group_row in (group.get("rows") or []) + if isinstance(group_row, dict) + ) + for group in groups + } + + +def _payload_key(payload) -> str: + """A stable identity for an event payload, for set comparison.""" + if not isinstance(payload, dict): + return str(payload) + for key in ("fact_id", "entity_id", "thread_id", "id", "predicate"): + if payload.get(key): + return f"{key}={payload[key]}" + return ",".join(f"{k}={payload[k]}" for k in sorted(payload)) diff --git a/backend/tests/test_bundle_v2.py b/backend/tests/test_bundle_v2.py index 7e5db58..b6da10c 100644 --- a/backend/tests/test_bundle_v2.py +++ b/backend/tests/test_bundle_v2.py @@ -568,10 +568,29 @@ def test_the_action_cap_counts_the_rows_a_v1_file_expands_into(client, monkeypat assert _adventure_count() == before, "and nothing was written" -def test_an_unknown_format_is_refused(client): - r = _import(client, {"format": "ai-dnd-adventure-v3", "title": "From the future"}) +def test_a_format_from_a_later_build_is_refused(client): + """A version this build has never heard of is refused, not guessed at. + + The placeholder version here has to stay ahead of `bundle.FORMAT`. It was + `v3` until M9 made v3 real, at which point this test started importing a + bundle it meant to reject — the failure mode a hard-coded "next version" + always eventually has, and the reason the message is asserted against + `bundle.FORMAT` rather than against a literal. + """ + r = _import(client, {"format": "ai-dnd-adventure-v99", "title": "From the future"}) assert r.status_code == 400, r.text - assert bundle.FORMAT in r.json()["detail"] + detail = r.json()["detail"] + assert bundle.FORMAT in detail + assert "ai-dnd-adventure-v99" in detail + + +def test_something_that_is_not_an_export_at_all_is_refused(client): + r = _import(client, {"title": "A file of some other kind"}) + assert r.status_code == 400, r.text + # Every version it can read is named, so the reader can tell whether the + # file they have is one of them. + for readable in bundle.READABLE: + assert readable in r.json()["detail"] # ------------------------------------------------------- the persona (Phase 18) diff --git a/backend/tests/test_imported_knowledge.py b/backend/tests/test_imported_knowledge.py index 28a036f..468b030 100644 --- a/backend/tests/test_imported_knowledge.py +++ b/backend/tests/test_imported_knowledge.py @@ -1660,7 +1660,22 @@ def test_an_edited_content_hash_is_recomputed_and_reported(client): def test_historical_prompt_evidence_survives_an_export_round_trip(client): - """§33: the round trip does not turn provenance into dangling ids.""" + """§33: the round trip does not turn provenance into dangling ids. + + Written in M7 and rewritten in M9, and the rewrite is the point of it. + + In M7 the bundle carried no context snapshots at all, so this test pinned + the *absence*: there were no ids to dangle because there was no evidence, + and the imported campaign's turns simply had no snapshot. That was recorded + at the time as a limit owned by M9 rather than as a property worth keeping — + `V1-ACCEPTANCE-TESTS.md` I05 said so in as many words, and the M8 report + made it handoff question B. + + M9 answered it: the evidence travels. So the assertion inverts, and what it + now pins is the thing M7 was worried about and could not check — that the + provenance arriving on the other side names *this* campaign's sources rather + than the ids it had on the machine that wrote the file. + """ ids = import_fixture(client) play(client, "Aldric asks about the Old Abbey and its broken-circle symbol.") actions = client.get(f"/api/adventures/{client.adv_id}/actions").json()["actions"] @@ -1673,17 +1688,33 @@ def test_historical_prompt_evidence_survives_an_export_round_trip(client): bundle = client.get(f"/api/adventures/{client.adv_id}/export").json() restored = client.post("/api/adventures/import", json=bundle).json() - # The bundle carries no context snapshots at all — it never has, by the rule - # at the top of `bundle.py` — so there are no ids to dangle. The imported - # campaign's turns simply have no snapshot, which is what a pre-M7 bundle - # already did for every other component of the inspector. new_actions = client.get( f"/api/adventures/{restored['id']}/actions" ).json()["actions"] new_ai = next(a for a in reversed(new_actions) if a["type"] == "ai") - assert client.get( + moved = client.get( f"/api/adventures/{restored['id']}/actions/{new_ai['id']}/context" - ).status_code == 404 + ) + assert moved.status_code == 200, moved.text[:300] + moved = moved.json() + + # The evidence itself is identical: the same passages, the same text, the + # same prompt the turn was actually assembled from. + assert [(r["title"], r["text"]) for r in moved["knowledge"]["used"]] == \ + [(r["title"], r["text"]) for r in before["knowledge"]["used"]] + assert moved["prompt"] == before["prompt"] + + # And the one pointer that is not evidence has been translated, so the + # inspector's "open this source" reaches the restored library rather than + # whatever holds that id here. + theirs = { + source["id"] for source in + client.get(f"/api/adventures/{restored['id']}/knowledge").json() + } + named = {r["source_id"] for r in moved["knowledge"]["used"] + if r["source_id"] is not None} + assert named and named <= theirs + # And the original campaign's evidence is untouched by having been exported. after = client.get( f"/api/adventures/{client.adv_id}/actions/{ai_action['id']}/context" diff --git a/backend/tests/test_m9_backup.py b/backend/tests/test_m9_backup.py new file mode 100644 index 0000000..0989121 --- /dev/null +++ b/backend/tests/test_m9_backup.py @@ -0,0 +1,451 @@ +"""M9: a consistent copy of the whole database, taken while it is being written. + +`app/backup.py` explains why a plain file copy is not a backup. This file is the +evidence for the claim, and the shape of it matters: **every test below opens the +backup as its own database and reads what is in it.** A test that only checked a +file appeared, or that the endpoint returned 201, would pass against a `cp` — and +a `cp` is exactly what this replaces. + +The load test is the one that separates the two. It writes to the source +database *while* the backup is being taken, from a second thread, and then asks +the copy for a story it can check turn by turn. A page-torn copy would show a +transcript with a hole in it, a campaign whose head points past its own story, or +a `quick_check` failure — and would show none of those on a quiet database, which +is why the quiet case is not the interesting one. + + python -m pytest tests/test_m9_backup.py -v +""" + +import os +import sqlite3 +import tempfile +import threading +import time +from pathlib import Path + +import pytest +from fastapi import Depends +from fastapi.testclient import TestClient + +from app import auth, backup, limits, models +from app.database import Base, SessionLocal, engine, get_db +from app.main import app +from app.routers import adventures + +from fakes import ScriptedProvider, tally_of, tally_reply + + +@pytest.fixture() +def client(monkeypatch): + """The app, and a campaign with enough in it to recognise afterwards.""" + Base.metadata.create_all(bind=engine) + setup = SessionLocal() + user = models.User(is_guest=False, email="backup@example.com") + setup.add(user) + setup.flush() + setup.add(models.Settings(user_id=user.id, model="test-model")) + adventure = models.Adventure(user_id=user.id, title="Backed up") + setup.add(adventure) + setup.flush() + setup.add(models.Action( + adventure_id=adventure.id, type="start", text="The story opens.", + )) + setup.commit() + adv_id, user_id = adventure.id, user.id + setup.close() + + monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None) + monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider) + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id) + ) + test_client = TestClient(app) + test_client.adv_id = adv_id + try: + yield test_client + finally: + app.dependency_overrides.clear() + adventures.turns._active_turns.clear() + Base.metadata.drop_all(bind=engine) + + +@pytest.fixture() +def elsewhere(tmp_path, monkeypatch): + """Backups land under a temporary directory, not beside the real database.""" + fake_db = tmp_path / "campaign.db" + fake_db.write_bytes(Path(str(engine.url.database)).read_bytes()) + return fake_db + + +def _play(client, text, total): + ScriptedProvider.replies = [tally_reply(f"Beat {total // 10}.", total)] + response = client.post(f"/api/adventures/{client.adv_id}/actions", + json={"type": "do", "text": text}) + assert response.status_code == 200, response.text[:300] + + +def _open(path) -> sqlite3.Connection: + """The backup, as its own database, read-only.""" + connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True) + connection.row_factory = sqlite3.Row + return connection + + +# ------------------------------------------------------------------ the copy + +def test_the_backup_is_a_database_that_passes_its_own_integrity_check(client): + for turn in range(1, 4): + _play(client, f"turn {turn}", turn * 10) + result = backup.create() + try: + assert result.integrity == "ok" + assert result.pages > 0 + assert result.bytes > 0 + with _open(result.path) as db: + assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok" + assert db.execute("PRAGMA foreign_key_check").fetchall() == [] + finally: + result.path.unlink(missing_ok=True) + + +def test_the_backup_holds_the_schema_and_every_family_of_row(client): + """Not "the file exists": the copy is opened and asked what is in it.""" + for turn in range(1, 4): + _play(client, f"turn {turn}", turn * 10) + checkpoint = client.post(f"/api/adventures/{client.adv_id}/checkpoints", + json={"name": "Here", "note": "A position."}) + assert checkpoint.status_code == 201 + upload = client.post( + f"/api/adventures/{client.adv_id}/knowledge", + files={"file": ("canon.md", b"# Rule\n\nThe dead do not return.\n", + "text/markdown")}, + data={"classification": "canon"}, + ) + assert upload.status_code == 201, upload.text[:300] + + result = backup.create() + try: + with _open(result.path) as db: + tables = { + row["name"] for row in + db.execute("SELECT name FROM sqlite_master WHERE type='table'") + } + for expected in ("adventures", "actions", "branches", "checkpoints", + "knowledge_sources", "knowledge_chunks", + "state_events", "summaries", "settings"): + assert expected in tables, f"{expected} is missing from the backup" + + campaign = db.execute( + "SELECT * FROM adventures WHERE id = ?", (client.adv_id,) + ).fetchone() + assert campaign["title"] == "Backed up" + # The head, which is the thing a restore has to reproduce. + assert campaign["head_depth"] >= 0 + assert campaign["head_branch_id"] is not None + + texts = [row["text"] for row in db.execute( + "SELECT text FROM actions WHERE adventure_id = ? ORDER BY id", + (client.adv_id,), + )] + assert "The story opens." in texts + assert any("Beat 3." in text for text in texts) + + assert db.execute( + "SELECT name FROM checkpoints WHERE adventure_id = ?", + (client.adv_id,), + ).fetchone()["name"] == "Here" + assert db.execute( + "SELECT COUNT(*) c FROM knowledge_sources WHERE adventure_id = ?", + (client.adv_id,), + ).fetchone()["c"] == 1 + assert db.execute( + "SELECT COUNT(*) c FROM state_events WHERE adventure_id = ?", + (client.adv_id,), + ).fetchone()["c"] > 0 + # And the head names a turn that is actually in the copy. + assert db.execute( + "SELECT COUNT(*) c FROM actions WHERE adventure_id = ? " + "AND branch_id = ? AND depth = ?", + (client.adv_id, campaign["head_branch_id"], campaign["head_depth"]), + ).fetchone()["c"] > 0 + finally: + result.path.unlink(missing_ok=True) + + +def test_the_state_in_the_backup_is_the_state_the_campaign_had(client): + """The authoritative document, read out of the copy and compared.""" + for turn in range(1, 5): + _play(client, f"turn {turn}", turn * 10) + live = client.get(f"/api/adventures/{client.adv_id}/state").json()["document"] + result = backup.create() + try: + with _open(result.path) as db: + from app import compression + + blob = db.execute( + "SELECT narrative_state FROM adventures WHERE id = ?", + (client.adv_id,), + ).fetchone()["narrative_state"] + assert tally_of(compression.unpack(blob)) == tally_of(live) == 40 + finally: + result.path.unlink(missing_ok=True) + + +# ------------------------------------------------------------ while it is live + +def test_a_backup_taken_during_writes_is_consistent(client): + """The claim a plain file copy cannot make. + + Turns are played from a second thread throughout the copy. The backup that + comes out is a snapshot of *some* committed point — which point is not + determined, and asserting on a particular one would be asserting on a race — + so what is checked is that it is a coherent one: `quick_check` passes, no + foreign key dangles, the transcript has no gap in it, and the head names a + turn that exists. + """ + stop = threading.Event() + written: list[int] = [] + failures: list[Exception] = [] + + def keep_writing(): + turn = 0 + while not stop.is_set() and turn < 40: + turn += 1 + try: + _play(client, f"concurrent {turn}", turn * 10) + written.append(turn) + except Exception as exc: # noqa: BLE001 - reported to the test + failures.append(exc) + return + time.sleep(0.005) + + writer = threading.Thread(target=keep_writing, daemon=True) + writer.start() + # Let a few turns land, so the copy is taken over a database that is moving + # rather than one that has not started. + while len(written) < 3 and writer.is_alive(): + time.sleep(0.01) + + result = backup.create() + stop.set() + writer.join(timeout=30) + assert not failures, f"the writer failed: {failures[0]}" + assert written, "no turn was written during the backup" + + try: + with _open(result.path) as db: + assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok" + assert db.execute("PRAGMA foreign_key_check").fetchall() == [] + + rows = db.execute( + "SELECT depth, type FROM actions WHERE adventure_id = ? " + "AND live = 1 ORDER BY depth", + (client.adv_id,), + ).fetchall() + depths = [row["depth"] for row in rows] + assert depths == list(range(len(depths))), ( + f"the transcript in the backup has a gap: {depths}" + ) + campaign = db.execute( + "SELECT head_branch_id, head_depth FROM adventures WHERE id = ?", + (client.adv_id,), + ).fetchone() + assert db.execute( + "SELECT COUNT(*) c FROM actions WHERE adventure_id = ? " + "AND branch_id = ? AND depth = ?", + (client.adv_id, campaign["head_branch_id"], campaign["head_depth"]), + ).fetchone()["c"] > 0, "the head points past the story in the backup" + finally: + result.path.unlink(missing_ok=True) + + +def test_the_source_database_is_untouched_by_a_backup(client): + """Opened read-only, so this is a guarantee rather than an observation.""" + _play(client, "one", 10) + source = Path(str(engine.url.database)) + before = source.read_bytes() + result = backup.create() + try: + assert source.read_bytes() == before + assert client.get(f"/api/adventures/{client.adv_id}").status_code == 200 + finally: + result.path.unlink(missing_ok=True) + + +# ---------------------------------------------------------------- the rules + +def test_an_existing_backup_is_never_overwritten(client): + """Yesterday's backup surviving today's mistake is most of the point.""" + first = backup.create() + second = backup.create() + try: + assert first.path != second.path + assert first.path.exists() and second.path.exists() + finally: + first.path.unlink(missing_ok=True) + second.path.unlink(missing_ok=True) + + +def test_two_backups_in_the_same_second_do_not_collide(client, monkeypatch): + from datetime import datetime + + fixed = datetime(2026, 9, 7, 4, 30, 0) + first = backup.create(now=fixed) + second = backup.create(now=fixed) + try: + assert first.path != second.path + assert first.path.exists() and second.path.exists() + finally: + first.path.unlink(missing_ok=True) + second.path.unlink(missing_ok=True) + + +def test_a_failed_verification_leaves_nothing_behind(client, monkeypatch): + """A backup nobody verified is a belief, and one that fails is not kept.""" + monkeypatch.setattr( + backup, "_verify", + lambda path: (_ for _ in ()).throw(backup.BackupError("bad pages")), + ) + root = backup.directory() + before = set(root.iterdir()) + with pytest.raises(backup.BackupError, match="bad pages"): + backup.create() + assert set(root.iterdir()) == before, "a failed backup left a file behind" + + +def test_a_failed_copy_leaves_nothing_behind_and_reports_the_reason( + client, monkeypatch +): + monkeypatch.setattr( + backup, "_copy", + lambda source, working: (_ for _ in ()).throw(OSError("disk full")), + ) + root = backup.directory() + before = set(root.iterdir()) + with pytest.raises(backup.BackupError, match="disk full"): + backup.create() + assert set(root.iterdir()) == before + + +def test_a_missing_source_database_is_reported_rather_than_guessed_at(tmp_path): + with pytest.raises(backup.BackupError, match="no database"): + backup.create(tmp_path / "not-here.db") + + +def test_the_partial_file_is_never_left_wearing_a_backups_name(client, monkeypatch): + """The rename is the last step, so an interrupted run is invisible.""" + seen: list[Path] = [] + real_copy = backup._copy + + def watch(source, working): + seen.append(Path(working)) + return real_copy(source, working) + + monkeypatch.setattr(backup, "_copy", watch) + result = backup.create() + try: + assert seen and seen[0].name.endswith(".partial") + assert not seen[0].exists(), "the temporary file survived" + assert result.path.exists() + assert not result.path.name.endswith(".partial") + finally: + result.path.unlink(missing_ok=True) + + +# --------------------------------------------------------------- the endpoint + +def test_the_endpoint_takes_a_backup_and_says_where_it_went(client): + response = client.post("/api/backups") + assert response.status_code == 201, response.text[:300] + body = response.json() + path = Path(body["directory"]) / body["filename"] + try: + assert body["integrity"] == "ok" + assert body["bytes"] > 0 + assert path.exists() + with _open(path) as db: + assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok" + finally: + path.unlink(missing_ok=True) + + +def test_the_endpoint_lists_what_is_there_newest_first(client): + """Ordered by when the backup was taken, which is what its name records. + + Both files here are written in the same instant, so their modification times + are indistinguishable and only the stamp in the name says which is which. + That is not a contrived case: copying a backup to another disk or restoring + one from an archive rewrites its mtime, and a list that reordered itself + afterwards would report when the file was last handled rather than when the + backup was taken. + """ + from datetime import datetime + + older = backup.create(now=datetime(2026, 9, 1, 10, 0, 0)) + newer = backup.create(now=datetime(2026, 9, 6, 10, 0, 0)) + try: + listed = client.get("/api/backups") + assert listed.status_code == 200 + rows = listed.json()["backups"] + names = [row["filename"] for row in rows] + assert names.index(newer.path.name) < names.index(older.path.name) + by_name = {row["filename"]: row["taken_at"] for row in rows} + assert by_name[newer.path.name].startswith("2026-09-06T10:00") + assert by_name[older.path.name].startswith("2026-09-01T10:00") + finally: + older.path.unlink(missing_ok=True) + newer.path.unlink(missing_ok=True) + + +def test_a_backup_this_build_did_not_name_still_lists(client): + """A file in the directory whose name carries no stamp is still shown. + + The modification time answers instead. The fallback exists to keep a + hand-renamed or third-party file visible rather than silently absent from + the list a reader uses to find their backups. + """ + stray = backup.directory() / f"{backup.PREFIX}-handwritten.db" + stray.write_bytes(b"SQLite format 3\x00") + try: + rows = client.get("/api/backups").json()["backups"] + listed = {row["filename"]: row for row in rows} + assert stray.name in listed + assert listed[stray.name]["taken_at"] + finally: + stray.unlink(missing_ok=True) + + +def test_the_endpoint_accepts_no_path_from_the_caller(client): + """H08. There is no field to attempt a traversal in. + + The destination is derived from the database the application already has + open and the name from the clock, so a body is not merely ignored — there is + nothing for one to name. + """ + from app.main import app as application + + schema = application.openapi()["paths"]["/api/backups"]["post"] + assert "requestBody" not in schema + assert not schema.get("parameters") + # And sending one anyway changes nothing about where the file lands. + response = client.post("/api/backups", json={"path": "../../../tmp/escape.db"}) + assert response.status_code == 201, response.text[:300] + body = response.json() + path = Path(body["directory"]) / body["filename"] + try: + assert path.parent == backup.directory() + assert ".." not in body["filename"] + finally: + path.unlink(missing_ok=True) + + +def test_a_failure_is_a_clear_error_rather_than_a_silent_success( + client, monkeypatch +): + monkeypatch.setattr( + backup, "create", + lambda *a, **k: (_ for _ in ()).throw(backup.BackupError("no space left")), + ) + response = client.post("/api/backups") + assert response.status_code == 500 + assert "no space left" in response.json()["detail"] diff --git a/backend/tests/test_m9_clean_import.py b/backend/tests/test_m9_clean_import.py new file mode 100644 index 0000000..3fcfe46 --- /dev/null +++ b/backend/tests/test_m9_clean_import.py @@ -0,0 +1,460 @@ +"""M9: the campaign moves to a machine that has never seen it. + +This is the milestone's Definition of Done, and it is the one claim the rest of +the M9 suite cannot make. `test_m9_portability.py` imports beside the original, +in one process, against one database — which is the right place to check the +*contract* and the wrong place to check *portability*. A shared id space, a +warm cache, a row the exporter forgot to scope, a session still holding the +original: every one of those would pass there and fail here. + +So each test below: + +1. starts a real server process against database A, and plays a campaign; +2. exports it over HTTP and stops that process; +3. starts a **second** server process against database B, **a file that has + never existed before**, in a different directory; +4. imports the file over HTTP, and asks the second process what it has. + +Nothing crosses between them but the bundle. Migrations run on B from nothing, +because it is a new file — so this is also the fresh-install path, and the +"clean data directory" in the Definition of Done is a directory, not a metaphor. + +The final test restarts the *importing* server, which is L03 after a move: a +Save Point restored in the third process must reach the same position and the +same state as it did in the second. + + python -m pytest tests/test_m9_clean_import.py -v +""" +import json +import os +import shutil +import sqlite3 +import subprocess +import sys +import tempfile +import urllib.error +import urllib.request +from pathlib import Path + +import pytest + +from fakes import TALLY_PER_TURN, tally_of +from test_process_restart import Server, _free_port + +HERE = Path(__file__).resolve().parent + + +@pytest.fixture() +def machines(): + """Two directories, each with its own database, and the servers on them. + + Two directories rather than two filenames, because the backup directory and + anything else the application derives from the database's location must land + in the importing machine's own space rather than beside the exporter's. + """ + root = tempfile.mkdtemp(prefix="m9-clean-") + started: list[Server] = [] + + def start(name: str) -> Server: + directory = os.path.join(root, name) + os.makedirs(directory, exist_ok=True) + server = Server(os.path.join(directory, "campaign.db"), _free_port()) + started.append(server) + server.wait_until_ready() + return server + + def path_of(name: str) -> str: + return os.path.join(root, name, "campaign.db") + + try: + yield start, path_of + finally: + for server in started: + server.stop() + shutil.rmtree(root, ignore_errors=True) + + +# ------------------------------------------------------------------ building + +def _campaign(server: Server) -> int: + """A campaign with everything a move has to carry, played over HTTP. + + Deliberately not `m9_fixture`: that builds through a `TestClient` and this + file exists to avoid one. What it reproduces is the same shape — a retry, a + Save Point, an imported source that a turn actually used, an undone head and + a retained future. + """ + adventure = server.call("POST", "/adventures", { + "title": "Moved between machines", + "canon_rules": ["The dead do not return."], + "opening": "Aldric sits in the Crooked Lantern with Mara.", + }, expect=201) + adv_id = adventure["id"] + + _upload(server, adv_id, "canon.md", "canon", ( + "# Westhaven\n\n## The Old Abbey\n\nThe abbey above Westhaven has stood " + "since the founding. Its crypt is sealed, its door is oak, and the seal " + "on it has never been broken.\n" + )) + _upload(server, adv_id, "secret.md", "canon", ( + "# The seal\n\nIt was broken once, sixty years ago.\n" + ), visibility="hidden") + disabled = _upload(server, adv_id, "draft.md", "reference", ( + "# Discarded draft\n\nAn earlier version, switched off.\n" + )) + server.call("PATCH", f"/adventures/{adv_id}/knowledge/{disabled}", + {"enabled": False}, expect=200) + + # The spawned narrator writes "Beat N." and nothing else, so every term the + # retrieval has to work with comes from the player's own words. They are + # written to name things the Canon file names. + server.play(adv_id, "ask Mara about the abbey crypt in Westhaven") + server.play(adv_id, "walk up the hill to the abbey") + server.play(adv_id, "try the sealed crypt door of the abbey") + _retry(server, adv_id) + server.call("POST", f"/adventures/{adv_id}/checkpoints", + {"name": "At the door", "note": "Before deciding."}, expect=201) + server.play(adv_id, "force the door") + server.play(adv_id, "go down the stair") + server.call("POST", f"/adventures/{adv_id}/state/corrections", { + "events": [{"type": "add_fact", "predicate": "keeper", "value": "Mara", + "fact_id": "keeper"}], + "note": "Established in play before the state system saw it.", + }, expect=201) + # Two Undos, so the export is taken behind the retained tip. + server.call("POST", f"/adventures/{adv_id}/undo", expect=200) + server.call("POST", f"/adventures/{adv_id}/undo", expect=200) + return adv_id + + +def _retry(server: Server, adv_id: int) -> None: + """Retries the newest turn, over the streaming endpoint it actually uses. + + `Server.call` parses JSON, and `/retry` answers with an SSE stream as + `/actions` does — so calling it as JSON reads `data: {...}` as a document and + fails on the first character. Draining the stream is what the browser does. + """ + request = urllib.request.Request( + f"http://127.0.0.1:{server.port}/api/adventures/{adv_id}/retry", + data=b"{}", method="POST", + headers={"Content-Type": "application/json"}, + ) + with urllib.request.urlopen(request, timeout=120) as response: + body = response.read() + assert b'"type": "error"' not in body, body[:300] + + +def _upload(server: Server, adv_id: int, name: str, classification: str, + body: str, **fields) -> int: + """A multipart knowledge upload over real HTTP, without a client library.""" + boundary = "----m9cleanimport" + parts = [] + for key, value in {"classification": classification, **fields}.items(): + parts.append( + f"--{boundary}\r\nContent-Disposition: form-data; name=\"{key}\"\r\n" + f"\r\n{value}\r\n" + ) + parts.append( + f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; " + f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n" + ) + payload = ("".join(parts) + f"--{boundary}--\r\n").encode() + request = urllib.request.Request( + f"http://127.0.0.1:{server.port}/api/adventures/{adv_id}/knowledge", + data=payload, method="POST", + headers={"Content-Type": f"multipart/form-data; boundary={boundary}"}, + ) + with urllib.request.urlopen(request, timeout=60) as response: + return json.loads(response.read())["id"] + + +def _snapshot(server: Server, adv_id: int) -> dict: + """What a reader can see, read over HTTP through the API they read.""" + page = server.call("GET", f"/adventures/{adv_id}", expect=200) + return { + "title": page["title"], + "canon_rules": page["canon_rules"], + "transcript": [(a["type"], a["text"]) for a in page["actions"]], + "can_undo": page["can_undo"], + "can_redo": page["can_redo"], + "state": server.call("GET", f"/adventures/{adv_id}/state", expect=200)["document"], + "checkpoints": sorted( + (c["name"], c["note"], c["depth"]) + for c in server.call("GET", f"/adventures/{adv_id}/checkpoints", expect=200) + ), + "knowledge": sorted( + (k["title"], k["classification"], k["enabled"], k["visibility"], + k["content_hash"], k["index_state"], k["chunk_count"] > 0) + for k in server.call("GET", f"/adventures/{adv_id}/knowledge", expect=200) + ), + "events": sorted( + (e["event_type"], e["source"], json.dumps(e["payload"], sort_keys=True)) + for e in server.call("GET", f"/adventures/{adv_id}/state/events?limit=500", + expect=200) + ), + "rows": server.total_rows(adv_id), + } + + +# ------------------------------------------------------------------- the move + +@pytest.fixture() +def moved(machines): + """The campaign, exported from machine A and imported into a clean B.""" + start, path_of = machines + source = start("a") + adv_id = _campaign(source) + before = _snapshot(source, adv_id) + # What the source machine retrieves at this position, recorded while it is + # still running. It is the only thing the copy can honestly be compared to. + retrieved = { + record["filename"] for record in + source.call("GET", f"/adventures/{adv_id}/context", expect=200) + ["knowledge"]["used"] + } + bundle = source.call("GET", f"/adventures/{adv_id}/export", expect=200) + source.stop() + + assert not os.path.exists(path_of("b")), "machine B must not exist yet" + target = start("b") + assert target.call("GET", "/adventures", expect=200) == [], \ + "machine B is not empty" + + imported = target.call("POST", "/adventures/import", bundle, expect=201) + return { + "bundle": bundle, "before": before, "target": target, + "retrieved": retrieved, + "copy_id": imported["id"], "imported": imported, + "path": path_of, "start": start, + } + + +def test_the_campaign_arrives_whole_on_a_machine_that_never_had_it(moved): + """The Definition of Done, in one assertion per family.""" + after = _snapshot(moved["target"], moved["copy_id"]) + before = moved["before"] + assert after["transcript"] == before["transcript"] + assert after["state"] == before["state"] + assert after["canon_rules"] == before["canon_rules"] + assert after["checkpoints"] == before["checkpoints"] + assert after["knowledge"] == before["knowledge"] + assert after["events"] == before["events"] + assert after["rows"] == before["rows"], "the retained tree is a different size" + + +def test_it_opens_at_the_exact_head_it_was_exported_at(moved): + """I07, across the boundary the acceptance test names. + + The export was taken two Undos behind the tip, so a machine that opened the + campaign at its newest retained turn would show a story two turns longer + than the one that was saved. + """ + after = _snapshot(moved["target"], moved["copy_id"]) + assert after["transcript"] == moved["before"]["transcript"] + assert after["can_redo"] is True, "the retained future is not reachable" + assert moved["imported"]["can_redo"] is True, ( + "the response that opens the campaign says Redo is unavailable" + ) + assert after["rows"] > len(after["transcript"]), ( + "the retained future is not in the database" + ) + + +def test_the_state_audit_arrives_and_still_names_its_author(moved): + """The manual correction is still a manual correction on the new machine.""" + events = moved["target"].call( + f"GET", f"/adventures/{moved['copy_id']}/state/events?limit=500", expect=200 + ) + manual = [e for e in events if e["source"] == "manual_correction"] + assert len(manual) == 1 + assert manual[0]["payload"]["predicate"] == "keeper" + assert any(e["source"] == "accepted_story" for e in events), ( + "and the story's own events are there beside it" + ) + + +def test_the_knowledge_works_with_no_access_to_the_original_machine(moved): + """§11. The exporting machine is stopped; nothing may reach back to it. + + Its process is dead and its directory holds a database this server has never + opened. If retrieval works here, it works from the content the file carried. + + The comparison is against what the *source* retrieved, recorded before that + process was killed, and the source's own result is asserted first. A test + that only checked the copy retrieved something would pass by accident on a + day the fixture happened to match, and — worse — would report a portability + failure when what had actually happened is that neither side retrieved + anything. That is M8's finding 10: assert your own precondition. + """ + assert moved["retrieved"], ( + "the source campaign retrieved nothing, so this proves nothing about " + "the copy" + ) + report = moved["target"].call( + "GET", f"/adventures/{moved['copy_id']}/context", expect=200 + ) + used = {record["filename"] for record in report["knowledge"]["used"]} + assert used == moved["retrieved"], ( + f"the copy retrieved {used} where the source retrieved {moved['retrieved']}" + ) + assert "draft.md" not in used, "the disabled source was re-enabled by the move" + assert "canon.md" in used + + +def test_a_historical_turn_still_shows_what_it_was_given(moved): + """The M8 handoff, across the boundary that made it a handoff. + + Inspect Context on an old narrator turn works on a machine that never + assembled that prompt and could not reassemble it — the sources are here but + the state, the head and the canon have all moved on since. + """ + target, copy_id = moved["target"], moved["copy_id"] + page = target.call("GET", f"/adventures/{copy_id}/actions?limit=200", expect=200) + narrator = [a for a in page["actions"] if a["type"] == "ai"] + assert narrator, "the imported campaign has no narrator turn" + inspected = 0 + for action in narrator: + response = target.call( + "GET", f"/adventures/{copy_id}/actions/{action['id']}/context" + ) + if response is None: + continue + assert response["prompt"]["system"], "a restored prompt is empty" + assert response["sections"], "a restored prompt has no sections" + inspected += 1 + assert inspected, "no turn on the new machine can say what it was told" + + +def test_no_secret_and_no_path_from_the_old_machine_travelled(moved): + """I06, and the private-detail half of it. + + The bundle is checked as text, because that is what actually left the + machine — a field added to a model the exporter walks would reach the file + without any test of a column noticing. + """ + text = json.dumps(moved["bundle"]) + assert "api_key" not in text + assert "11434" not in text, "an inference endpoint travelled with the campaign" + assert "/tmp/" not in text and "campaign.db" not in text, ( + "a filesystem path from the exporting machine travelled" + ) + + +def test_the_importing_machine_keeps_its_own_settings(moved): + """§15. A campaign is not a way to reconfigure the destination. + + The bundle carries per-turn model provenance, which is a record of what + happened. It does not carry the endpoint, the model or the context budget, + because those describe the machine rather than the campaign — and importing + a campaign must not silently repoint the destination's inference at the + source's. + """ + settings = moved["target"].call("GET", "/settings", expect=200) + assert settings["endpoint_url"] == "http://localhost:11434/v1", ( + "the import changed the destination's inference endpoint" + ) + assert settings["context_token_budget"] == 16384 + + +def test_a_missing_model_does_not_stop_the_campaign_arriving(moved): + """§15. The campaign and its data are portable independently of a model. + + The importing server has no model configured at all — nothing has ever + written a `model` into its settings — and the import still succeeds, opens, + and shows its state. Play would fail; recovery does not. + """ + settings = moved["target"].call("GET", "/settings", expect=200) + assert settings["model"] == "", "this test needs an unconfigured destination" + after = _snapshot(moved["target"], moved["copy_id"]) + assert after["transcript"] == moved["before"]["transcript"] + + +# ------------------------------------------------- L03, after the campaign moved + +def test_l03_a_save_point_restored_on_the_new_machine_survives_its_restart(moved): + """L03, with the move in front of it. + + Restore a Save Point in the second process, record the position and the + state, kill the process, start a **third** against the same file, and ask + again. What crosses is bytes on disk. + """ + target, copy_id = moved["target"], moved["copy_id"] + points = target.call("GET", f"/adventures/{copy_id}/checkpoints", expect=200) + assert points, "the Save Point did not survive the move" + point = points[0] + assert point["resolved"] is True + + target.call("POST", f"/adventures/{copy_id}/checkpoints/{point['id']}/restore", + expect=200) + restored = _snapshot(target, copy_id) + rows_before = restored["rows"] + target.stop() + assert not target.is_listening() + + third = moved["start"]("b") + again = _snapshot(third, copy_id) + assert again["transcript"] == restored["transcript"] + assert again["state"] == restored["state"] + assert again["rows"] == rows_before, "restoring deleted later history" + + +# ------------------------------------------------------ the database it wrote + +def test_the_importing_machines_database_passes_its_own_integrity_check(moved): + """A campaign written by an import is a database SQLite is happy with.""" + moved["target"].stop() + connection = sqlite3.connect(moved["path"]("b")) + try: + assert connection.execute("PRAGMA quick_check").fetchone()[0] == "ok" + assert connection.execute("PRAGMA foreign_key_check").fetchall() == [] + finally: + connection.close() + + +def test_the_import_left_no_orphan_behind(moved): + """§17's list, checked against the database rather than against the API. + + Every one of these would be invisible from the outside until the moment it + mattered: a Save Point pointing at a turn that is not there, knowledge owned + by a campaign that does not exist, an action on a branch belonging to + something else. + """ + moved["target"].stop() + connection = sqlite3.connect(moved["path"]("b")) + try: + def one(sql): + return connection.execute(sql).fetchone()[0] + + assert one(""" + SELECT COUNT(*) FROM checkpoints c + LEFT JOIN actions a + ON a.branch_id = c.branch_id AND a.depth = c.depth + AND a.adventure_id = c.adventure_id + WHERE a.id IS NULL + """) == 0, "a Save Point names a position with no turn at it" + assert one(""" + SELECT COUNT(*) FROM actions a + LEFT JOIN branches b ON b.id = a.branch_id + WHERE a.branch_id IS NOT NULL + AND (b.id IS NULL OR b.adventure_id <> a.adventure_id) + """) == 0, "an action sits on another campaign's branch" + assert one(""" + SELECT COUNT(*) FROM knowledge_sources k + LEFT JOIN adventures adv ON adv.id = k.adventure_id + WHERE adv.id IS NULL + """) == 0, "knowledge owned by no campaign" + assert one(""" + SELECT COUNT(*) FROM state_events e + LEFT JOIN actions a ON a.id = e.action_id + WHERE e.action_id IS NOT NULL + AND (a.id IS NULL OR a.adventure_id <> e.adventure_id) + """) == 0, "a state event names a turn in another campaign" + assert one(""" + SELECT COUNT(*) FROM adventures adv + LEFT JOIN actions a + ON a.branch_id = adv.head_branch_id AND a.depth = adv.head_depth + AND a.adventure_id = adv.id + WHERE adv.head_depth >= 0 AND a.id IS NULL + """) == 0, "the head points outside the retained story" + finally: + connection.close() diff --git a/backend/tests/test_m9_corrupt_bundles.py b/backend/tests/test_m9_corrupt_bundles.py new file mode 100644 index 0000000..2d2db29 --- /dev/null +++ b/backend/tests/test_m9_corrupt_bundles.py @@ -0,0 +1,747 @@ +"""M9: what a broken bundle does, and what it must never do. + +A campaign bundle is a file on a disk. It can be truncated by a full volume, +mangled by a text editor, hand-written by somebody curious, or produced by a +build that does not exist yet. Every case below starts from a real export of the +M9 fixture and breaks exactly one thing about it, so what each test measures is +that one break rather than a fixture nobody would recognise. + +## The two rules + +**Nothing lands.** A refused import leaves no campaign, no branch, no orphan +action, no Save Point pointing at nothing, and no knowledge owned by a campaign +that does not exist. `bundle.plan` has no side effects and runs before a row is +written, and the endpoint commits once, so a refusal is a refusal — checked here +by counting rows before and after rather than by trusting the status code. + +**Nothing is fetched, read or run.** A bundle is data. A URL in it is text, a +filename in it is text, and a path in it is text. No test here needs a network +guard to pass, which is the point: there is no code path that would use one. + +## Refuse or repair, and why each is which + +The two are not interchangeable and the choice is made per field, on one +question — *does a wrong value here make the rest of the campaign wrong?* + + refuse the head, the tree, the audit trail + a head past the story misplaces every read of it; a node on a + branch that is not listed is a story with a hole; an audit record + naming a turn that is not there leaves state nobody can explain + repair a knowledge classification that is unreadable, a filename with a + path in it, a live flag nobody set + the value is not load-bearing for anything but itself + drop a Save Point that names no turn, a summary with no coordinate + a bookmark costs a bookmark; refusing the campaign to save it + would lose the story + +What none of them ever is: **retarget**. A Save Point whose position is not in +the file does not get moved to a nearby one, because the reader named a position +and no other position is the one they named. + + python -m pytest tests/test_m9_corrupt_bundles.py -v +""" + +import copy +import json + +import pytest +from fastapi import Depends +from fastapi.testclient import TestClient + +from app import auth, limits, memorybank, models +from app.database import Base, SessionLocal, engine, get_db +from app.knowledge import embeddings +from app.main import app +from app.routers import adventures + +import m9_fixture +from fakes import ScriptedProvider +from test_m9_portability import StubDerived + + +@pytest.fixture() +def client(monkeypatch): + Base.metadata.create_all(bind=engine) + memorybank._vector_cache.clear() + embeddings._cache.clear() + setup = SessionLocal() + user = models.User(is_guest=False, email="corrupt@example.com") + setup.add(user) + setup.flush() + setup.add(models.Settings( + user_id=user.id, model="test-model", embedding_model="stub-embed", + context_token_budget=4000, max_output_tokens=400, + )) + adventure = models.Adventure( + user_id=user.id, title="Source campaign", + campaign_canon=m9_fixture.CAMPAIGN_CANON, + ) + setup.add(adventure) + setup.flush() + setup.add(models.Action( + adventure_id=adventure.id, type="start", text=m9_fixture.OPENING, + )) + setup.commit() + adv_id, user_id = adventure.id, user.id + setup.close() + + monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None) + monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider) + monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived()) + monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived()) + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id) + ) + test_client = TestClient(app) + test_client.adv_id = adv_id + try: + yield test_client + finally: + app.dependency_overrides.clear() + adventures.turns._active_turns.clear() + memorybank._vector_cache.clear() + embeddings._cache.clear() + Base.metadata.drop_all(bind=engine) + + +@pytest.fixture(scope="module") +def _cache(): + """One place to keep the exported fixture between tests in this module.""" + return {} + + +@pytest.fixture() +def good(client): + """A real, valid export of the M9 fixture, ready to be broken.""" + m9_fixture.build(client, client.adv_id) + response = client.get(f"/api/adventures/{client.adv_id}/export") + assert response.status_code == 200 + return response.json() + + +# ------------------------------------------------------------------ the rules + +def _counts() -> dict: + """Every row that an import can create, per table.""" + with SessionLocal() as db: + return { + model.__name__: db.query(model).count() + for model in ( + models.Adventure, models.Branch, models.Action, models.Memory, + models.Summary, models.Checkpoint, models.StateEvent, + models.StateProposal, models.KnowledgeSource, + models.KnowledgeChunk, models.StoryCard, + ) + } + + +def refused(client, payload, *, status=(400, 409, 413, 422)) -> str: + """Imports expecting a refusal, and asserts that nothing at all landed.""" + before = _counts() + response = client.post("/api/adventures/import", json=payload) + assert response.status_code in status, ( + f"expected a refusal, got {response.status_code}: {response.text[:400]}" + ) + assert _counts() == before, ( + "a refused import wrote rows: " + f"{ {k: (before[k], v) for k, v in _counts().items() if before[k] != v} }" + ) + body = response.json() + return str(body.get("detail", body)) + + +def accepted(client, payload) -> int: + response = client.post("/api/adventures/import", json=payload) + assert response.status_code == 201, response.text[:500] + return response.json()["id"] + + +def broken(good: dict, **changes) -> dict: + return dict(copy.deepcopy(good), **changes) + + +# -------------------------------------------------------- format and version + +def test_a_payload_that_is_not_an_object_is_refused(client): + for payload in ([], "a string", 7): + response = client.post("/api/adventures/import", json=payload) + assert response.status_code in (400, 422), response.text[:200] + + +def test_an_empty_object_is_refused(client): + assert "format" in refused(client, {}).lower() or "export" in refused(client, {}) + + +def test_a_missing_format_is_refused(client, good): + payload = copy.deepcopy(good) + del payload["format"] + refused(client, payload) + + +def test_a_format_of_the_wrong_type_is_refused(client, good): + for wrong in (3, None, ["ai-dnd-adventure-v3"], {"v": 3}): + refused(client, broken(good, format=wrong)) + + +def test_an_unsupported_future_version_is_refused_with_its_name(client, good): + detail = refused(client, broken(good, format="ai-dnd-adventure-v42")) + assert "ai-dnd-adventure-v42" in detail + + +# ------------------------------------------------------------- the tree graph + +def test_an_action_on_a_branch_the_file_does_not_list_is_refused(client, good): + payload = copy.deepcopy(good) + payload["actions"][0]["branch"] = 99 + assert "99" in refused(client, payload) + + +def test_a_branch_forking_from_one_listed_after_it_is_refused(client, good): + """Which is also how a cycle is made impossible rather than detected. + + A branch may only fork from a branch listed before it, so the graph is + acyclic by construction. Without it a lineage walk on a hand-edited file + would not terminate. + """ + payload = copy.deepcopy(good) + payload["branches"][0] = {"parent": 1, "forkDepth": 0} + refused(client, payload) + + +def test_a_branch_that_forks_from_itself_is_refused(client, good): + payload = copy.deepcopy(good) + payload["branches"][1] = {"parent": 1, "forkDepth": 3} + refused(client, payload) + + +def test_a_fork_with_no_depth_is_refused(client, good): + payload = copy.deepcopy(good) + payload["branches"][1] = {"parent": 0} + assert "depth" in refused(client, payload) + + +def test_an_action_with_no_depth_is_refused(client, good): + payload = copy.deepcopy(good) + payload["actions"][1]["depth"] = None + assert "depth" in refused(client, payload) + + +def test_an_action_with_a_negative_depth_is_refused(client, good): + payload = copy.deepcopy(good) + payload["actions"][1]["depth"] = -4 + refused(client, payload) + + +def test_a_head_past_the_story_is_refused(client, good): + assert "ends at" in refused(client, broken(good, headDepth=10_000)) + + +def test_a_head_depth_of_the_wrong_type_is_refused(client, good): + for wrong in ("3", 3.5, True, [3]): + refused(client, broken(good, headDepth=wrong)) + + +def test_a_head_branch_that_is_not_listed_falls_back_to_the_root(client, good): + """Repaired rather than refused, and the repair is the safe direction. + + The head *depth* is checked against the story and refused when it disagrees, + because a wrong depth silently moves the reader. A head *branch* that names + nothing cannot be read at all, so there is no wrong position to land at — + the root is where a campaign with no chosen branch is read. + """ + payload = copy.deepcopy(good) + payload["headBranch"] = 77 + payload.pop("headDepth") # the depth belongs to the branch it names + copy_id = accepted(client, payload) + with SessionLocal() as db: + adventure = db.get(models.Adventure, copy_id) + root = ( + db.query(models.Branch) + .filter(models.Branch.adventure_id == copy_id, + models.Branch.parent_branch_id.is_(None)) + .first() + ) + assert adventure.head_branch_id == root.id + + +def test_two_actions_claiming_one_identity_are_refused(client, good): + """Take parentage and the whole audit trail hang off these ids.""" + payload = copy.deepcopy(good) + payload["actions"][1]["id"] = payload["actions"][0]["id"] + assert "both call themselves" in refused(client, payload) + + +def test_a_turn_whose_takes_are_all_dead_still_tells_one(client, good): + """Repaired, because a turn with no live attempt disappears from the story.""" + payload = copy.deepcopy(good) + for action in payload["actions"]: + action["live"] = False + copy_id = accepted(client, payload) + with SessionLocal() as db: + rows = ( + db.query(models.Action) + .filter(models.Action.adventure_id == copy_id) + .all() + ) + per_turn = {} + for row in rows: + per_turn.setdefault((row.branch_id, row.depth), []).append(row) + for group in per_turn.values(): + assert sum(1 for row in group if row.live) == 1 + + +def test_a_parent_naming_a_node_the_file_does_not_hold_is_ignored(client, good): + """Dropped, not refused: a wrong parent costs a pager, not a campaign.""" + payload = copy.deepcopy(good) + for action in payload["actions"]: + if action.get("parentId") is not None: + action["parentId"] = 999_999 + copy_id = accepted(client, payload) + story = client.get(f"/api/adventures/{copy_id}").json() + assert story["actions"], "the campaign did not import" + + +def test_a_node_that_is_its_own_parent_does_not_loop(client, good): + payload = copy.deepcopy(good) + for action in payload["actions"]: + if action.get("id") is not None: + action["parentId"] = action["id"] + copy_id = accepted(client, payload) + with SessionLocal() as db: + assert db.query(models.Action).filter( + models.Action.adventure_id == copy_id, + models.Action.parent_id == models.Action.id, + ).count() == 0 + # And the pager still resolves rather than recursing. + assert client.get(f"/api/adventures/{copy_id}").status_code == 200 + + +# ----------------------------------------------------------------- save points + +def test_a_save_point_beyond_the_retained_story_is_dropped_not_retargeted( + client, good +): + payload = copy.deepcopy(good) + original = payload["checkpoints"][0]["name"] + payload["checkpoints"][0]["depth"] = 5_000 + copy_id = accepted(client, payload) + landed = client.get(f"/api/adventures/{copy_id}/checkpoints").json() + assert original not in {point["name"] for point in landed} + assert all(point["depth"] < 5_000 for point in landed) + assert landed, "the good Save Point was lost with the bad one" + + +def test_a_save_point_on_a_branch_that_is_not_listed_is_dropped(client, good): + payload = copy.deepcopy(good) + payload["checkpoints"][0]["branch"] = 44 + copy_id = accepted(client, payload) + landed = client.get(f"/api/adventures/{copy_id}/checkpoints").json() + assert len(landed) == len(good["checkpoints"]) - 1 + + +def test_a_save_point_with_a_blank_name_is_dropped(client, good): + payload = copy.deepcopy(good) + payload["checkpoints"][0]["name"] = " " + copy_id = accepted(client, payload) + assert len(client.get(f"/api/adventures/{copy_id}/checkpoints").json()) == \ + len(good["checkpoints"]) - 1 + + +def test_a_checkpoints_section_that_is_not_a_list_costs_the_bookmarks_only( + client, good +): + copy_id = accepted(client, broken(good, checkpoints={"nope": 1})) + assert client.get(f"/api/adventures/{copy_id}/checkpoints").json() == [] + assert client.get(f"/api/adventures/{copy_id}").json()["actions"] + + +# ------------------------------------------------------------ state and audit + +def test_a_state_section_that_is_not_a_list_is_refused(client, good): + assert "list" in refused(client, broken(good, stateEvents={"a": 1})) + assert "list" in refused(client, broken(good, stateProposals="events")) + + +def test_a_state_event_that_is_not_an_object_is_refused(client, good): + payload = copy.deepcopy(good) + payload["stateEvents"][0] = "an event" + refused(client, payload) + + +def test_a_state_event_with_no_type_is_refused(client, good): + payload = copy.deepcopy(good) + payload["stateEvents"][0]["eventType"] = "" + assert "type" in refused(client, payload) + + +def test_an_event_naming_a_turn_the_file_does_not_hold_is_refused(client, good): + payload = copy.deepcopy(good) + payload["stateEvents"][0]["action"] = 424_242 + assert "424242" in refused(client, payload).replace(",", "") + + +def test_a_proposal_naming_a_turn_the_file_does_not_hold_is_refused(client, good): + payload = copy.deepcopy(good) + payload["stateProposals"][0]["action"] = 424_242 + refused(client, payload) + + +def test_an_event_naming_a_proposal_that_is_gone_keeps_its_coordinate(client, good): + """`ON DELETE SET NULL`, as a file. The event is the accepted change. + + A proposal can be deleted while the event it produced stands — the schema + says so — so an event whose proposal is not in the file is not a broken + file. It loses the pointer and keeps everything that makes it an audit + record: what changed, where, and who asserted it. + """ + payload = copy.deepcopy(good) + payload["stateProposals"] = [] + copy_id = accepted(client, payload) + events = client.get( + f"/api/adventures/{copy_id}/state/events?limit=500" + ).json() + assert len(events) == len(good["stateEvents"]) + assert any(e["source"] == "manual_correction" for e in events) + with SessionLocal() as db: + assert db.query(models.StateEvent).filter( + models.StateEvent.adventure_id == copy_id, + models.StateEvent.proposal_id.isnot(None), + ).count() == 0 + + +def test_a_malformed_narrative_state_costs_the_state_and_not_the_campaign( + client, good +): + """M5's rule, unchanged: a malformed document is normalised, not fatal. + + The story is the valuable thing. A state section that arrives as nonsense + becomes an empty document — which is honest, because nothing in it can be + trusted — and every turn still imports. + """ + copy_id = accepted(client, broken(good, narrativeState={"entities": "wrong"})) + story = client.get(f"/api/adventures/{copy_id}").json() + assert len(story["actions"]) == len( + client.get(f"/api/adventures/{client.adv_id}").json()["actions"] + ) + state = client.get(f"/api/adventures/{copy_id}/state").json() + assert state["document"]["entities"] == {} + + +def test_a_per_position_snapshot_that_is_not_an_object_is_dropped(client, good): + payload = copy.deepcopy(good) + for action in payload["actions"]: + if "narrativeStateAfter" in action: + action["narrativeStateAfter"] = "not a document" + copy_id = accepted(client, payload) + assert client.get(f"/api/adventures/{copy_id}").status_code == 200 + # Arriving at such a position gives the empty document rather than a + # later position's state, which is M5's finding 3. + client.post(f"/api/adventures/{copy_id}/undo") + assert client.get(f"/api/adventures/{copy_id}/state").json()["document"]["facts"] == [] + + +# -------------------------------------------------------------- knowledge + +def test_a_knowledge_section_that_is_not_a_list_is_refused(client, good): + assert "list" in refused(client, broken(good, knowledge={"a": 1})) + + +def test_a_source_with_no_content_is_refused(client, good): + """Refused rather than dropped, and M7 chose that deliberately. + + A campaign whose imported Canon quietly did not arrive is a campaign whose + narrator has stopped being told the rules, and the reader has no way to + notice. + """ + payload = copy.deepcopy(good) + payload["knowledge"][0]["content"] = "" + assert "content" in refused(client, payload) + + +def test_a_source_with_an_unknown_classification_is_refused(client, good): + payload = copy.deepcopy(good) + payload["knowledge"][0]["classification"] = "gospel" + assert "classification" in refused(client, payload) + + +def test_a_source_that_is_not_an_object_is_refused(client, good): + payload = copy.deepcopy(good) + payload["knowledge"][0] = "canon.md" + refused(client, payload) + + +def test_an_unreadable_visibility_becomes_normal_rather_than_hidden(client, good): + """Repaired, and in the direction that reveals rather than conceals. + + Visibility is not a permission system — the person who imported the file can + always read it — so a source that should have been narrator-only and lands + as normal costs a spoiler in the prompt framing. The other direction would + silently withhold material the reader expects the narrator to use, with + nothing saying so. + """ + payload = copy.deepcopy(good) + for source in payload["knowledge"]: + source["visibility"] = "invisible" + copy_id = accepted(client, payload) + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + assert all(source["visibility"] == "normal" for source in library) + + +def test_a_content_hash_that_disagrees_is_recomputed_and_reported(client, good): + """The one derived value in the file, and the only reason it is there. + + The stored hash is recomputed from what actually arrived, so it always + describes the content. The file's own claim is not silently discarded + either: a mismatch means the file was edited after it was written, and the + reader is told on the source itself. + """ + payload = copy.deepcopy(good) + payload["knowledge"][0]["contentHash"] = "0" * 64 + copy_id = accepted(client, payload) + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + edited = [s for s in library if s["content_hash"] != "0" * 64] + assert len(edited) == len(library) + detail = client.get( + f"/api/adventures/{copy_id}/knowledge/{library[0]['id']}" + ).json() + assert "did not match" in detail["notes"] + + +def test_more_sources_than_the_cap_is_refused(client, good, monkeypatch): + from app.knowledge import importer + + monkeypatch.setattr(importer, "MAX_SOURCES_PER_ADVENTURE", 2) + assert "limit" in refused(client, good) + + +def test_an_oversized_source_is_refused(client, good, monkeypatch): + from app.knowledge import importer + + monkeypatch.setattr(importer, "MAX_SOURCE_BYTES", 32) + assert "larger than" in refused(client, good) + + +# ------------------------------------------------------------- provenance + +def test_a_context_snapshot_that_is_not_an_object_is_dropped(client, good): + """Evidence is restored verbatim or not at all. It is never guessed at.""" + payload = copy.deepcopy(good) + payload["actions"] = [ + {k: v for k, v in action.items() if k != "contextSnapshotZ"} + | ({"contextSnapshot": "the prompt was long"} + if m9_fixture.snapshot_in(action) else {}) + for action in payload["actions"] + ] + copy_id = accepted(client, payload) + page = client.get(f"/api/adventures/{copy_id}").json() + narrator = [a for a in page["actions"] if a["type"] == "ai"] + assert narrator + for action in narrator: + response = client.get( + f"/api/adventures/{copy_id}/actions/{action['id']}/context" + ) + assert response.status_code == 404, "a mangled snapshot was restored" + + +def test_a_snapshot_whose_knowledge_block_is_nonsense_does_not_break_the_import( + client, good +): + payload = copy.deepcopy(good) + rewritten = [] + for action in payload["actions"]: + snapshot = m9_fixture.snapshot_in(action) + if isinstance(snapshot, dict) and "knowledge" in snapshot: + snapshot["knowledge"] = ["not", "a", "report"] + rewritten.append(m9_fixture.with_snapshot(action, snapshot)) + else: + rewritten.append(action) + payload["actions"] = rewritten + copy_id = accepted(client, payload) + assert client.get(f"/api/adventures/{copy_id}").status_code == 200 + + +def test_a_snapshot_naming_an_impossible_source_is_relinked_to_nothing( + client, good +): + payload = copy.deepcopy(good) + rewritten = [] + for action in payload["actions"]: + snapshot = m9_fixture.snapshot_in(action) + if not isinstance(snapshot, dict): + rewritten.append(action) + continue + for record in (snapshot.get("knowledge") or {}).get("used") or []: + record["source_id"] = -1 + rewritten.append(m9_fixture.with_snapshot(action, snapshot)) + payload["actions"] = rewritten + copy_id = accepted(client, payload) + with SessionLocal() as db: + from sqlalchemy.orm import undefer + + for row in ( + db.query(models.Action) + .filter(models.Action.adventure_id == copy_id) + .options(undefer(models.Action.context_snapshot)) + ): + snapshot = row.context_snapshot + if not isinstance(snapshot, dict): + continue + for record in (snapshot.get("knowledge") or {}).get("used") or []: + assert record["source_id"] is None + + +# ------------------------------------------------------------ summaries + +def test_a_summary_with_no_coordinate_is_dropped_not_placed(client, good): + """Placing it at a guess is how E03's leak would arrive by a new route.""" + payload = copy.deepcopy(good) + payload["summaries"][0]["depth"] = None + copy_id = accepted(client, payload) + with SessionLocal() as db: + landed = db.query(models.Summary).filter( + models.Summary.adventure_id == copy_id + ).count() + assert landed == len(good["summaries"]) - 1 + + +def test_a_summaries_section_that_is_not_a_list_costs_the_summaries_only( + client, good +): + copy_id = accepted(client, broken(good, summaries="a paragraph")) + assert client.get(f"/api/adventures/{copy_id}").json()["actions"] + with SessionLocal() as db: + assert db.query(models.Summary).filter( + models.Summary.adventure_id == copy_id + ).count() == 0 + + +# ------------------------------------------------------------ caps and size + +def test_more_actions_than_the_cap_is_refused(client, good, monkeypatch): + monkeypatch.setattr(limits, "MAX_ACTIONS_PER_ADVENTURE", 3) + monkeypatch.setitem(limits._BUNDLE_LIST_CAPS, "actions", 3) + assert "limit" in refused(client, good) + + +def test_more_branches_than_the_cap_is_refused(client, good, monkeypatch): + monkeypatch.setattr(limits, "MAX_BRANCHES_PER_ADVENTURE", 1) + monkeypatch.setitem(limits._BUNDLE_LIST_CAPS, "branches", 1) + assert "limit" in refused(client, good) + + +def test_a_body_past_the_import_ceiling_is_refused_before_it_is_parsed(client): + """413 from the middleware, on the declared length, before any read.""" + padding = "x" * (limits.MAX_IMPORT_BODY_BYTES + 1024) + response = client.post( + "/api/adventures/import", + content=json.dumps({"format": "ai-dnd-adventure-v3", "title": padding}), + headers={"Content-Type": "application/json"}, + ) + assert response.status_code == 413 + assert "too large" in response.json()["detail"].lower() + + +# ------------------------------------------------- the transaction, not the plan + +def test_a_failure_deep_inside_the_write_leaves_nothing_behind( + client, good, monkeypatch +): + """The other half of atomicity, and the half the planner cannot provide. + + Every test above is refused by `bundle.plan`, which has no side effects — so + they prove the *planner*, and a passing planner would look identical if the + write phase left debris. This one breaks something the planner has already + approved, half way through writing: the branches, the nodes, their + parentage, the memories, the head and the Save Points are all in the session + by then. + + What must survive that is the whole transaction rolling back — every table, + not merely the adventure row. A half-written campaign is the outcome L01 + forbids for a turn, and an import is the other place it could happen. + """ + from app import bundle as bundle_module + + def explode(*args, **kwargs): + raise RuntimeError("simulated failure deep inside the write") + + monkeypatch.setattr(bundle_module, "_write_summaries", explode) + before = _counts() + with pytest.raises(RuntimeError, match="simulated failure"): + client.post("/api/adventures/import", json=good) + assert _counts() == before, ( + "a failed write left rows behind: " + f"{ {k: (before[k], v) for k, v in _counts().items() if before[k] != v} }" + ) + + +def test_the_session_is_usable_after_a_failed_import(client, good, monkeypatch): + """The rollback is explicit, so the next request is not poisoned by it. + + Left to the session closing, a failure would leave the request's session in + a state the next caller inherits only by luck of pooling. `bundle_io` rolls + back and re-raises, so the very next import succeeds. + """ + from app import bundle as bundle_module + + calls = {"n": 0} + original = bundle_module._write_summaries + + def once(*args, **kwargs): + calls["n"] += 1 + if calls["n"] == 1: + raise RuntimeError("simulated, once") + return original(*args, **kwargs) + + monkeypatch.setattr(bundle_module, "_write_summaries", once) + with pytest.raises(RuntimeError): + client.post("/api/adventures/import", json=good) + copy_id = accepted(client, good) + assert client.get(f"/api/adventures/{copy_id}").json()["actions"] + + +# --------------------------------------------------------------- inert data + +def test_a_url_in_a_bundle_stays_text(client, good): + """H01/G08 for the import path: nothing in a file is ever fetched. + + There is no allowlist to test and no request to intercept, which is the + result rather than a gap — the import has no code that could make one. What + is asserted is that the text arrives as text. + """ + payload = copy.deepcopy(good) + payload["knowledge"][0]["content"] = ( + "# Sources\n\nSee https://example.invalid/secret.txt and " + "file:///etc/passwd and ![map](https://example.invalid/map.png)\n" + ) + copy_id = accepted(client, payload) + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + detail = client.get( + f"/api/adventures/{copy_id}/knowledge/{library[0]['id']}" + ).json() + assert "https://example.invalid/secret.txt" in detail["content"] + + +def test_a_path_in_a_bundle_never_becomes_a_path(client, good): + """H08. `originalFilename` is metadata; the import stores no file.""" + payload = copy.deepcopy(good) + for hostile in ("../../../etc/passwd", "/etc/shadow", "C:\\Windows\\hosts", + "....//....//etc/passwd"): + payload["knowledge"][0]["originalFilename"] = hostile + copy_id = accepted(client, payload) + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + for source in library: + assert "/" not in source["original_filename"] + assert "\\" not in source["original_filename"] + assert ".." not in source["original_filename"] + + +def test_a_title_that_looks_like_a_command_is_stored_as_a_title(client, good): + payload = broken(good, title="; rm -rf / #") + copy_id = accepted(client, payload) + assert client.get(f"/api/adventures/{copy_id}").json()["title"] == "; rm -rf / #" + + +def test_an_over_long_title_is_truncated_rather_than_refused(client, good): + copy_id = accepted(client, broken(good, title="A" * 5_000)) + title = client.get(f"/api/adventures/{copy_id}").json()["title"] + assert 0 < len(title) <= 200 diff --git a/backend/tests/test_m9_legacy_bundles.py b/backend/tests/test_m9_legacy_bundles.py new file mode 100644 index 0000000..dc29fc6 --- /dev/null +++ b/backend/tests/test_m9_legacy_bundles.py @@ -0,0 +1,365 @@ +"""M9: every older bundle still imports, and none is reinterpreted. + +A backup that stops importing is not a backup, so the importer keeps every +version it has ever written. That is the easy half. The hard half is the rule +`V1-ACCEPTANCE-TESTS.md` I07 states about the head and this file generalises: + +> Do not reinterpret missing legacy data using modern assumptions that did not +> exist when the file was written. + +An older file is missing things because its **format** could not carry them, not +because the campaign lacked them, and the two demand opposite treatment. A file +written before the head was carried opens at its tip, because tip was the only +position that format could represent — reproducing what it recorded. A file +written before state events existed opens with no state events, because +manufacturing an audit trail from the snapshots it does carry would be this +build's reading of a history it never saw, handed to a reader as the record of +what happened. + +Each seam below is built by taking a real v3 export and removing exactly what +the older format could not hold. That is deliberate: a checked-in fixture file +drifts, and a hand-written one tests a shape nothing ever wrote. + + python -m pytest tests/test_m9_legacy_bundles.py -v +""" + +import copy + +import pytest +from fastapi import Depends +from fastapi.testclient import TestClient + +from app import auth, bundle, limits, memorybank, models +from app.database import Base, SessionLocal, engine, get_db +from app.knowledge import embeddings +from app.main import app +from app.routers import adventures + +import m9_fixture +from fakes import ScriptedProvider +from test_m9_portability import StubDerived + + +@pytest.fixture() +def client(monkeypatch): + Base.metadata.create_all(bind=engine) + memorybank._vector_cache.clear() + embeddings._cache.clear() + setup = SessionLocal() + user = models.User(is_guest=False, email="legacy@example.com") + setup.add(user) + setup.flush() + setup.add(models.Settings( + user_id=user.id, model="test-model", embedding_model="stub-embed", + context_token_budget=4000, max_output_tokens=400, + )) + adventure = models.Adventure( + user_id=user.id, title="Source", + campaign_canon=m9_fixture.CAMPAIGN_CANON, + ) + setup.add(adventure) + setup.flush() + setup.add(models.Action( + adventure_id=adventure.id, type="start", text=m9_fixture.OPENING, + )) + setup.commit() + adv_id, user_id = adventure.id, user.id + setup.close() + + monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None) + monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider) + monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived()) + monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived()) + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id) + ) + test_client = TestClient(app) + test_client.adv_id = adv_id + try: + yield test_client + finally: + app.dependency_overrides.clear() + adventures.turns._active_turns.clear() + memorybank._vector_cache.clear() + embeddings._cache.clear() + Base.metadata.drop_all(bind=engine) + + +@pytest.fixture() +def current(client): + """A real v3 export of the M9 fixture, to age backwards from.""" + m9_fixture.build(client, client.adv_id) + response = client.get(f"/api/adventures/{client.adv_id}/export") + assert response.status_code == 200 + return response.json() + + +# ---------------------------------------------------- ageing a bundle backwards + +def as_of(payload: dict, era: str) -> dict: + """The same campaign as an export from an earlier era. + + Each step removes only what that era's format genuinely could not carry, so + the result is the file a build of that vintage would have produced from this + campaign — not a mutilated modern one. + """ + older = copy.deepcopy(payload) + eras = ("pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head") + assert era in eras, era + reached = eras.index(era) + + # M9 (v3): the evidence sections and the node identities. + older["format"] = bundle.TREE_FORMAT + for key in ("stateEvents", "stateProposals", "summaries"): + older.pop(key, None) + for action in older["actions"]: + for key in ("contextSnapshot", "contextSnapshotZ", "id", "parentId"): + action.pop(key, None) + for memory in older.get("memories") or []: + memory.pop("authority", None) + for source in older.get("knowledge") or []: + for key in ("sourceId", "parserVersion", "chunkingVersion"): + source.pop(key, None) + if reached == 0: + return older + + # M7: the imported knowledge library. + older.pop("knowledge", None) + if reached == 1: + return older + + # M5: the authoritative narrative state, its per-position snapshots, and + # the campaign's own canon. + for key in ("narrativeState", "campaignCanon"): + older.pop(key, None) + for action in older["actions"]: + for key in ("narrativeStateAfter", "stateChanges"): + action.pop(key, None) + if reached == 2: + return older + + # M4: named Save Points. + older.pop("checkpoints", None) + if reached == 3: + return older + + # M3: the chosen head. Such a file could only ever be read at its tip. + older.pop("headDepth", None) + return older + + +def bring_back(client, payload) -> int: + response = client.post("/api/adventures/import", json=payload) + assert response.status_code == 201, response.text[:500] + return response.json()["id"] + + +def _rows(adv_id, model) -> int: + with SessionLocal() as db: + return db.query(model).filter(model.adventure_id == adv_id).count() + + +def _tree_size(client, adv_id) -> int: + """Every retained row, which is what "no accepted story was lost" means.""" + return len(client.get(f"/api/adventures/{adv_id}/export").json()["actions"]) + + +# --------------------------------------------------------------- every era + +@pytest.mark.parametrize("era", [ + "pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head", +]) +def test_no_accepted_story_is_lost_at_any_seam(client, current, era): + """The floor under every case below: the turns all arrive. + + Counted over the whole retained tree rather than the active path, because + the head moves between eras and a count of what is on screen would move + with it. + """ + copy_id = bring_back(client, as_of(current, era)) + assert _tree_size(client, copy_id) == len(current["actions"]) + + +@pytest.mark.parametrize("era", [ + "pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head", +]) +def test_nothing_is_invented_to_fill_a_gap_the_format_left(client, current, era): + """Absent means the format could not say. It never means "make one up". + + Each era is checked against what that era's files could hold: a pre-M9 file + gets no audit trail and no summaries, a pre-M7 file no knowledge, a pre-M5 + file no state, a pre-Save-Point file no Save Points. + """ + copy_id = bring_back(client, as_of(current, era)) + reached = ("pre-m9", "pre-m7", "pre-m5", "pre-save-points", + "pre-active-head").index(era) + + assert _rows(copy_id, models.StateEvent) == 0 + assert _rows(copy_id, models.StateProposal) == 0 + assert _rows(copy_id, models.Summary) == 0 + if reached >= 1: + assert _rows(copy_id, models.KnowledgeSource) == 0 + assert _rows(copy_id, models.KnowledgeChunk) == 0 + if reached >= 2: + state = client.get(f"/api/adventures/{copy_id}/state").json() + assert state["document"]["facts"] == [] + assert state["document"]["entities"] == {} + assert client.get(f"/api/adventures/{copy_id}").json()["canon_rules"] == [] + if reached >= 3: + assert _rows(copy_id, models.Checkpoint) == 0 + + +# ------------------------------------------------------ the head, era by era + +def test_a_pre_m9_file_still_opens_at_the_head_it_recorded(client, current): + """v2 carried the head, so it is honoured exactly as before.""" + copy_id = bring_back(client, as_of(current, "pre-m9")) + with SessionLocal() as db: + assert db.get(models.Adventure, copy_id).head_depth == current["headDepth"] + + +def test_a_pre_active_head_file_opens_at_its_tip(client, current): + """I07's compatibility clause. Not a degraded path. + + Such a file was written when the head could not be anywhere but the tip, so + opening it there reproduces the position it recorded. An import that refused + it, or that guessed some other position, would be the failure. + """ + copy_id = bring_back(client, as_of(current, "pre-active-head")) + with SessionLocal() as db: + adventure = db.get(models.Adventure, copy_id) + tip = max( + row.depth for row in + db.query(models.Action).filter( + models.Action.adventure_id == copy_id, + models.Action.branch_id == adventure.head_branch_id, + ) + ) + assert adventure.head_depth == tip + assert adventure.head_depth > current["headDepth"], ( + "the fixture's head must really be behind its tip, or this proves nothing" + ) + + +def test_a_pre_active_head_file_offers_no_redo_because_it_is_at_the_tip( + client, current +): + copy_id = bring_back(client, as_of(current, "pre-active-head")) + page = client.get(f"/api/adventures/{copy_id}").json() + assert page["can_redo"] is False + assert page["can_undo"] is True + + +# ----------------------------------------------------- what each era can do + +def test_a_pre_m5_campaign_can_be_played_on_and_gains_state_from_there( + client, current +): + """The M5 rule, applied to an import: no backfill, and no obstacle either. + + An old campaign starts with an empty state because its narration was never + read by a state extractor. The next turn fills it in, which is what makes + "no backfill" a decision rather than a loss. + """ + copy_id = bring_back(client, as_of(current, "pre-m5")) + assert client.get(f"/api/adventures/{copy_id}/state").json()["empty"] is True + + ScriptedProvider.replies = [ + "The door gives at last.\n" + __import__("fakes").state_block([ + {"type": "add_fact", "predicate": "tally", "value": 500, + "fact_id": "tally-500"} + ]) + ] + played = client.post(f"/api/adventures/{copy_id}/actions", + json={"type": "do", "text": "push harder"}) + assert played.status_code == 200, played.text[:300] + after = client.get(f"/api/adventures/{copy_id}/state").json() + assert after["empty"] is False + assert any(f["predicate"] == "tally" for f in after["document"]["facts"]) + + +def test_a_pre_m7_campaign_needs_no_source_and_can_import_one(client, current): + copy_id = bring_back(client, as_of(current, "pre-m7")) + assert client.get(f"/api/adventures/{copy_id}/knowledge").json() == [] + # It plays without one. + assert client.get(f"/api/adventures/{copy_id}/context").status_code == 200 + # And gains one. + landed = m9_fixture.upload( + client, copy_id, "canon.md", m9_fixture.CANON_MD, "canon", + ) + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + assert [s["id"] for s in library] == [landed] + assert library[0]["index_state"] == "ready" + + +def test_a_pre_save_point_campaign_can_be_given_one(client, current): + copy_id = bring_back(client, as_of(current, "pre-save-points")) + assert client.get(f"/api/adventures/{copy_id}/checkpoints").json() == [] + made = client.post(f"/api/adventures/{copy_id}/checkpoints", + json={"name": "From here", "note": ""}) + assert made.status_code == 201, made.text[:300] + assert made.json()["resolved"] is True + + +def test_a_pre_m9_campaign_re_exports_as_v3_without_gaining_evidence( + client, current +): + """Re-exporting an old campaign does not turn absence into presence. + + The file it writes is a v3 file, because that is what this build writes. Its + evidence sections are empty, because the campaign genuinely has none — and a + later reader can therefore trust a v3 file's empty `stateEvents` to mean + "this campaign has no audit trail" rather than "the file could not say". + """ + copy_id = bring_back(client, as_of(current, "pre-m9")) + again = client.get(f"/api/adventures/{copy_id}/export").json() + assert again["format"] == bundle.FORMAT + assert again["stateEvents"] == [] + assert again["stateProposals"] == [] + assert again["summaries"] == [] + assert not any(a.get("contextSnapshotZ") for a in again["actions"]) + # And the story it does have survives a second round trip unchanged. + twice = bring_back(client, again) + assert _tree_size(client, twice) == _tree_size(client, copy_id) + + +def test_a_v1_file_still_imports_and_reads_in_order(client): + """The flat format, with its retries as a repeating group.""" + copy_id = bring_back(client, { + "format": bundle.LEGACY_FORMAT, + "title": "An old flat file", + "memory": "Kept from before the tree.", + "actions": [ + {"index": 0, "type": "start", "text": "It begins."}, + {"index": 1, "type": "do", "text": "look around"}, + {"index": 2, "type": "ai", "text": "Take two.", + "variants": [{"text": "Take one."}, {"text": "Take two."}], + "variantIndex": 1}, + ], + }) + page = client.get(f"/api/adventures/{copy_id}").json() + assert [a["text"] for a in page["actions"]] == [ + "It begins.", "look around", "Take two.", + ] + assert page["memory"] == "Kept from before the tree." + # Both attempts arrived; only one is the story. + with SessionLocal() as db: + rows = db.query(models.Action).filter( + models.Action.adventure_id == copy_id, models.Action.type == "ai", + ).all() + assert sorted(r.text for r in rows) == ["Take one.", "Take two."] + assert sum(1 for r in rows if r.live) == 1 + + +def test_a_pre_m2_file_with_scripting_still_imports(client, current): + """M2 removed campaign scripting. Its keys are ignored, not rejected. + + The story, the tree and everything else in such a file are still worth + importing, and refusing the campaign over a subsystem that no longer exists + would lose all of it to reject one key. + """ + payload = as_of(current, "pre-m5") + payload["scripts"] = [{"name": "onTurn", "code": "state.gold += 10"}] + payload["scriptState"] = {"gold": 70} + copy_id = bring_back(client, payload) + assert _tree_size(client, copy_id) == len(current["actions"]) diff --git a/backend/tests/test_m9_portability.py b/backend/tests/test_m9_portability.py new file mode 100644 index 0000000..ccca2f2 --- /dev/null +++ b/backend/tests/test_m9_portability.py @@ -0,0 +1,1186 @@ +"""M9: a campaign survives being moved, and can say why it is what it is. + +The Definition of Done is one sentence — *a campaign can be safely exported, +imported into a clean data directory, and reopened at the exact intended active +position with authoritative history/state intact* — and this file is the part of +its evidence that runs in the suite. Two things it deliberately does not claim: + +* **"A clean data directory" is not proved here.** Every test in this file runs + against one database, and importing beside the original is a weaker check than + importing where the original has never existed — a shared row, a shared id + space or a shared cache would go unnoticed. `test_m9_clean_import.py` does that + across a genuine second process and a genuinely empty database, and this file + says so rather than implying otherwise. +* **A green run here is not the milestone.** The browser workflow, the process + restart and the real-narrator run are in their own places, for the reason M8 + wrote down: a suite that stubs a boundary is not evidence about that boundary. + +What this file does prove is the shape of the contract. `m9_fixture` builds one +campaign that holds every portable data family at once — an undone head, a +retained future, an abandoned line with its own state and derived data, two +takes at one coordinate, Save Points on two branches, a manual state correction, +five imported sources across three classes and both lifecycle states, and stored +prompts for turns that used them. Every assertion below is a comparison between +that campaign and its copy, read through the API a reader reads. + + python -m pytest tests/test_m9_portability.py -v +""" + +import asyncio +import copy + +import pytest +from fastapi import Depends +from fastapi.testclient import TestClient + +from app import auth, bundle, limits, memorybank, models, summaries +from app.context import lineage +from app.database import Base, SessionLocal, engine, get_db +from app.knowledge import embeddings +from app.main import app +from app.routers import adventures + +import m9_fixture +from fakes import ScriptedProvider + + +class StubDerived: + """A deterministic embedder and summariser in one object. + + Both factories are stubbed from it, which is M6's finding M6-F3: replacing + only one leaves the other building a real provider against the default + endpoint, and every turn in the file opens a socket. + """ + + written = 0 + + async def complete(self, system, prompt, **kwargs): + StubDerived.written += 1 + return f"Memory {StubDerived.written}: what the story had established." + + async def embed(self, texts): + out = [] + for text in texts: + lowered = text.lower() + out.append([ + 1.0, + 1.0 if "abbey" in lowered or "crypt" in lowered else 0.0, + 1.0 if "tavern" in lowered or "lantern" in lowered else 0.0, + 1.0 if "rain" in lowered else 0.0, + 1.0 if "seal" in lowered else 0.0, + ]) + return out + + +@pytest.fixture() +def client(monkeypatch): + Base.metadata.create_all(bind=engine) + memorybank._vector_cache.clear() + embeddings._cache.clear() + StubDerived.written = 0 + setup = SessionLocal() + user = models.User(is_guest=False, email="m9@example.com") + setup.add(user) + setup.flush() + setup.add(models.Settings( + user_id=user.id, model="test-model", embedding_model="stub-embed", + context_token_budget=4000, max_output_tokens=400, memory_top_k=3, + )) + adventure = models.Adventure( + user_id=user.id, title="M9 Portability Fixture", + campaign_canon=m9_fixture.CAMPAIGN_CANON, + ) + setup.add(adventure) + setup.flush() + setup.add(models.Action( + adventure_id=adventure.id, type="start", text=m9_fixture.OPENING, + )) + # A neighbour campaign. A bundle that reached past its own rows would bring + # this one's story back with it, and every count below would still add up. + neighbour = models.Adventure(user_id=user.id, title="Neighbour") + setup.add(neighbour) + setup.flush() + setup.add(models.Action( + adventure_id=neighbour.id, type="start", + text="A different story entirely, which must not travel.", + )) + setup.commit() + adv_id, other_id, user_id = adventure.id, neighbour.id, user.id + setup.close() + + monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None) + monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider) + monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived()) + monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived()) + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id) + ) + test_client = TestClient(app) + test_client.adv_id = adv_id + test_client.other_id = other_id + try: + yield test_client + finally: + app.dependency_overrides.clear() + adventures.turns._active_turns.clear() + memorybank._vector_cache.clear() + embeddings._cache.clear() + Base.metadata.drop_all(bind=engine) + + +# ----------------------------------------------------------------- helpers + +def export(client, adv_id=None) -> dict: + response = client.get(f"/api/adventures/{adv_id or client.adv_id}/export") + assert response.status_code == 200, response.text[:400] + return response.json() + + +def bring_back(client, payload) -> int: + """Imports a bundle and returns the new campaign's id.""" + response = client.post("/api/adventures/import", json=payload) + assert response.status_code == 201, response.text[:600] + return response.json()["id"] + + +def refuses(client, payload) -> str: + """Imports a bundle expecting a refusal, and returns the reason given.""" + before = _campaign_count() + response = client.post("/api/adventures/import", json=payload) + assert response.status_code in (400, 409), ( + f"expected a refusal, got {response.status_code}: {response.text[:300]}" + ) + assert _campaign_count() == before, "and nothing was written" + return response.json()["detail"] + + +def _campaign_count() -> int: + with SessionLocal() as db: + return db.query(models.Adventure).count() + + +def rows(adv_id, model, **filters): + with SessionLocal() as db: + query = db.query(model).filter(model.adventure_id == adv_id) + for field, value in filters.items(): + query = query.filter(getattr(model, field) == value) + return query.order_by(model.id).all() + + +@pytest.fixture() +def moved(client): + """The fixture campaign, its bundle, and the copy the bundle produced. + + Built once per test that needs it, because building it is a second of + scripted turns and most of the assertions below are about the same round + trip seen from different angles. + """ + source = m9_fixture.build(client, client.adv_id) + payload = export(client) + copy_id = bring_back(client, payload) + return { + "source": source, + "bundle": payload, + "copy_id": copy_id, + "copy": m9_fixture.snapshot_of(client, copy_id), + } + + +# --------------------------------------------------- the format and its versions + +def test_the_export_declares_version_three(moved): + assert moved["bundle"]["format"] == "ai-dnd-adventure-v3" + + +def test_every_earlier_version_still_imports(client): + """A backup that stops importing is not a backup. + + Both older versions are exercised, and each is checked for the thing its own + format could carry: a v1 file has a story, and a v2 file has a story, a tree + and a chosen head. + """ + m9_fixture.build(client, client.adv_id) + v3 = export(client) + + v2 = _as_version_two(v3) + copy_id = bring_back(client, v2) + assert _texts(client, copy_id) == _texts(client, client.adv_id) + + v1 = { + "format": bundle.LEGACY_FORMAT, + "title": "An old file", + "actions": [ + {"index": 0, "type": "start", "text": "It begins."}, + {"index": 1, "type": "do", "text": "look around"}, + {"index": 2, "type": "ai", "text": "Nothing moves."}, + ], + } + v1_id = bring_back(client, v1) + assert _texts(client, v1_id) == ["It begins.", "look around", "Nothing moves."] + + +def test_a_version_two_file_gets_no_invented_evidence(client): + """Absent is read as "the format could not say", never as "there was none". + + The distinction is the whole reason the version was bumped. A v2 file + carries no state events, and the import must leave the campaign with none + rather than reconstructing an audit trail from the snapshots it does carry — + a reconstructed trail would be this build's reading of a history it never + saw, presented to a reader as the record of what happened. + """ + m9_fixture.build(client, client.adv_id) + copy_id = bring_back(client, _as_version_two(export(client))) + + assert rows(copy_id, models.StateEvent) == [] + assert rows(copy_id, models.StateProposal) == [] + assert rows(copy_id, models.Summary) == [] + assert all( + row.context_snapshot is None + for row in _all_actions(copy_id) + ), "a v2 file carries no prompts, and none may be invented" + # What v2 *could* say still arrives in full. + assert _texts(client, copy_id) == _texts(client, client.adv_id) + assert len(rows(copy_id, models.Checkpoint)) == 2 + + +def test_a_format_from_a_later_build_is_refused_rather_than_guessed_at(client): + detail = refuses(client, dict(export(client), format="ai-dnd-adventure-v9")) + assert "ai-dnd-adventure-v9" in detail + assert bundle.FORMAT in detail + + +def _as_version_two(payload: dict) -> dict: + """The same campaign as a version 2 file: no v3 sections, no v3 node keys. + + This is what an export taken before M9 looks like, built by removing exactly + what M9 added rather than by keeping a fixture file that would drift. + """ + older = copy.deepcopy(payload) + older["format"] = bundle.TREE_FORMAT + for key in ("stateEvents", "stateProposals", "summaries"): + older.pop(key, None) + for action in older["actions"]: + for key in ("contextSnapshot", "contextSnapshotZ", "id", "parentId"): + action.pop(key, None) + for memory in older.get("memories") or []: + memory.pop("authority", None) + for source in older.get("knowledge") or []: + for key in ("sourceId", "parserVersion", "chunkingVersion"): + source.pop(key, None) + return older + + +# ------------------------------------------------------------------ I01, I02 + +def test_i01_the_export_holds_enough_to_restore_the_story(moved): + """I01. Not "the file is non-empty": the file names every family.""" + payload = moved["bundle"] + for section in ("actions", "branches", "checkpoints", "knowledge", + "memories", "summaries", "stateEvents", "stateProposals"): + assert payload[section], f"the bundle carries no {section}" + assert payload["headDepth"] is not None + assert payload["narrativeState"] + assert any(m9_fixture.snapshot_in(a) for a in payload["actions"]) + + +def test_i02_the_copy_reads_the_same_story_and_holds_the_same_state(moved): + """I02. The active transcript and the authoritative state, both restored.""" + assert moved["copy"]["transcript"] == moved["source"]["transcript"] + assert moved["copy"]["state"] == moved["source"]["state"] + assert moved["copy"]["state"], "and the state is not empty in both" + assert moved["copy"]["canon_rules"] == moved["source"]["canon_rules"] + + +def test_the_neighbouring_campaign_does_not_travel(moved, client): + """A bundle carries one campaign. Nothing from the one beside it.""" + everything = _texts(client, moved["copy_id"], every_branch=True) + assert not any("A different story entirely" in text for text in everything) + + +# ------------------------------------------------------- I07 and the head rules + +def test_i07_the_copy_opens_at_the_exported_head_not_at_the_tip(moved, client): + """I07. The head is where the reader left it, and the future is still there. + + The fixture's head is behind the retained tip of its own branch *and* behind + the abandoned line's, and the fixture ends with an Undo precisely so that + the newest row written is not the position being read. An importer that + took the tip, the newest row, or the deepest row would each land somewhere + else. + """ + source_head = _head(client.adv_id) + copy_head = _head(moved["copy_id"]) + assert copy_head["depth"] == source_head["depth"] + assert moved["copy"]["transcript"] == moved["source"]["transcript"] + + # The future past the head is still in the database, unread. + ahead = [ + row for row in _all_actions(moved["copy_id"]) + if row.branch_id == copy_head["branch_id"] and row.depth > copy_head["depth"] + ] + assert ahead, "the retained future did not survive the round trip" + + # And Redo is offered, rather than the story having silently been redone. + assert moved["copy"]["can_redo"] is True + assert moved["source"]["can_redo"] is True + + +def test_redo_after_import_walks_the_future_that_was_retained(moved, client): + """Redo is coherent in the copy: it reaches the same next turn.""" + before = client.post(f"/api/adventures/{moved['copy_id']}/redo") + assert before.status_code == 200, before.text[:300] + original = client.post(f"/api/adventures/{client.adv_id}/redo") + assert original.status_code == 200, original.text[:300] + assert _texts(client, moved["copy_id"]) == _texts(client, client.adv_id) + + +def test_a_file_written_before_the_head_was_carried_opens_at_its_tip(client): + """I07's compatibility clause, and why it is not a degraded path. + + Such a file was written when the head could not be anywhere but the tip, so + opening it there reproduces the position it actually recorded. Guessing some + other position for it would be the failure. + """ + m9_fixture.build(client, client.adv_id) + payload = export(client) + stated = payload.pop("headDepth") + copy_id = bring_back(client, payload) + landed = _head(copy_id) + tip = max( + row.depth for row in _all_actions(copy_id) + if row.branch_id == landed["branch_id"] + ) + assert landed["depth"] == tip + assert landed["depth"] > stated, "and the fixture's head really was behind it" + + +def test_a_head_beyond_the_story_the_file_carries_is_refused(client): + """A file disagreeing with itself is refused, not opened at a guess.""" + m9_fixture.build(client, client.adv_id) + payload = export(client) + detail = refuses(client, dict(payload, headDepth=payload["headDepth"] + 500)) + assert "ends at" in detail + + +# ------------------------------------------------------------- I03, takes + +def test_i03_both_futures_survive_and_stay_distinguishable(moved, client): + """I03. The abandoned line comes back, still marked as abandoned.""" + assert moved["copy"]["branch_count"] == moved["source"]["branch_count"] == 2 + + stories = _texts(client, moved["copy_id"], every_branch=True) + assert any("The crypt is dry" in text for text in stories), \ + "the abandoned future is gone" + assert any("windows are lit" in text for text in stories), \ + "the continuation the reader chose is gone" + + # Disposition, not merely presence: the line the story left is still the + # line the story left, at the depth it left it. + with SessionLocal() as db: + branches = ( + db.query(models.Branch) + .filter(models.Branch.adventure_id == moved["copy_id"]) + .order_by(models.Branch.id).all() + ) + superseded = [b for b in branches if b.superseded_at is not None] + assert len(superseded) == 1 + assert superseded[0].superseded_depth == 6 + + +def test_a_superseded_take_survives_and_the_selected_one_is_still_selected(moved): + """No accepted row disappears for being inactive.""" + source_takes = _takes(moved["source"]["id"]) + copy_takes = _takes(moved["copy_id"]) + assert copy_takes == source_takes + assert sum(1 for _, live in copy_takes if not live) == 1, \ + "the retained alternate take is gone" + assert sum(1 for _, live in copy_takes if live) == 1 + + +def test_the_takes_at_one_turn_are_still_grouped_as_one_turn(moved, client): + """M9's parentage. The pager reads the same in the copy as in the original. + + Version 2 exported no parentage, so every imported node landed parentless + and `attempts.group` fell back to the coordinate. That is right for a simple + retry and wrong as soon as two takes of one turn each have takes of their + own beneath them. + """ + assert _take_pagers(client, moved["copy_id"]) == \ + _take_pagers(client, client.adv_id) + with SessionLocal() as db: + parented = ( + db.query(models.Action) + .filter(models.Action.adventure_id == moved["copy_id"], + models.Action.parent_id.isnot(None)) + .count() + ) + assert parented > 0, "no imported node knows which turn it belongs to" + + +# ------------------------------------------------------------------------ I04 + +def test_i04_save_points_survive_with_their_names_notes_and_positions(moved): + assert moved["copy"]["checkpoints"] == moved["source"]["checkpoints"] + assert len(moved["copy"]["checkpoints"]) == 2 + + +def test_each_restored_save_point_reaches_the_position_it_names(moved, client): + """And restoring one uses M3's head movement, leaving later history alone.""" + copy_id = moved["copy_id"] + before = len(_all_actions(copy_id)) + for point in client.get(f"/api/adventures/{copy_id}/checkpoints").json(): + assert point["resolved"] is True, f"{point['name']} resolves to nothing" + restored = client.post( + f"/api/adventures/{copy_id}/checkpoints/{point['id']}/restore" + ) + assert restored.status_code == 200, restored.text[:300] + assert _head(copy_id)["depth"] == point["depth"] + assert len(_all_actions(copy_id)) == before, "restoring deleted history" + + +def test_the_two_save_points_restore_to_different_states(moved, client): + """They name different positions, so they must restore different states.""" + copy_id = moved["copy_id"] + seen = [] + for point in client.get(f"/api/adventures/{copy_id}/checkpoints").json(): + client.post(f"/api/adventures/{copy_id}/checkpoints/{point['id']}/restore") + seen.append(client.get(f"/api/adventures/{copy_id}/state").json()["document"]) + assert seen[0] != seen[1] + + +def test_a_save_point_naming_a_position_the_file_does_not_hold_is_dropped(client): + """Deliberate, documented, and never retargeted somewhere else. + + A bad bookmark costs the bookmark. Refusing the campaign over it would lose + the story to save the bookmark, and moving it to a nearby turn would be an + invention — the reader named a position, and if that position is not in the + file then no other position is the one they named. + """ + m9_fixture.build(client, client.adv_id) + payload = export(client) + payload["checkpoints"][0]["depth"] = 9999 + copy_id = bring_back(client, payload) + names = [c["name"] for c in + client.get(f"/api/adventures/{copy_id}/checkpoints").json()] + assert len(names) == 1, "the good Save Point did not survive beside the bad one" + assert payload["checkpoints"][0]["name"] not in names + assert not any( + c["depth"] == 9999 + for c in client.get(f"/api/adventures/{copy_id}/checkpoints").json() + ) + + +# ------------------------------------------------------- the state and its audit + +def test_the_accepted_state_events_come_back_with_their_before_values(moved): + source_events = _events(moved["source"]["id"]) + copy_events = _events(moved["copy_id"]) + assert copy_events == source_events + assert len(copy_events) == 12 + + +def test_a_manual_correction_is_still_identifiable_as_one(moved): + """C04's authority survives the move. + + This is the state change no narration explains, and with the events omitted + it was indistinguishable from something the story established — the copy + showed the fact and could not say who asserted it. + """ + corrections = [ + event for event in _events(moved["copy_id"]) + if event["source"] == "manual_correction" + ] + assert len(corrections) == 1 + assert corrections[0]["event_type"] == "add_fact" + assert corrections[0]["payload"]["predicate"] == "keeper_of_the_lantern" + + +def test_the_proposals_come_back_and_the_events_still_name_them(moved): + """The two tables arrive linked, not merely both present.""" + with SessionLocal() as db: + proposals = ( + db.query(models.StateProposal) + .filter(models.StateProposal.adventure_id == moved["copy_id"]) + .all() + ) + events = ( + db.query(models.StateEvent) + .filter(models.StateEvent.adventure_id == moved["copy_id"]) + .all() + ) + ids = {p.id for p in proposals} + assert len(proposals) == 12 + linked = [e for e in events if e.proposal_id is not None] + assert linked, "every event lost the proposal that produced it" + assert all(e.proposal_id in ids for e in linked), \ + "an event points at a proposal that is not in this campaign" + # And the proposals point at this campaign's own turns, not the source's. + theirs = {row.id for row in _all_actions(moved["copy_id"])} + assert all(p.action_id in theirs for p in proposals if p.action_id is not None) + + +def test_a_state_event_naming_a_turn_the_file_does_not_hold_is_refused(client): + """Unlike a Save Point, and for a stated reason. + + An audit record that quietly did not arrive leaves a campaign whose state + cannot be explained, and it is the explanation a reader goes looking for + exactly when something looks wrong. + """ + m9_fixture.build(client, client.adv_id) + payload = export(client) + payload["stateEvents"][0]["action"] = 999_999 + detail = refuses(client, payload) + assert "999999" in detail.replace(",", "") + + +def test_state_restoration_after_import_is_a_snapshot_read_not_a_replay(moved, client): + """L02, and the bound M4 made load-bearing. + + Undo, Redo and Save Point restore all resolve a coordinate and read the + state recorded there. If the import had dropped the per-position snapshots + and left only the events, every one of those would have to replay the + campaign — so the check is that each position in the copy holds the same + state as the same position in the original, walked the same way. + """ + copy_id = moved["copy_id"] + walked_copy, walked_source = [], [] + for _ in range(3): + for adv_id, out in ((copy_id, walked_copy), (client.adv_id, walked_source)): + assert client.post(f"/api/adventures/{adv_id}/undo").status_code == 200 + out.append(client.get(f"/api/adventures/{adv_id}/state").json()["document"]) + assert walked_copy == walked_source + for _ in range(3): + for adv_id, out in ((copy_id, walked_copy), (client.adv_id, walked_source)): + assert client.post(f"/api/adventures/{adv_id}/redo").status_code == 200 + out.append(client.get(f"/api/adventures/{adv_id}/state").json()["document"]) + assert walked_copy == walked_source + + +def test_the_abandoned_line_still_holds_its_own_different_state(moved, client): + """E01, after a move. Two futures, two states, and neither leaks.""" + copy_id = moved["copy_id"] + at_head = client.get(f"/api/adventures/{copy_id}/state").json()["document"] + assert at_head != moved["source"]["tip_state"], ( + "the fixture's two futures must differ, or this proves nothing" + ) + + +# ---------------------------------------------- historical prompt provenance (§10) + +def test_an_old_turn_can_still_show_what_it_was_actually_given(moved, client): + """The M8 handoff, closed. + + Inspect Context on a historical narrator turn in the *copy* returns the + prompt that turn was assembled from — the same sections, the same text — and + not a prompt rebuilt from the campaign as it stands now. + """ + source_turn, copy_turn = _first_narrator_turn(client, client.adv_id), \ + _first_narrator_turn(client, moved["copy_id"]) + original = _context_of(client, client.adv_id, source_turn) + restored = _context_of(client, moved["copy_id"], copy_turn) + assert restored["prompt"] == original["prompt"] + assert [s["label"] for s in restored["sections"]] == \ + [s["label"] for s in original["sections"]] + assert restored["tokens"] == original["tokens"] + + +def test_the_old_turn_still_names_the_passages_it_was_shown(moved, client): + """F06's provenance, restored. Including the text each passage supplied.""" + turn = _first_narrator_turn(client, moved["copy_id"]) + knowledge = _context_of(client, moved["copy_id"], turn)["knowledge"] + assert knowledge["used"], "the restored turn was shown no imported passage" + for record in knowledge["used"]: + assert record["text"], "a passage record arrived with no text" + assert record["classification"] in ("canon", "reference", "inspiration") + + +def test_the_evidence_outlives_the_source_being_deleted(moved, client): + """`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50, across a machine boundary. + + The record holds the text rather than a pointer to a row that can go away, + so deleting the source in the copy leaves the old turn still able to say + what it was told — and the "open this source" link honestly goes quiet + rather than pointing somewhere else. + """ + copy_id = moved["copy_id"] + turn = _first_narrator_turn(client, copy_id) + shown = _context_of(client, copy_id, turn)["knowledge"]["used"] + was_shown = {record["source_id"] for record in shown} + assert was_shown, "nothing was shown, so there is nothing to outlive" + + for source_id in was_shown: + assert client.delete( + f"/api/adventures/{copy_id}/knowledge/{source_id}" + ).status_code == 204 + + after = _context_of(client, copy_id, turn)["knowledge"]["used"] + assert [r["text"] for r in after] == [r["text"] for r in shown] + assert [r["title"] for r in after] == [r["title"] for r in shown] + + +def test_a_restored_retrieval_record_points_at_this_machines_source(moved, client): + """The one pointer the import translates, and why. + + `source_id` names a row on the machine that wrote the file. Left alone it + would point the inspector's "open this source" control at whatever holds + that id here — nothing, or somebody else's file. + """ + copy_id = moved["copy_id"] + theirs = { + source["id"] for source in + client.get(f"/api/adventures/{copy_id}/knowledge").json() + } + turn = _first_narrator_turn(client, copy_id) + used = _context_of(client, copy_id, turn)["knowledge"]["used"] + named = [r["source_id"] for r in used if r["source_id"] is not None] + assert named, "no restored record names a source at all" + assert set(named) <= theirs, "a record points outside this campaign's library" + + +def test_a_snapshot_naming_a_source_the_file_does_not_carry_goes_quiet(moved, client): + """A source deleted before the export was taken. + + The record keeps its text and its filename and says `null` for the source, + which is the honest answer: this passage came from a file that is no longer + here, and here is what it said. It must not be pointed at a different file. + """ + payload = moved["bundle"] + assert any(m9_fixture.snapshot_in(a) for a in payload["actions"]), \ + "the fixture carries no snapshot to edit" + edited = copy.deepcopy(payload) + rewritten = [] + for action in edited["actions"]: + snapshot = m9_fixture.snapshot_in(action) + if not isinstance(snapshot, dict): + rewritten.append(action) + continue + for record in (snapshot.get("knowledge") or {}).get("used") or []: + record["source_id"] = 88_888 + rewritten.append(m9_fixture.with_snapshot(action, snapshot)) + edited["actions"] = rewritten + copy_id = bring_back(client, edited) + for row in _all_actions(copy_id): + snapshot = row.context_snapshot + if not isinstance(snapshot, dict): + continue + for record in (snapshot.get("knowledge") or {}).get("used") or []: + assert record["source_id"] is None + assert record["text"], "and the evidence itself is untouched" + + +def test_a_retried_turn_does_not_multiply_the_prompt(moved): + """The prompt is stored once per turn, and travels once per turn. + + A superseded take keeps only its own slices — its reply, its state proposal, + its token accounting — so a campaign someone retried twenty times does not + carry twenty copies of the largest thing in the file. + """ + superseded = [ + snapshot for action in moved["bundle"]["actions"] + if not action.get("live", True) + for snapshot in [m9_fixture.snapshot_in(action)] if snapshot + ] + assert superseded, "the fixture's retry left no superseded take with a record" + for take in superseded: + assert "sections" not in take + assert "prompt" not in take + + +# ------------------------------------------------------------------------ I05 + +def test_i05_the_library_survives_with_its_classifications_and_lifecycle(moved): + assert moved["copy"]["knowledge"] == moved["source"]["knowledge"] + assert len(moved["copy"]["knowledge"]) == 5 + + +def test_the_disabled_source_is_still_disabled_and_the_hidden_one_still_hidden( + moved, client +): + library = client.get(f"/api/adventures/{moved['copy_id']}/knowledge").json() + by_file = {source["original_filename"]: source for source in library} + assert by_file["draft.md"]["enabled"] is False + assert by_file["secret.md"]["visibility"] == "hidden" + assert by_file["canon.md"]["always_include"] is True + assert by_file["canon.md"]["classification"] == "canon" + # And the filename really is only metadata: the title is the source's own, + # derived at import time from its content rather than from its path. + assert by_file["canon.md"]["title"] == "canon" + + +def test_the_copy_needs_no_original_file_and_is_searchable_at_once(moved, client): + """§11. Nothing here reopens a path on the exporting machine. + + The content came in the file, and the passages were rebuilt from it before + the import returned — so retrieval works with no reindex step and no + explanation owed to the reader. + """ + copy_id = moved["copy_id"] + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + assert all(source["index_state"] == "ready" for source in library) + assert all(source["chunk_count"] > 0 for source in library) + # The filename is metadata and nothing else: it never became a path. + assert all("/" not in source["original_filename"] for source in library) + + +def test_the_disabled_source_stays_out_of_retrieval_after_the_move(moved, client): + report = client.get(f"/api/adventures/{moved['copy_id']}/context").json() + used = {record["filename"] for record in report["knowledge"]["used"]} + assert "draft.md" not in used + + +# ----------------------------------------------------- summaries and memories + +def test_the_summaries_come_back_on_the_coordinates_that_hold_them(moved): + assert moved["copy"]["summaries"] == moved["source"]["summaries"] + assert len(moved["copy"]["summaries"]) == 2 + + +def test_an_abandoned_lines_summary_is_still_ineligible_after_the_move(moved): + """E03 does not get a second chance through the import. + + The fixture writes one summary on the line it later abandons and one at the + head it keeps. Both travel. Eligibility is not a stored flag — it is whether + the coordinate lies on the active capped lineage — so restoring the + coordinates restores the answer, including the "no". + """ + eligible = { + (text, trigger): is_eligible + for text, trigger, is_eligible in moved["copy"]["summaries"] + } + assert sorted(eligible.values()) == [False, True], ( + "the copy should hold exactly one eligible and one ineligible summary" + ) + assert eligible[("Aldric went back to the Lantern instead.", "manual")] is True + + +def test_the_abandoned_summary_does_not_reach_the_copys_next_prompt(moved, client): + """E03, measured where it matters: the prompt the copy would send next. + + The generated summary sits on the line the fixture abandoned, and the typed + one sits at the head. Only the second may reach a prompt. The provenance + record names the coordinate rather than carrying the text, so eligibility is + checked there and the leak is checked in the assembled prompt itself. + """ + copy_id = moved["copy_id"] + report = client.get(f"/api/adventures/{copy_id}/context").json() + prompt = report["prompt"]["system"] + report["prompt"]["story"] + + with SessionLocal() as db: + adventure = db.get(models.Adventure, copy_id) + entitled = summaries.current(db, adventure) + every = {row.id: row for row in summaries.all_for(db, adventure)} + assert entitled is not None, "the copy is entitled to no summary at all" + assert entitled.trigger == "manual", "the abandoned line's summary is eligible" + + abandoned = [row for row in every.values() if row.id != entitled.id] + assert abandoned, "the fixture's abandoned summary did not survive the move" + for row in abandoned: + assert row.text not in prompt, "an abandoned summary reached the prompt" + assert report["summary"]["id"] == entitled.id + + +def test_the_summary_mirror_agrees_with_the_lineage_after_import(moved, client): + """M6-F1's defect must not arrive by a new route.""" + copy_head = client.get(f"/api/adventures/{moved['copy_id']}").json() + with SessionLocal() as db: + adventure = db.get(models.Adventure, moved["copy_id"]) + entitled = summaries.current(db, adventure) + assert copy_head["story_summary"] == (entitled.text if entitled else "") + + +def test_memory_authority_survives_rather_than_being_promoted(client): + """F07. A heuristic memory must not become accepted story by being moved.""" + m9_fixture.build(client, client.adv_id) + with SessionLocal() as db: + memory = ( + db.query(models.Memory) + .filter(models.Memory.adventure_id == client.adv_id) + .order_by(models.Memory.id).first() + ) + memory.authority = memorybank.HEURISTIC + db.commit() + copy_id = bring_back(client, export(client)) + authorities = [row.authority for row in rows(copy_id, models.Memory)] + assert memorybank.HEURISTIC in authorities + assert authorities == [row.authority for row in rows(client.adv_id, models.Memory)] + + +def test_a_memory_comes_back_on_the_node_it_hangs_off(moved): + source_rows = [(m.text, m.depth) for m in rows(moved["source"]["id"], models.Memory)] + copy_rows = [(m.text, m.depth) for m in rows(moved["copy_id"], models.Memory)] + assert copy_rows == source_rows + + +# ---------------------------------------------------- L04, derived data rebuild + +def test_l04_deleting_the_derived_indexes_and_rebuilding_changes_no_story( + moved, client +): + """L04, on a copy rather than on the original. + + Every physically derived structure is destroyed — passages, the FTS rows and + the vectors — and rebuilt from the source content the bundle carried. The + authoritative campaign must be identical either side, and retrieval must + work again afterwards. + """ + copy_id = moved["copy_id"] + before = m9_fixture.snapshot_of(client, copy_id) + with SessionLocal() as db: + db.query(models.KnowledgeEmbedding).filter( + models.KnowledgeEmbedding.adventure_id == copy_id + ).delete(synchronize_session=False) + db.query(models.KnowledgeChunk).filter( + models.KnowledgeChunk.adventure_id == copy_id + ).delete(synchronize_session=False) + db.commit() + + empty = client.get(f"/api/adventures/{copy_id}/knowledge").json() + assert all(source["chunk_count"] == 0 for source in empty) + + rebuilt = client.post(f"/api/adventures/{copy_id}/knowledge/reindex") + assert rebuilt.status_code == 200, rebuilt.text[:300] + + after = client.get(f"/api/adventures/{copy_id}/knowledge").json() + assert all(source["chunk_count"] > 0 for source in after) + assert all(source["index_state"] == "ready" for source in after) + # Content, classification and lifecycle are untouched by a rebuild. + assert m9_fixture.snapshot_of(client, copy_id)["knowledge"] == before["knowledge"] + # And the authoritative campaign did not move. + assert m9_fixture.snapshot_of(client, copy_id)["transcript"] == before["transcript"] + assert m9_fixture.snapshot_of(client, copy_id)["state"] == before["state"] + # Retrieval works again. + report = client.get(f"/api/adventures/{copy_id}/context").json() + assert report["knowledge"]["used"] + + +def test_losing_the_semantic_half_leaves_the_lexical_half_working(moved, client): + """M7's rule, which M9 must not regress during a rebuild. + + Lexical retrieval is a supported production path, not a fallback, so a + campaign whose vectors are gone still finds its Canon. + """ + copy_id = moved["copy_id"] + with SessionLocal() as db: + db.query(models.KnowledgeEmbedding).filter( + models.KnowledgeEmbedding.adventure_id == copy_id + ).delete(synchronize_session=False) + db.commit() + report = client.get(f"/api/adventures/{copy_id}/context").json() + assert report["knowledge"]["used"], "lexical retrieval stopped with the vectors" + + +def test_deleting_a_campaign_takes_its_lexical_index_with_it(moved, client): + """M9 finding: the FTS index is a virtual table and nothing cascades into it. + + Deleting a campaign dropped its passages and left one index row per passage + behind. Nothing read them — the search joins through `knowledge_chunks`, and + those were gone — so the leak was invisible until the id came round again. + """ + from sqlalchemy import text as sql + + copy_id = moved["copy_id"] + with SessionLocal() as db: + before = db.execute(sql("SELECT COUNT(*) FROM knowledge_fts")).scalar() + assert before > 0 + assert client.delete(f"/api/adventures/{copy_id}").status_code == 204 + with SessionLocal() as db: + orphans = db.execute(sql(""" + SELECT COUNT(*) FROM knowledge_fts + WHERE rowid NOT IN (SELECT id FROM knowledge_chunks) + """)).scalar() + assert orphans == 0 + + +def test_an_orphaned_index_row_does_not_break_the_next_import(client): + """The other half, and the one that repairs a database already damaged. + + SQLite hands out the lowest free primary key, so an orphan left by an older + build is met head-on by the next campaign to import anything — in a campaign + with no connection to the one that leaked it. It used to be a 500 from an + ordinary upload, and Reindex could not clear it either, because + `clear_index` finds index rows through chunks that no longer exist. + """ + from sqlalchemy import text as sql + + # An orphan, written the way a pre-M9 build would have left one. + with SessionLocal() as db: + db.execute( + sql("INSERT OR REPLACE INTO knowledge_fts (rowid, text) " + "VALUES (1, 'left behind by a deleted campaign')") + ) + db.commit() + + landed = m9_fixture.upload( + client, client.adv_id, "canon.md", m9_fixture.CANON_MD, "canon", + ) + library = client.get(f"/api/adventures/{client.adv_id}/knowledge").json() + source = next(s for s in library if s["id"] == landed) + assert source["index_state"] == "ready" + assert source["chunk_count"] > 0 + # And what the orphan said is gone rather than searchable. + with SessionLocal() as db: + stale = db.execute(sql( + "SELECT COUNT(*) FROM knowledge_fts WHERE text LIKE '%left behind%'" + )).scalar() + assert stale == 0 + + +def test_reindex_repairs_an_index_that_lost_its_passages(client): + """L04's repair path, exercised against the damage it exists to repair. + + Deleting the chunk rows without the index rows is what a cascade used to do, + and Reindex is documented as the repair. It has to actually be one. + """ + m9_fixture.build(client, client.adv_id) + with SessionLocal() as db: + db.query(models.KnowledgeChunk).filter( + models.KnowledgeChunk.adventure_id == client.adv_id + ).delete(synchronize_session=False) + db.commit() + + rebuilt = client.post(f"/api/adventures/{client.adv_id}/knowledge/reindex") + assert rebuilt.status_code == 200, rebuilt.text[:400] + assert rebuilt.json()["failed"] == [] + library = client.get(f"/api/adventures/{client.adv_id}/knowledge").json() + assert all(source["index_state"] == "ready" for source in library) + assert all(source["chunk_count"] > 0 for source in library) + report = client.get(f"/api/adventures/{client.adv_id}/context").json() + assert report["knowledge"]["used"] + + +def test_a_rebuilt_index_does_not_reseed_an_abandoned_summary(moved, client): + """M6's leak must not return through the rebuild path either.""" + copy_id = moved["copy_id"] + client.post(f"/api/adventures/{copy_id}/knowledge/reindex") + eligible = { + (text, trigger): is_eligible + for text, trigger, is_eligible in + m9_fixture.snapshot_of(client, copy_id)["summaries"] + } + assert sorted(eligible.values()) == [False, True] + + +# ------------------------------------------------------------- story cards (§13) + +def test_story_cards_still_travel_in_both_directions(client): + """Compatibility-only, but compatibility means a round trip loses nothing.""" + with SessionLocal() as db: + db.add(models.StoryCard( + adventure_id=client.adv_id, type="character", name="Gwen", + keys="gwen, innkeeper", entry="Gwen keeps the road-house.", + notes="From an older campaign.", + )) + db.commit() + payload = export(client) + assert payload["storyCards"] == [{ + "type": "character", "name": "Gwen", "keys": "gwen, innkeeper", + "entry": "Gwen keeps the road-house.", "notes": "From an older campaign.", + }] + copy_id = bring_back(client, payload) + assert [(c.name, c.entry) for c in rows(copy_id, models.StoryCard)] == \ + [("Gwen", "Gwen keeps the road-house.")] + + +def test_a_story_card_no_longer_reaches_the_narrator(client): + """§73: not an alternate untracked path around knowledge authority. + + Before M9 a keyword match put `World Lore: ` in front of the narrator + with no class, no visibility, no source, no way to switch it off and no row + in the context inspector. The rows stay and the round trip stays; the + injection does not. + """ + with SessionLocal() as db: + db.add(models.StoryCard( + adventure_id=client.adv_id, type="location", name="The road-house", + keys="road-house, roadhouse", + entry="The road-house at Fen Cross has burned down.", + )) + db.commit() + m9_fixture._play(client, client.adv_id, "ask about the road-house", + "She shrugs and pours another.") + report = client.get(f"/api/adventures/{client.adv_id}/context").json() + prompt = report["prompt"]["system"] + report["prompt"]["story"] + assert "World Lore" not in prompt + assert "has burned down" not in prompt + assert report["cards"] == [] + + +def test_a_historical_snapshot_that_carried_cards_still_renders_them(client): + """The `cards` key stays in the report shape, for the old evidence. + + An old turn's record says story cards were included, and M9 has just made + that record portable. Removing the key would make a restored campaign's own + history unreadable. + """ + m9_fixture.build(client, client.adv_id) + payload = export(client) + for position, action in enumerate(payload["actions"]): + snapshot = m9_fixture.snapshot_in(action) + if isinstance(snapshot, dict) and "cards" in snapshot: + snapshot["cards"] = [ + {"id": 7, "name": "Gwen", "keyword": "gwen", "included": True} + ] + payload["actions"][position] = m9_fixture.with_snapshot(action, snapshot) + break + else: + pytest.fail("no snapshot in the fixture bundle to edit") + copy_id = bring_back(client, payload) + found = [ + row.context_snapshot["cards"] for row in _all_actions(copy_id) + if isinstance(row.context_snapshot, dict) + and row.context_snapshot.get("cards") + ] + assert found == [[{"id": 7, "name": "Gwen", "keyword": "gwen", "included": True}]] + + +# ----------------------------------------------- scene / media metadata (§14) + +def test_the_scene_section_of_the_state_document_round_trips(moved): + """There are no media tables in this build, and none were invented. + + What exists is the `scene` section of the authoritative narrative state + document — `SPECIFICATION.md` §14's scene snapshot as this build represents + it — and it travels with the document, per position, like the rest of it. + """ + for action in moved["bundle"]["actions"]: + state = action.get("narrativeStateAfter") + if isinstance(state, dict): + assert "scene" in state + break + else: + pytest.fail("no per-position state document in the bundle") + assert "scene" in moved["bundle"]["narrativeState"] + + +# ------------------------------------------------------- I06 and local-only + +def test_i06_the_export_carries_no_credential_of_any_kind(moved, client): + """I06. Checked against the file's text, not against a list of columns. + + The application needs no cloud key, which makes this easy — and is exactly + why it is worth testing rather than assuming. The inert `api_key` column is + still in the schema, and a future field could reach the bundle by being + added to a model the exporter walks. + """ + import json + + with SessionLocal() as db: + settings = db.query(models.Settings).first() + settings.api_key = "enc:this-must-never-be-exported" + settings.endpoint_url = "http://127.0.0.1:11434/v1" + db.commit() + + text = json.dumps(export(client)) + assert "enc:this-must-never-be-exported" not in text + assert "api_key" not in text + assert "apiKey" not in text + # And no endpoint, model host or absolute filesystem path travels either. + assert "11434" not in text + assert "/home/" not in text + + +def test_the_bundle_names_no_path_that_could_become_one(moved): + """H08, for the import direction. A filename is metadata, never a path.""" + for source in moved["bundle"]["knowledge"]: + assert "/" not in source["originalFilename"] + assert "\\" not in source["originalFilename"] + assert ".." not in source["originalFilename"] + + +def test_a_bundle_carrying_a_traversal_filename_is_sanitised(client): + m9_fixture.build(client, client.adv_id) + payload = export(client) + payload["knowledge"][0]["originalFilename"] = "../../../etc/passwd" + copy_id = bring_back(client, payload) + library = client.get(f"/api/adventures/{copy_id}/knowledge").json() + assert all(".." not in source["original_filename"] for source in library) + assert all("/" not in source["original_filename"] for source in library) + + +# ------------------------------------------------------------------- helpers + +def _texts(client, adv_id, every_branch=False) -> list[str]: + if not every_branch: + return [a["text"] for a in + client.get(f"/api/adventures/{adv_id}").json()["actions"]] + out = [] + for branch in client.get(f"/api/adventures/{adv_id}/branches").json(): + client.post(f"/api/adventures/{adv_id}/branches/{branch['id']}/switch") + out.extend(_texts(client, adv_id)) + return out + + +def _head(adv_id) -> dict: + with SessionLocal() as db: + adventure = db.get(models.Adventure, adv_id) + return {"branch_id": adventure.head_branch_id, "depth": adventure.head_depth} + + +def _all_actions(adv_id) -> list[models.Action]: + with SessionLocal() as db: + from sqlalchemy.orm import undefer + return ( + db.query(models.Action) + .filter(models.Action.adventure_id == adv_id) + .options(undefer(models.Action.context_snapshot)) + .order_by(models.Action.branch_id, models.Action.depth, models.Action.id) + .all() + ) + + +def _takes(adv_id) -> list[tuple[str, bool]]: + """Every attempt at the retried turn, as text and liveness.""" + with SessionLocal() as db: + adventure = db.get(models.Adventure, adv_id) + rows_at = ( + db.query(models.Action) + .filter(models.Action.adventure_id == adv_id, + models.Action.type == "ai") + .order_by(models.Action.branch_id, models.Action.depth, + models.Action.id) + .all() + ) + seen = {} + for row in rows_at: + seen.setdefault((row.branch_id, row.depth), []).append(row) + for group in seen.values(): + if len(group) > 1: + return [(row.text, bool(row.live)) for row in group] + return [] + + +def _take_pagers(client, adv_id) -> list[tuple[int, int]]: + """The take pager under every narrator message, as (position, total).""" + page = client.get(f"/api/adventures/{adv_id}").json() + return [ + (action.get("take_index"), action.get("take_count")) + for action in page["actions"] if action["type"] == "ai" + ] + + +def _events(adv_id) -> list[dict]: + """The accepted state events, with the ids left out.""" + with SessionLocal() as db: + return [ + {"event_type": row.event_type, "payload": row.payload, + "before": row.before, "source": row.source, "depth": row.depth, + "sequence": row.sequence} + for row in db.query(models.StateEvent) + .filter(models.StateEvent.adventure_id == adv_id) + .order_by(models.StateEvent.id).all() + ] + + +def _first_narrator_turn(client, adv_id) -> int: + """The id of the earliest narrator action that has a stored prompt.""" + for row in _all_actions(adv_id): + if row.type == "ai" and row.live and isinstance(row.context_snapshot, dict) \ + and row.context_snapshot.get("prompt"): + return row.id + raise AssertionError(f"campaign {adv_id} has no turn with a stored prompt") + + +def _context_of(client, adv_id, action_id) -> dict: + response = client.get(f"/api/adventures/{adv_id}/actions/{action_id}/context") + assert response.status_code == 200, response.text[:300] + return response.json() diff --git a/backend/tests/test_retry_variants.py b/backend/tests/test_retry_variants.py index 4e09237..f9fd498 100644 --- a/backend/tests/test_retry_variants.py +++ b/backend/tests/test_retry_variants.py @@ -350,13 +350,17 @@ def test_export_and_import_round_trips_variants(client): assert [a["type"] for a in actions] == ["start", "do", "ai"] assert actions[-1]["text"] == "Two." - # The pager reads 1/1 on the copy, because the import writes no - # `parent_id` and `annotate_takes` groups on it. The attempts are both - # there, at one coordinate, and `GET .../variants` still lists them. This - # is a gap in the import rather than in the drop: `take_count` has been the - # only number the client reads since SP9, and the import has never set the - # column it is derived from. - assert actions[-1]["take_count"] == 1 + # The pager reads 2/2 on the copy, as it does on the original. + # + # It read 1/1 until M9, and this test recorded that as a gap in the import + # rather than in the export: the attempts were both there at one coordinate + # and `GET .../variants` listed them, but the import wrote no `parent_id`, + # so `annotate_takes` grouped on the coordinate instead. That is right for a + # plain retry and wrong the moment two takes of one turn each have takes of + # their own beneath them, which is why M9 carried the parentage rather than + # leaving the pager to a fallback. See `bundle._link_take_parents`. + assert actions[-1]["take_count"] == 2 + assert actions[-1]["take_index"] == 1 variants = client.get( f"/api/adventures/{imported}/actions/{actions[-1]['id']}/variants").json() assert [v["text"] for v in variants] == ["One.", "Two."] diff --git a/backend/tests/test_starter_adventure.py b/backend/tests/test_starter_adventure.py index 3ca393e..3470b37 100644 --- a/backend/tests/test_starter_adventure.py +++ b/backend/tests/test_starter_adventure.py @@ -50,10 +50,20 @@ def payload() -> dict: def test_the_shipped_file_is_a_bundle_this_build_can_import(): - """The file is written by an export, so a format change can strand it.""" + """The file is written by an export, so a format change can strand it. + + It is checked against every version the importer reads rather than against + the newest one it writes, which is the property that actually matters and + the one the shipped file has to keep. M9 bumped the format to v3 and did not + regenerate this asset: the starter is a linear story with no state events, + no summaries and no stored prompts, so a v3 rewrite of it would differ from + the v2 file in the version string alone — and rewriting a shipped asset to + keep a test's equality holding would be changing the evidence to fit the + test. What it does need is to go on importing, which is asserted below. + """ data = payload() version = bundle.check_format(data) - assert version == bundle.FORMAT + assert version in bundle.READABLE story = bundle.plan(data, version) assert story["nodes"] diff --git a/backend/tests/test_story_tree_baseline.py b/backend/tests/test_story_tree_baseline.py index 8d9caef..7e6d028 100644 --- a/backend/tests/test_story_tree_baseline.py +++ b/backend/tests/test_story_tree_baseline.py @@ -21,7 +21,7 @@ import pytest from fastapi import Depends from fastapi.testclient import TestClient -from app import auth, limits, models +from app import auth, bundle as bundle_module, limits, models from app.database import Base, SessionLocal, engine, get_db from app.main import app from app.routers import adventures @@ -425,6 +425,13 @@ def test_export_carries_the_whole_story(client): array in `test_export_keeps_retry_attempts`. That change reflects the same fact: a bundle that stores coordinates has no use for a repeating group. Everything else here still passes unmodified. + + M9 changed the same one line again, for the same kind of reason — the + version now says that the file can carry state events and historical + prompts as well as a tree. It is asserted against `bundle.FORMAT` this + time, so the next writer of a new version does not have to find this line: + what the test is about is that an export declares its version, not which + version this build happens to write. """ ScriptedProvider.replies = [gold_reply(t) for t in ["One.", "Two."]] _play(client, "go north") @@ -433,7 +440,7 @@ def test_export_carries_the_whole_story(client): r = client.get(f"/api/adventures/{client.adv_id}/export") assert r.status_code == 200, r.text bundle = r.json() - assert bundle["format"] == "ai-dnd-adventure-v2" + assert bundle["format"] == bundle_module.FORMAT assert bundle["title"] == "Cave" assert [a["text"] for a in bundle["actions"]] == [ OPENING, "> You go north.", "One.", "> You go south.", "Two.", diff --git a/backend/tools/m9_migration_proof.py b/backend/tools/m9_migration_proof.py new file mode 100644 index 0000000..02ec114 --- /dev/null +++ b/backend/tools/m9_migration_proof.py @@ -0,0 +1,331 @@ +"""M9's migration claim, proved against a database M8's own code wrote. + + # from the M8 worktree, using M8's interpreter: + python -m tools.m9_migration_proof build + + # from the M9 tree, using M9's interpreter: + python -m tools.m9_migration_proof open + +M9 claims to add no schema change. `git diff` proves that nothing in +`migrations.py` or `models.py` moved, which is necessary and not sufficient: a +migration can also be *missing*, and the failure then is a database that opens +and quietly answers wrongly. The M8 report set the standard here — a database +created by today's code and read by today's code proves nothing — so the +campaign below is built by a server running the signed M8 commit, from a git +worktree, and read back by M9. + +`build` writes a campaign that touches every family M9 changed the handling of: +story with an alternate take, a Save Point, a manual state correction, memories +and a summary, imported knowledge including a disabled and a narrator-only +source, and per-turn context snapshots. It prints what it wrote, as JSON. + +`open` opens that file with the current code, runs the migration path, and +checks every one of those against what `build` reported. It also asserts the +schema version did not move and that a second open is a no-op, which is what +"no migration" means in practice: the stamp is the same number before and after. + +Neither half imports anything from the other. What crosses is the database file +and one JSON report on stdout, which is the only way the two builds can be made +to talk without one of them importing the other's code. +""" + +from __future__ import annotations + +import argparse +import json +import os +import sys +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tests")) + + +def _app(db_path: str): + """Imports the application against `db_path`. Must run before any app import.""" + os.environ["AIDND_DB_PATH"] = db_path + os.environ.pop("AIDND_DATABASE_URL", None) + os.environ.pop("DATABASE_URL", None) + + from fastapi import Depends + from fastapi.testclient import TestClient + + from app import auth, limits, memorybank, models + from app.database import Base, SessionLocal, engine, get_db + from app.main import app + from app.routers import adventures + from fakes import ScriptedProvider, state_block + + class Stub: + async def complete(self, system, prompt, **kwargs): + return "A memory of what had happened by then." + + async def embed(self, texts): + return [[1.0, 0.5, 0.25] for _ in texts] + + adventures.turns.OpenAICompatibleProvider = ScriptedProvider + memorybank.embedding_provider = lambda s: Stub() + memorybank.summary_provider = lambda s: Stub() + limits.check_row_cap = lambda *a, **k: None + return { + "Base": Base, "SessionLocal": SessionLocal, "engine": engine, + "models": models, "app": app, "auth": auth, "get_db": get_db, + "Depends": Depends, "TestClient": TestClient, + "ScriptedProvider": ScriptedProvider, "state_block": state_block, + "memorybank": memorybank, + } + + +def _client(ctx, user_id: int): + ctx["app"].dependency_overrides[ctx["auth"].get_current_user] = ( + lambda db=ctx["Depends"](ctx["get_db"]): db.get(ctx["models"].User, user_id) + ) + return ctx["TestClient"](ctx["app"]) + + +# ------------------------------------------------------------------- building + +def build(db_path: str) -> dict: + ctx = _app(db_path) + ctx["Base"].metadata.create_all(bind=ctx["engine"]) + models, SessionLocal = ctx["models"], ctx["SessionLocal"] + + db = SessionLocal() + try: + user = models.User(is_guest=False, email="m9mig@example.com") + db.add(user) + db.flush() + db.add(models.Settings( + user_id=user.id, model="m8-model", embedding_model="stub", + context_token_budget=4000, max_output_tokens=400, + )) + adventure = models.Adventure( + user_id=user.id, title="Built by M8", auto_summarize=True, + memory_bank_enabled=True, + campaign_canon={"rules": ["The dead do not return."]}, + ) + db.add(adventure) + db.flush() + db.add(models.Action( + adventure_id=adventure.id, type="start", + text="Aldric sits in the Crooked Lantern with Mara.", + )) + db.commit() + adv_id, user_id = adventure.id, user.id + finally: + db.close() + + client = _client(ctx, user_id) + upload = client.post( + f"/api/adventures/{adv_id}/knowledge", + files={"file": ("canon.md", ( + "# Westhaven\n\n## The Old Abbey\n\nThe abbey crypt is sealed.\n" + ).encode(), "text/markdown")}, + data={"classification": "canon", "always_include": "true"}, + ) + assert upload.status_code == 201, upload.text[:300] + hidden = client.post( + f"/api/adventures/{adv_id}/knowledge", + files={"file": ("secret.md", ( + "# The seal\n\nIt was broken once, sixty years ago.\n" + ).encode(), "text/markdown")}, + data={"classification": "canon", "visibility": "hidden"}, + ) + assert hidden.status_code == 201, hidden.text[:300] + disabled = client.post( + f"/api/adventures/{adv_id}/knowledge", + files={"file": ("draft.md", b"# Draft\n\nAn earlier version.\n", + "text/markdown")}, + data={"classification": "reference"}, + ) + assert disabled.status_code == 201, disabled.text[:300] + assert client.patch( + f"/api/adventures/{adv_id}/knowledge/{disabled.json()['id']}", + json={"enabled": False}, + ).status_code == 200 + + state_block = ctx["state_block"] + for turn in range(1, 8): + ctx["ScriptedProvider"].replies = [ + f"The rain keeps on, and Mara says nothing for a while. [{turn}]\n" + + state_block([{"type": "add_fact", "predicate": "tally", + "value": turn * 10, "fact_id": f"tally-{turn * 10}"}]) + ] + response = client.post(f"/api/adventures/{adv_id}/actions", + json={"type": "do", "text": f"ask about turn {turn}"}) + assert response.status_code == 200, response.text[:300] + if turn == 3: + assert client.post(f"/api/adventures/{adv_id}/retry").status_code == 200 + point = client.post(f"/api/adventures/{adv_id}/checkpoints", + json={"name": "Third turn", "note": "A position."}) + assert point.status_code == 201, point.text[:300] + + correction = client.post(f"/api/adventures/{adv_id}/state/corrections", json={ + "events": [{"type": "add_fact", "predicate": "keeper", "value": "Mara", + "fact_id": "keeper"}], + "note": "Established in play.", + }) + assert correction.status_code == 201, correction.text[:300] + + import asyncio + asyncio.run(ctx["memorybank"].run_post_turn(adv_id)) + + # Two Undos, so the head is behind the retained tip when M9 opens it. + for _ in range(2): + assert client.post(f"/api/adventures/{adv_id}/undo").status_code == 200 + + report = _describe(ctx, client, adv_id) + ctx["app"].dependency_overrides.clear() + return report + + +# -------------------------------------------------------------------- reading + +def _describe(ctx, client, adv_id: int) -> dict: + """Everything the other build has to agree with, read through the API.""" + models, SessionLocal = ctx["models"], ctx["SessionLocal"] + page = client.get(f"/api/adventures/{adv_id}").json() + with SessionLocal() as db: + adventure = db.get(models.Adventure, adv_id) + version = db.execute(_pragma()).scalar() + counts = { + name: db.query(model).filter(model.adventure_id == adv_id).count() + for name, model in ( + ("actions", models.Action), ("memories", models.Memory), + ("summaries", models.Summary), ("checkpoints", models.Checkpoint), + ("state_events", models.StateEvent), + ("state_proposals", models.StateProposal), + ("knowledge_sources", models.KnowledgeSource), + ("knowledge_chunks", models.KnowledgeChunk), + ) + } + head = {"branch_id": adventure.head_branch_id, "depth": adventure.head_depth} + snapshots = db.query(models.Action).filter( + models.Action.adventure_id == adv_id, + models.Action.context_snapshot.isnot(None), + ).count() + return { + "adventure_id": adv_id, + "schema_version": version, + "title": page["title"], + "canon_rules": page["canon_rules"], + "transcript": [a["text"] for a in page["actions"]], + "can_undo": page["can_undo"], + "can_redo": page["can_redo"], + "head": head, + "counts": counts, + "snapshot_rows": snapshots, + "state": client.get(f"/api/adventures/{adv_id}/state").json()["document"], + "checkpoints": sorted( + (c["name"], c["depth"]) + for c in client.get(f"/api/adventures/{adv_id}/checkpoints").json() + ), + "knowledge": sorted( + (k["original_filename"], k["classification"], k["enabled"], + k["visibility"], k["always_include"], k["content_hash"], + k["index_state"]) + for k in client.get(f"/api/adventures/{adv_id}/knowledge").json() + ), + "events": sorted( + (e["event_type"], e["source"], json.dumps(e["payload"], sort_keys=True)) + for e in client.get( + f"/api/adventures/{adv_id}/state/events?limit=500").json() + ), + } + + +def _comparable(value): + """`value` as it survives a JSON round trip, so the two builds compare like.""" + return json.loads(json.dumps(value, sort_keys=True, default=str)) + + +def _pragma(): + from sqlalchemy import text + + return text("PRAGMA user_version") + + +def open_it(db_path: str, expected: dict) -> dict: + """Opens an existing database with this build, and checks it against `expected`.""" + ctx = _app(db_path) + from app import migrations + + # This is the migration run. `main` already called `bootstrap` at import. + with ctx["engine"].begin() as conn: + after_first = conn.execute(_pragma()).scalar() + # And again, to prove idempotence: a second run must change nothing. + migrations.bootstrap(ctx["engine"]) + with ctx["engine"].begin() as conn: + after_second = conn.execute(_pragma()).scalar() + + adv_id = expected["adventure_id"] + with ctx["SessionLocal"]() as db: + user = db.query(ctx["models"].User).first() + user_id = user.id + client = _client(ctx, user_id) + actual = _describe(ctx, client, adv_id) + + problems = [] + for key in ("title", "canon_rules", "transcript", "head", "counts", + "snapshot_rows", "state", "checkpoints", "knowledge", "events", + "can_undo", "can_redo"): + # Compared through JSON, because that is how the other build's answer + # arrived: a tuple written by `_describe` comes back as a list, and a + # comparison that called that a difference would report ten differences + # in a database nothing had changed. + if _comparable(actual[key]) != _comparable(expected[key]): + problems.append(f"{key}: expected {expected[key]!r}, got {actual[key]!r}") + if expected["schema_version"] != after_first: + problems.append( + f"the schema version moved: {expected['schema_version']} -> {after_first}" + ) + if after_first != after_second: + problems.append( + f"a second open migrated again: {after_first} -> {after_second}" + ) + + # And the campaign still works, rather than merely reading correctly. + exported = client.get(f"/api/adventures/{adv_id}/export") + if exported.status_code != 200: + problems.append(f"export failed: {exported.status_code}") + else: + imported = client.post("/api/adventures/import", json=exported.json()) + if imported.status_code != 201: + problems.append(f"round trip failed: {imported.text[:300]}") + elif exported.json()["format"] != "ai-dnd-adventure-v3": + problems.append("the M8 database did not export as v3") + redo = client.post(f"/api/adventures/{adv_id}/redo") + if redo.status_code != 200: + problems.append(f"Redo failed on the migrated campaign: {redo.status_code}") + + ctx["app"].dependency_overrides.clear() + return { + "schema_version_before": expected["schema_version"], + "schema_version_after": after_first, + "schema_version_second_open": after_second, + "problems": problems, + "checked": { + "families": 12, "snapshot_rows": actual["snapshot_rows"], + "counts": actual["counts"], + }, + } + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("mode", choices=("build", "open")) + parser.add_argument("db_path") + parser.add_argument("--expected", help="the JSON `build` printed (open only)") + args = parser.parse_args() + + if args.mode == "build": + print(json.dumps(build(args.db_path), sort_keys=True)) + return 0 + + expected = json.loads(Path(args.expected).read_text()) + result = open_it(args.db_path, expected) + print(json.dumps(result, indent=2, sort_keys=True)) + return 1 if result["problems"] else 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/backend/tools/m9_portability_report.py b/backend/tools/m9_portability_report.py new file mode 100644 index 0000000..10c81ac --- /dev/null +++ b/backend/tools/m9_portability_report.py @@ -0,0 +1,354 @@ +"""Measures what a campaign bundle preserves, omits and rebuilds. + + python -m tools.m9_portability_report # human-readable + python -m tools.m9_portability_report --json # machine-readable + +Run from `backend/`, with the virtualenv on the path. The script builds the M9 +portability fixture in a throwaway database, exports it, imports it into a second +throwaway database, and then compares the two campaigns family by family. + +It exists because the M9 brief asks for the baseline to be **measured** rather +than assumed. Running it on the M8 commit produces the inventory M9 started from; +running it on the M9 tree produces the one M9 finished with, and the difference +between the two files is the milestone's portability claim in a form a reviewer +can reproduce rather than take on trust. + +The comparison is by data family rather than by row count. "12 actions in, 12 +actions out" is the check that misses a bundle carrying every turn and none of +its state, so each family below reports what a reader could still see afterwards. + +Nothing here touches the developer's own database: two temporary files are +created and removed, and no network call is made — the narrator, the summariser +and the embedder are all local fakes. +""" + +from __future__ import annotations + +import argparse +import json +import os +import sys +import tempfile +import time +from pathlib import Path + +# The test harness owns the fixture and the fakes. Both live under `tests/`, +# which is not a package, so the path is extended rather than imported from. +_HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(_HERE.parent / "tests")) + +# `app.database` reads this at import and builds the engine once, exactly as +# `tests/conftest.py` explains. It has to be set before the first `app` import. +_SOURCE_DB = tempfile.NamedTemporaryFile(suffix="-m9-source.db", delete=False) +_SOURCE_DB.close() +os.environ["AIDND_DB_PATH"] = _SOURCE_DB.name +os.environ.pop("AIDND_DATABASE_URL", None) +os.environ.pop("DATABASE_URL", None) + +from fastapi import Depends # noqa: E402 +from fastapi.testclient import TestClient # noqa: E402 + +import m9_fixture # noqa: E402 +from app import auth, limits, memorybank, models # noqa: E402 +from app.database import Base, SessionLocal, engine, get_db # noqa: E402 +from app.main import app # noqa: E402 +from app.routers import adventures # noqa: E402 +from fakes import ScriptedProvider # noqa: E402 + + +class _StubDerivedProvider: + """Deterministic vectors and prose, so the report needs no model at all. + + One object serves as both the embedder and the summariser, because the + memory pass builds each from the same factory and stubbing only one of them + is the M6 finding M6-F3 mistake: the unstubbed factory opens a socket + against the default endpoint on every turn. + """ + + _written = 0 + + async def complete(self, system, prompt, **kwargs): + _StubDerivedProvider._written += 1 + return ( + f"Memory {_StubDerivedProvider._written}: what the story had " + f"established by this point." + ) + + async def embed(self, texts): + out = [] + for text in texts: + lowered = text.lower() + out.append([ + 1.0, + 1.0 if "abbey" in lowered or "crypt" in lowered else 0.0, + 1.0 if "tavern" in lowered or "lantern" in lowered else 0.0, + 1.0 if "rain" in lowered else 0.0, + ]) + return out + + +def _install_fakes() -> None: + adventures.turns.OpenAICompatibleProvider = ScriptedProvider + memorybank.embedding_provider = lambda s: _StubDerivedProvider() + memorybank.summary_provider = lambda s: _StubDerivedProvider() + limits.check_row_cap = lambda *a, **k: None + + +def _new_user_and_campaign(title: str) -> tuple[int, int]: + db = SessionLocal() + try: + user = models.User(is_guest=False, email=f"m9-{title}@example.com") + db.add(user) + db.flush() + db.add(models.Settings( + user_id=user.id, model="report-model", embedding_model="stub-embed", + context_token_budget=4000, max_output_tokens=400, + )) + adventure = models.Adventure( + user_id=user.id, title=title, + campaign_canon=m9_fixture.CAMPAIGN_CANON, + ) + db.add(adventure) + db.flush() + db.add(models.Action( + adventure_id=adventure.id, type="start", text=m9_fixture.OPENING, + )) + # A neighbour, so a bundle that reached past its own campaign would + # bring back rows this report can see. + neighbour = models.Adventure(user_id=user.id, title="Neighbour") + db.add(neighbour) + db.flush() + db.add(models.Action( + adventure_id=neighbour.id, type="start", text="A different story.", + )) + db.commit() + return adventure.id, user.id + finally: + db.close() + + +def _client(user_id: int) -> TestClient: + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id) + ) + return TestClient(app) + + +# ----------------------------------------------------------------- the families +# One entry per data family the M9 brief asks the baseline to classify. Each +# `present` function answers "did this survive into the file?" from the bundle +# alone, because that is the question the classification is about. + +def _actions(bundle: dict) -> list[dict]: + return [a for a in (bundle.get("actions") or []) if isinstance(a, dict)] + + +def _snapshots(bundle: dict) -> list[dict]: + """Every stored prompt in the file, decoded. + + The export compresses them (`bundle._packed`), so a report that looked for a + plain dict would say the evidence was omitted when it is merely encoded — + which is the mistake this whole tool exists to avoid making about anything. + """ + from app import bundle as bundle_module + + out = [] + for action in _actions(bundle): + snapshot = ( + action.get("contextSnapshot") + if isinstance(action.get("contextSnapshot"), dict) + else bundle_module._unpacked(action.get("contextSnapshotZ")) + ) + if isinstance(snapshot, dict): + out.append(snapshot) + return out + + +FAMILIES: list[tuple[str, str, callable]] = [ + ("campaign identity", + "title, instructions, persona, canon, the campaign's own settings", + lambda b: bool(b.get("title"))), + ("transcript", + "every accepted player and narrator action, live and superseded", + lambda b: bool(_actions(b))), + ("branches", + "the retained tree, its fork points and its names", + lambda b: bool(b.get("branches"))), + ("branch disposition", + "which lines the story left behind, and where", + lambda b: any("supersededAt" in x for x in (b.get("branches") or []))), + ("active head", + "the branch and depth the campaign is being read at", + lambda b: b.get("headDepth") is not None), + ("alternate takes", + "every attempt at a turn, and which one is the story", + lambda b: any(not a.get("live", True) for a in _actions(b))), + ("take grouping", + "which attempts belong to the same turn across a fork (SP9 parentage)", + lambda b: any("parentId" in a for a in _actions(b))), + ("save points", + "named coordinates, their notes and their positions", + lambda b: bool(b.get("checkpoints"))), + ("narrative state (current)", + "the authoritative document at the exported head", + lambda b: b.get("narrativeState") is not None), + ("narrative state (per position)", + "the snapshot every position restores from", + lambda b: any("narrativeStateAfter" in a for a in _actions(b))), + ("state events", + "the accepted typed events: the audit half of the hybrid", + lambda b: bool(b.get("stateEvents"))), + ("state proposals", + "what the model proposed and what the application did about it", + lambda b: bool(b.get("stateProposals"))), + ("manual corrections", + "state the user asserted, distinguishable from state the story did", + lambda b: any(e.get("source") == "manual_correction" + for e in (b.get("stateEvents") or []))), + ("historical prompt/context", + "the exact prompt each turn was given", + lambda b: bool(_snapshots(b))), + ("retrieval provenance", + "which passages a historical turn was shown, and their text", + lambda b: any((s.get("knowledge") or {}).get("used") for s in _snapshots(b))), + ("per-turn model settings", + "the model and generation settings a historical turn ran under", + lambda b: any(s.get("settings") for s in _snapshots(b))), + ("imported knowledge", + "source content, class, lifecycle, visibility and hash", + lambda b: bool(b.get("knowledge"))), + ("knowledge parser versions", + "what produced the chunks the source last had", + lambda b: any("parserVersion" in k for k in (b.get("knowledge") or []))), + ("summaries", + "the generated rolling summaries and the story they cover", + lambda b: bool(b.get("summaries"))), + ("memories", + "long-term memories and the coordinate each hangs off", + lambda b: bool(b.get("memories"))), + ("memory authority", + "whether a memory is accepted story or a heuristic reading of it", + lambda b: any("authority" in m for m in (b.get("memories") or []))), + ("scene metadata", + "the scene section of the authoritative state document", + lambda b: isinstance(b.get("narrativeState"), dict) + and "scene" in b["narrativeState"]), + ("story cards (legacy)", + "the inherited lore primitive, which has no v1 browser surface", + lambda b: "storyCards" in b), +] + +#: Families that are deliberately rebuilt rather than carried, with the reason. +REBUILDABLE = { + "knowledge passages": "a deterministic function of the source content", + "lexical (FTS) index": "rebuilt from the passages on import", + "knowledge embeddings": "belong to the importing machine's embedding model", + "memory embeddings": "the same, for the memory bank", + "branch lineage cache": "computed from parent plus fork depth", + "derived status": "describes the last run of a background pass, not the story", +} + + +def measure(json_out: bool) -> dict: + _install_fakes() + Base.metadata.create_all(bind=engine) + adv_id, user_id = _new_user_and_campaign("M9 Portability Fixture") + client = _client(user_id) + + built = time.perf_counter() + source = m9_fixture.build(client, adv_id) + build_seconds = time.perf_counter() - built + + started = time.perf_counter() + response = client.get(f"/api/adventures/{adv_id}/export") + export_seconds = time.perf_counter() - started + response.raise_for_status() + bundle = response.json() + encoded = json.dumps(bundle, ensure_ascii=False).encode("utf-8") + + started = time.perf_counter() + imported = client.post("/api/adventures/import", json=bundle) + import_seconds = time.perf_counter() - started + import_status = imported.status_code + copy = ( + m9_fixture.snapshot_of(client, imported.json()["id"]) + if import_status == 201 else None + ) + + source_db_bytes = Path(_SOURCE_DB.name).stat().st_size + app.dependency_overrides.clear() + + families = [ + {"family": name, "what": what, + "verdict": "PRESERVED" if present(bundle) else "OMITTED"} + for name, what, present in FAMILIES + ] + report = { + "format": bundle.get("format"), + "families": families, + "rebuildable": REBUILDABLE, + "sizes": { + "source_database_bytes": source_db_bytes, + "bundle_bytes": len(encoded), + "bundle_actions": len(_actions(bundle)), + "bundle_keys": sorted(bundle), + }, + "timings_seconds": { + "fixture_build": round(build_seconds, 3), + "export": round(export_seconds, 3), + "import": round(import_seconds, 3), + }, + "round_trip": { + "import_status": import_status, + "agrees": _agreement(source, copy) if copy else None, + }, + } + return report + + +def _agreement(source: dict, copy: dict) -> dict: + """Which of the reader-visible families match between original and copy.""" + keys = ("title", "canon_rules", "transcript", "branch_count", "checkpoints", + "knowledge", "state", "state_events", "memories", "summaries", + "can_undo", "can_redo") + return {key: source.get(key) == copy.get(key) for key in keys} + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--json", action="store_true", + help="print the report as JSON") + args = parser.parse_args() + try: + report = measure(args.json) + finally: + for path in (_SOURCE_DB.name,): + try: + os.unlink(path) + except OSError: + pass + if args.json: + print(json.dumps(report, indent=2, sort_keys=True)) + return 0 + print(f"bundle format: {report['format']}") + print(f"bundle size: {report['sizes']['bundle_bytes']:,} bytes " + f"across {report['sizes']['bundle_actions']} actions") + print(f"source db: {report['sizes']['source_database_bytes']:,} bytes") + print(f"timings: {report['timings_seconds']}") + print() + width = max(len(name) for name, _, _ in FAMILIES) + for row in report["families"]: + print(f" {row['verdict']:<10} {row['family']:<{width}} {row['what']}") + print() + print(" DERIVED/REBUILDABLE (deliberately not carried)") + for name, why in REBUILDABLE.items(): + print(f" {name:<24} {why}") + print() + print(f"round trip: HTTP {report['round_trip']['import_status']}") + for key, agreed in (report["round_trip"]["agrees"] or {}).items(): + print(f" {'same' if agreed else 'DIFFERS':<8} {key}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/backend/tools/m9_scale_report.py b/backend/tools/m9_scale_report.py new file mode 100644 index 0000000..bed16dd --- /dev/null +++ b/backend/tools/m9_scale_report.py @@ -0,0 +1,264 @@ +"""How a campaign bundle grows with the campaign, measured rather than reasoned. + + python -m tools.m9_scale_report [--turns 120] [--budget 16384] + +Run from `backend/`. Plays a campaign of `--turns` turns against a scripted +narrator with the real prompt builder and a realistic context budget, then +exports it and reports where the bytes are. + +The question it exists to answer is the one M9's decision to carry historical +prompts raises: **a per-turn prompt contains the story so far, so storing one per +turn is quadratic in campaign length.** That is already true of the database — +`compression.py` records the column as 89% of production storage — and M9 makes +it true of the export as well. Reasoning about it gives the wrong number, because +the prompt is bounded by the context budget rather than by the transcript: once +the history window is full, each turn's snapshot stops growing and the total +becomes linear again. Where that knee falls is a measurement. + +It also watches for the accidental costs §26 names: a query per row, a +duplicated body of knowledge content, or a snapshot written more than once per +turn. +""" + +from __future__ import annotations + +import argparse +import json +import os +import sys +import tempfile +import time +from pathlib import Path + +_HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(_HERE.parent / "tests")) + +_DB = tempfile.NamedTemporaryFile(suffix="-m9-scale.db", delete=False) +_DB.close() +os.environ["AIDND_DB_PATH"] = _DB.name +os.environ.pop("AIDND_DATABASE_URL", None) +os.environ.pop("DATABASE_URL", None) + +from fastapi import Depends # noqa: E402 +from fastapi.testclient import TestClient # noqa: E402 + +import m9_fixture # noqa: E402 +from app import auth, limits, memorybank, models # noqa: E402 +from app.database import Base, SessionLocal, engine, get_db # noqa: E402 +from app.main import app # noqa: E402 +from app.routers import adventures # noqa: E402 +from fakes import ScriptedProvider, state_block # noqa: E402 + +#: Prose long enough that a turn is a turn rather than a sentence. The history +#: window is what fills the prompt, so a fixture of three-word replies would +#: measure a campaign nobody plays. +PROSE = ( + "The rain came harder off the fen and the lantern light shivered on the wet " + "boards. Mara set down the cloth she had been folding and looked at him for " + "a while without saying anything, the way she did when the answer was going " + "to cost her something. Outside, somebody crossed the yard and did not stop." +) + + +class _Stub: + async def complete(self, system, prompt, **kwargs): + return "The story had established a good deal by this point." + + async def embed(self, texts): + return [[1.0, 0.5, 0.25, 0.125] for _ in texts] + + +def _setup(budget: int) -> tuple[TestClient, int]: + adventures.turns.OpenAICompatibleProvider = ScriptedProvider + memorybank.embedding_provider = lambda s: _Stub() + memorybank.summary_provider = lambda s: _Stub() + limits.check_row_cap = lambda *a, **k: None + Base.metadata.create_all(bind=engine) + db = SessionLocal() + try: + user = models.User(is_guest=False, email="scale@example.com") + db.add(user) + db.flush() + db.add(models.Settings( + user_id=user.id, model="scale-model", embedding_model="stub", + context_token_budget=budget, max_output_tokens=800, + )) + adventure = models.Adventure( + user_id=user.id, title="Scale", auto_summarize=True, + memory_bank_enabled=True, + campaign_canon=m9_fixture.CAMPAIGN_CANON, + ) + db.add(adventure) + db.flush() + db.add(models.Action( + adventure_id=adventure.id, type="start", text=m9_fixture.OPENING, + )) + db.commit() + adv_id, user_id = adventure.id, user.id + finally: + db.close() + app.dependency_overrides[auth.get_current_user] = ( + lambda db=Depends(get_db): db.get(models.User, user_id) + ) + return TestClient(app), adv_id + + +def _bundle_bytes(client, adv_id) -> tuple[int, dict, float]: + started = time.perf_counter() + response = client.get(f"/api/adventures/{adv_id}/export") + seconds = time.perf_counter() - started + response.raise_for_status() + payload = response.json() + return len(json.dumps(payload).encode("utf-8")), payload, seconds + + +def _snapshot_bytes(payload: dict) -> int: + """What the stored prompts cost **in the file**, which is the encoded size. + + Measured as they appear rather than decoded first: the question this report + answers is how large the file gets and how close it comes to the import + ceiling, so what counts is the bytes that actually travel. + """ + return sum( + len(json.dumps(action[key]).encode("utf-8")) + for action in payload["actions"] + for key in ("contextSnapshotZ", "contextSnapshot") + if action.get(key) + ) + + +def _decoded_snapshot_bytes(payload: dict) -> int: + """What the same prompts would cost uncompressed, for the ratio.""" + from app import bundle as bundle_module + + total = 0 + for action in payload["actions"]: + snapshot = ( + action.get("contextSnapshot") + if isinstance(action.get("contextSnapshot"), dict) + else bundle_module._unpacked(action.get("contextSnapshotZ")) + ) + if isinstance(snapshot, dict): + total += len(json.dumps(snapshot).encode("utf-8")) + return total + + +def _section_bytes(payload: dict) -> dict: + """What each v3 addition costs in the file, separately. + + Needed because "the snapshots are 21% of the file" does not answer "what did + M9 add": the state events, the proposals and the summaries are v3 additions + too, and a claim about M9's cost that counted only the prompts would be + understating it. + """ + def size(value) -> int: + return len(json.dumps(value).encode("utf-8")) + + per_node = {"contextSnapshotZ": 0, "id": 0, "parentId": 0} + for action in payload["actions"]: + for key in per_node: + if key in action: + per_node[key] += size(action[key]) + len(key) + 4 + return { + "prompts (contextSnapshotZ)": per_node["contextSnapshotZ"], + "state events": size(payload.get("stateEvents") or []), + "state proposals": size(payload.get("stateProposals") or []), + "summaries": size(payload.get("summaries") or []), + "node ids + parentage": per_node["id"] + per_node["parentId"], + "per-position state (v2 already)": sum( + size(a["narrativeStateAfter"]) for a in payload["actions"] + if "narrativeStateAfter" in a + ), + } + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--turns", type=int, default=120) + parser.add_argument("--budget", type=int, default=16384) + parser.add_argument("--every", type=int, default=20, + help="report the running size every N turns") + args = parser.parse_args() + + client, adv_id = _setup(args.budget) + for name, body, kind in ( + ("canon.md", m9_fixture.CANON_MD, "canon"), + ("reference.md", m9_fixture.REFERENCE_MD, "reference"), + ("secret.md", m9_fixture.SECRET_MD, "canon"), + ): + m9_fixture.upload(client, adv_id, name, body, kind) + + print(f"budget {args.budget} tokens, {args.turns} turns\n") + print(f"{'turns':>6} {'actions':>8} {'bundle B':>12} {'snapshots B':>13} " + f"{'B/turn':>9} {'export s':>9} {'import s':>9}") + rows = [] + sections: dict[int, dict] = {} + for turn in range(1, args.turns + 1): + ScriptedProvider.replies = [ + f"{PROSE} [{turn}]\n" + + state_block([{"type": "add_fact", "predicate": "tally", + "value": turn * 10, "fact_id": f"tally-{turn}"}]) + ] + response = client.post(f"/api/adventures/{adv_id}/actions", + json={"type": "do", "text": f"press on, {turn}"}) + assert response.status_code == 200, response.text[:300] + if turn % args.every and turn != args.turns: + continue + m9_fixture.settle_derived(adv_id) + size, payload, export_seconds = _bundle_bytes(client, adv_id) + started = time.perf_counter() + imported = client.post("/api/adventures/import", json=payload) + import_seconds = time.perf_counter() - started + assert imported.status_code == 201, imported.text[:300] + client.delete(f"/api/adventures/{imported.json()['id']}") + snapshots = _snapshot_bytes(payload) + plain = _decoded_snapshot_bytes(payload) + rows.append((turn, size, snapshots, plain)) + sections[turn] = _section_bytes(payload) + print(f"{turn:>6} {len(payload['actions']):>8} {size:>12,} " + f"{snapshots:>13,} {snapshots // turn:>9,} " + f"{export_seconds:>9.3f} {import_seconds:>9.3f}") + + db_bytes = Path(_DB.name).stat().st_size + last_turn, last_size, last_snapshots, last_plain = rows[-1] + per_turn = last_snapshots // last_turn + cap = limits.MAX_IMPORT_BODY_BYTES + print() + print(f"database on disk: {db_bytes:,} bytes") + print(f"snapshot share of file: {100 * last_snapshots // last_size}%") + print(f"stored uncompressed: {last_plain:,} bytes " + f"({last_plain / max(last_snapshots, 1):.1f}x the encoded size)") + print(f"import body cap: {cap:,} bytes") + print(f"turns before the cap: ~{cap // max(per_turn, 1):,} " + f"at the marginal rate above") + # Growth between the last two samples says whether the per-turn cost has + # settled. It should: once the history window fills the budget, a prompt + # stops growing with the transcript and the total becomes linear. + if len(rows) >= 2: + (t0, _, s0, _p0), (t1, _, s1, _p1) = rows[-2], rows[-1] + print(f"marginal cost, last {t1 - t0} turns: " + f"{(s1 - s0) // max(t1 - t0, 1):,} bytes/turn") + print() + print("where the bytes are, at the last sample:") + last = sections[last_turn] + added = sum(v for k, v in last.items() if not k.endswith("(v2 already)")) + for name, value in sorted(last.items(), key=lambda kv: -kv[1]): + print(f" {name:34} {value:>12,} {100 * value / last_size:5.1f}%") + print(f" {'--- everything v3 added':34} {added:>12,} " + f"{100 * added / last_size:5.1f}%") + without = last_size - added + print(f" a v2 file of the same campaign {without:>12,}") + print(f" ceiling with v3 additions: ~{int((cap / (last_size / last_turn**2)) ** 0.5):,} turns") + print(f" ceiling without them: ~{int((cap / (without / last_turn**2)) ** 0.5):,} turns") + app.dependency_overrides.clear() + return 0 + + +if __name__ == "__main__": + try: + raise SystemExit(main()) + finally: + try: + os.unlink(_DB.name) + except OSError: + pass diff --git a/frontend/src/api.js b/frontend/src/api.js index 9ba41e4..235a893 100644 --- a/frontend/src/api.js +++ b/frontend/src/api.js @@ -153,6 +153,14 @@ export const api = { retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal), exportAdventure: (id) => request(`/adventures/${id}/export`), importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }), + + // M9. A verified copy of the whole database, which is a different tool from + // exporting one campaign: the export moves a campaign between installations, + // and this is a safety copy of everything on this machine. Neither takes a + // path — the server derives the destination from the database it already has + // open, so there is nothing here for a caller to point somewhere else. + listBackups: () => request('/backups'), + createBackup: () => request('/backups', { method: 'POST' }), undo: (advId) => request(`/adventures/${advId}/undo`, { method: 'POST' }), // Undo moves the story back without deleting it, so there is somewhere to // move forward to again (M3). Both answer with the newest window. diff --git a/frontend/src/components.jsx b/frontend/src/components.jsx index 50dc126..92e15a0 100644 --- a/frontend/src/components.jsx +++ b/frontend/src/components.jsx @@ -15,17 +15,53 @@ export function pickJSONFile() { const input = document.createElement('input') input.type = 'file' input.accept = '.json,application/json' + // M9: in the document, and driveable, rather than detached. + // + // A detached input is what this was, and `.click()` on one opens the + // browser's file dialog in Firefox and Chrome today — but it is not + // something the HTML spec requires, and it made the one control that + // recovers a campaign impossible to drive from a browser test: there is no + // element for WebDriver to hand a path to, so the import workflow could + // only ever be checked by calling the API underneath it. + // + // Taken out of layout rather than marked `hidden`, and the difference is + // load-bearing. A `hidden` input is non-interactable, and WebDriver will + // set `files` on one without dispatching `change` — so the file lands and + // nothing happens, which is a worse failure than the detached input was + // because it looks like it worked. This is the ordinary visually-hidden + // file-input pattern: off-screen, zero-sized, out of the accessibility + // tree and out of the tab order, so no reader meets a stray "Choose file" + // control while the browser's own dialog is what they are looking at. + input.setAttribute('aria-hidden', 'true') + input.tabIndex = -1 + input.style.cssText = + 'position:fixed;left:-9999px;width:1px;height:1px;opacity:0;pointer-events:none' + input.dataset.testid = 'import-file' + const done = (settle) => (value) => { input.remove(); settle(value) } + const ok = done(resolve) + const bad = done(reject) + // M9: a cancelled dialog settles the promise. + // + // It did not before. `onchange` does not fire when the reader closes the + // picker without choosing anything, so the promise stayed pending forever + // — and the Campaigns screen awaits it, so its `finally` never ran and the + // Import button sat disabled reading "Importing…" until the page was + // reloaded. The rejection carries an empty message, because that screen + // already treats a message-less error as "they changed their mind" and + // says nothing: a cancelled dialog is not a failure to report. + input.oncancel = () => bad(Object.assign(new Error(), { message: '' })) input.onchange = () => { const file = input.files[0] - if (!file) return reject(new Error('No file selected')) + if (!file) return bad(new Error('No file selected')) const reader = new FileReader() reader.onload = () => { - try { resolve(JSON.parse(reader.result)) } - catch { reject(new Error('Not valid JSON')) } + try { ok(JSON.parse(reader.result)) } + catch { bad(new Error('Not valid JSON')) } } - reader.onerror = () => reject(new Error('Could not read file')) + reader.onerror = () => bad(new Error('Could not read file')) reader.readAsText(file) } + document.body.appendChild(input) input.click() }) } diff --git a/frontend/src/filePicker.test.jsx b/frontend/src/filePicker.test.jsx new file mode 100644 index 0000000..0ad490e --- /dev/null +++ b/frontend/src/filePicker.test.jsx @@ -0,0 +1,110 @@ +/* M9: the file picker that recovers a campaign. + * + * `pickJSONFile` is four lines of DOM and was the only control in the product + * with no test at all, for a structural reason: it built a detached + * `` and clicked it, so there was no element for a test — or + * for WebDriver — to hand a file to. The import workflow could therefore only + * ever be checked by calling the API underneath it, which is not the workflow. + * + * Appending the input made it testable, and writing the test found a real bug + * that had been there since the picker was written: closing the dialog without + * choosing anything never settled the promise, so the Campaigns screen's + * `finally` never ran and its Import button stayed disabled reading + * "Importing…" until the page was reloaded. The screen's own comment says a + * cancelled picker is not worth a message — it had just never received one. + */ + +import { fireEvent } from '@testing-library/react' +import { beforeEach, describe, expect, it } from 'vitest' +import { pickJSONFile } from './components' + +function theInput() { + return document.querySelector('input[type="file"]') +} + +/** A `File` the way the browser hands one to a change event. */ +function jsonFile(name, contents) { + return new File([JSON.stringify(contents)], name, { type: 'application/json' }) +} + +/** Puts `files` on the input, since `files` is read-only in jsdom. */ +function choose(input, files) { + Object.defineProperty(input, 'files', { value: files, configurable: true }) + fireEvent.change(input) +} + +beforeEach(() => { document.body.innerHTML = '' }) + +describe('pickJSONFile', () => { + it('puts a findable input in the document rather than a detached one', () => { + pickJSONFile().catch(() => {}) + const input = theInput() + expect(input).toBeInTheDocument() + expect(input.dataset.testid).toBe('import-file') + expect(input.accept).toContain('json') + // Out of sight, out of the tab order and out of the accessibility tree, + // because the reader is looking at the browser's own dialog — but **not** + // `hidden`, which would make it non-interactable and stop the browser + // dispatching `change` when a file is chosen programmatically. + expect(input.hidden).toBe(false) + expect(input.getAttribute('aria-hidden')).toBe('true') + expect(input.tabIndex).toBe(-1) + expect(input.style.position).toBe('fixed') + }) + + it('resolves with the parsed bundle', async () => { + const promise = pickJSONFile() + choose(theInput(), [jsonFile('c.json', { format: 'ai-dnd-adventure-v3' })]) + await expect(promise).resolves.toEqual({ format: 'ai-dnd-adventure-v3' }) + }) + + it('rejects a file that is not JSON, with a message worth showing', async () => { + const promise = pickJSONFile() + const input = theInput() + Object.defineProperty(input, 'files', { + value: [new File(['not json at all'], 'c.json')], configurable: true, + }) + fireEvent.change(input) + await expect(promise).rejects.toThrow(/not valid json/i) + }) + + it('settles even when the file cannot be read, rather than hanging', async () => { + // A real case, not a defensive one. A browser can hand the page a `File` + // whose contents it will not then let the page read — a sandboxed Firefox + // does exactly that for a path outside its confinement, and reports + // `NotFoundError` from the FileReader with the name and size intact. + // Whatever happens, the promise must settle: leaving it pending is what + // left the Import button disabled reading "Importing…". + const promise = pickJSONFile() + const input = theInput() + Object.defineProperty(input, 'files', { + value: [jsonFile('c.json', { ok: true })], configurable: true, + }) + fireEvent.change(input) + await expect(Promise.race([ + promise.then(() => 'settled', () => 'settled'), + new Promise((r) => { setTimeout(() => r('hung'), 300) }), + ])).resolves.toBe('settled') + }) + + it('settles when the dialog is cancelled, instead of hanging forever', async () => { + const promise = pickJSONFile() + fireEvent(theInput(), new Event('cancel')) + // Rejected, so the caller's `finally` runs — and with no message, so the + // caller shows nothing. Both halves matter: a hang leaves the button + // disabled, and a message would report a decision as a failure. + await expect(promise).rejects.toSatisfy((err) => err.message === '') + }) + + it('takes the input back out of the document however it settles', async () => { + const resolved = pickJSONFile() + choose(theInput(), [jsonFile('c.json', { ok: true })]) + await resolved + expect(theInput()).toBeNull() + + const cancelled = pickJSONFile() + fireEvent(theInput(), new Event('cancel')) + await cancelled.catch(() => {}) + expect(theInput()).toBeNull() + }) +}) diff --git a/frontend/src/pages/Settings.jsx b/frontend/src/pages/Settings.jsx index 132e270..17abcce 100644 --- a/frontend/src/pages/Settings.jsx +++ b/frontend/src/pages/Settings.jsx @@ -85,6 +85,105 @@ function DebugLog() { ) } + +/* M9. A verified copy of the whole database, on this machine. + * + * Deliberately small, and deliberately not a second export. The campaign export + * on the Campaigns screen is the tool for moving one campaign to another + * installation; this is the tool for keeping a copy of everything before doing + * something risky. Conflating them would leave a reader guessing which one + * answers "how do I not lose my campaigns". + * + * There is no restore button and no download link, and both absences are + * decisions rather than gaps: + * + * * **Restore** means replacing the file the running application has open, + * which is how somebody loses both copies at once. `DEVELOPMENT.md` carries + * the procedure — stop the app, move the file, start it — and it is a + * procedure precisely because each step needs the app to be stopped. + * * **Download** would put a copy of every campaign on the machine into the + * browser's download directory and its cache. For an application whose + * premise is that the story does not leave the machine, a path the reader + * can copy is the better default. + */ +function DatabaseBackup() { + const toast = useToast() + const [state, setState] = useState(null) + const [busy, setBusy] = useState(false) + + const load = () => { + api.listBackups() + .then(setState) + .catch((err) => toast(err.message, 'error')) + } + + const take = async () => { + setBusy(true) + try { + const result = await api.createBackup() + toast(`Backup written: ${result.filename} (${formatBytes(result.bytes)}).`) + load() + } catch (err) { + toast(err.message, 'error') + } finally { + setBusy(false) + } + } + + return ( +
{ if (e.currentTarget.open && state === null) load() }} + > + Back up everything on this machine +

+ Writes a verified copy of the whole database — every campaign, every + imported file, every setting — beside the database itself. The copy is + checked before it is kept, and an existing backup is never overwritten. + To export a single campaign so it can be opened somewhere else, use + Export on the campaign instead. +

+
+ +
+ {state && ( + <> +

+ Backups are written to {state.directory}. To restore + one, stop the application, put the file in place of the database, and + start it again. +

+ {state.backups.length === 0 ? ( +

No backups yet.

+ ) : ( +
    + {state.backups.map((b) => ( +
  • + {b.filename} + + {' '}{formatBytes(b.bytes)} · {new Date(b.taken_at).toLocaleString()} + +
  • + ))} +
+ )} + + )} +
+ ) +} + +function formatBytes(bytes) { + if (!Number.isFinite(bytes)) return '' + if (bytes < 1024) return `${bytes} B` + if (bytes < 1024 * 1024) return `${Math.round(bytes / 1024)} kB` + return `${(bytes / (1024 * 1024)).toFixed(1)} MB` +} + + export default function Settings() { const [settings, setSettings] = useState(null) const [testResult, setTestResult] = useState(null) @@ -347,6 +446,8 @@ export default function Settings() { + + diff --git a/frontend/src/pages/backup.test.jsx b/frontend/src/pages/backup.test.jsx new file mode 100644 index 0000000..338454d --- /dev/null +++ b/frontend/src/pages/backup.test.jsx @@ -0,0 +1,116 @@ +/* M9: the database backup control, and what it must not become. + * + * The control itself is three lines of state, so the interesting assertions are + * about the boundaries around it rather than about the button: + * + * * it is **not** an export. The campaign export moves one campaign to + * another installation; this copies everything on this machine. A reader + * who cannot tell them apart has no way to answer "how do I not lose my + * campaigns", so the panel says which is which. + * * it offers **no restore and no download**, and both are decisions. Restore + * means replacing the file the running application has open; download means + * putting every campaign on the machine into the browser's cache. + * * a failure **says so**. A backup that silently did not happen is worse + * than no backup, because the reader believes they have one. + */ + +import { screen, waitFor } from '@testing-library/react' +import userEvent from '@testing-library/user-event' +import { beforeEach, describe, expect, it, vi } from 'vitest' +import { api } from '../api' +import Settings from './Settings' +import { mockModelStatus, renderWith } from '../test/helpers' + +const LISTING = { + directory: '/home/reader/.adventure/backups', + backups: [ + { filename: 'adventure-storyteller-20260907-043000.db', bytes: 2_400_000, + taken_at: '2026-09-07T04:30:00' }, + { filename: 'adventure-storyteller-20260901-101500.db', bytes: 2_100_000, + taken_at: '2026-09-01T10:15:00' }, + ], +} + +beforeEach(() => { vi.restoreAllMocks() }) + +async function openTheBackupPanel() { + mockModelStatus(api) + vi.spyOn(api, 'listBackups').mockResolvedValue(LISTING) + await renderWith() + const panel = screen.getByTestId('database-backup') + await userEvent.click(screen.getByText(/Back up everything on this machine/i)) + return panel +} + +describe('the backup control', () => { + it('lists the backups already on disk, and where they are', async () => { + await openTheBackupPanel() + await waitFor(() => { + expect(screen.getByText('adventure-storyteller-20260907-043000.db')).toBeInTheDocument() + }) + expect(screen.getByText('adventure-storyteller-20260901-101500.db')).toBeInTheDocument() + expect(screen.getByText('/home/reader/.adventure/backups')).toBeInTheDocument() + }) + + it('takes a backup and reports what was written', async () => { + const panel = await openTheBackupPanel() + const create = vi.spyOn(api, 'createBackup').mockResolvedValue({ + directory: LISTING.directory, + filename: 'adventure-storyteller-20260907-050000.db', + bytes: 2_500_000, pages: 610, seconds: 0.02, integrity: 'ok', + }) + await userEvent.click(screen.getByRole('button', { name: /Back up now/i })) + await waitFor(() => expect(create).toHaveBeenCalled()) + expect(await screen.findByText(/Backup written/)).toBeInTheDocument() + expect(screen.getByText(/adventure-storyteller-20260907-050000\.db/)).toBeInTheDocument() + expect(panel).toBeInTheDocument() + }) + + it('reloads the list afterwards, so the new file is visible', async () => { + await openTheBackupPanel() + vi.spyOn(api, 'createBackup').mockResolvedValue({ + directory: LISTING.directory, filename: 'new.db', bytes: 1, integrity: 'ok', + }) + await userEvent.click(screen.getByRole('button', { name: /Back up now/i })) + await waitFor(() => expect(api.listBackups).toHaveBeenCalledTimes(2)) + }) + + it('reports a failure rather than looking as though it worked', async () => { + await openTheBackupPanel() + vi.spyOn(api, 'createBackup').mockRejectedValue( + new Error('The backup was written but did not verify: page 4 missing.'), + ) + await userEvent.click(screen.getByRole('button', { name: /Back up now/i })) + expect(await screen.findByText(/did not verify/)).toBeInTheDocument() + expect(screen.queryByText(/Backup written/)).not.toBeInTheDocument() + }) + + it('says which tool this is, and which one moves one campaign', async () => { + await openTheBackupPanel() + expect(screen.getByText(/every campaign, every/i)).toBeInTheDocument() + expect( + screen.getByText(/To export a single campaign so it can be opened somewhere else/i), + ).toBeInTheDocument() + }) + + it('offers no restore button and no download link', async () => { + const panel = await openTheBackupPanel() + await waitFor(() => { + expect(screen.getByText(LISTING.directory)).toBeInTheDocument() + }) + const labels = [...panel.querySelectorAll('button')].map((b) => b.textContent) + expect(labels.some((label) => /restore/i.test(label))).toBe(false) + expect(labels.some((label) => /download/i.test(label))).toBe(false) + expect(panel.querySelectorAll('a[download]')).toHaveLength(0) + expect(panel.querySelectorAll('a[href]')).toHaveLength(0) + // And the procedure is stated instead, so the absence is an answer. + expect(screen.getByText(/stop the application/i)).toBeInTheDocument() + }) + + it('does not fetch anything until the panel is opened', async () => { + mockModelStatus(api) + const list = vi.spyOn(api, 'listBackups').mockResolvedValue(LISTING) + await renderWith() + expect(list).not.toHaveBeenCalled() + }) +}) diff --git a/frontend/src/styles/library.css b/frontend/src/styles/library.css index cd70963..92150bb 100644 --- a/frontend/src/styles/library.css +++ b/frontend/src/styles/library.css @@ -204,6 +204,21 @@ } .advanced-block > summary:hover { color: var(--text); } .advanced-block h4 { margin: 10px 0 4px; font-size: 0.74rem; color: var(--text-dim); } +/* M9: the list of database backups already on disk. A filename, a size and a + time — enough to recognise one, and nothing that needs a table. */ +.backup-list { + margin: 8px 0 0; + padding: 0; + list-style: none; + font-size: 0.78rem; +} +.backup-list li { + padding: 3px 0; + border-top: 1px solid var(--border); + /* A long filename wraps rather than pushing the panel sideways. */ + overflow-wrap: anywhere; +} +.backup-list li:first-child { border-top: none; } .advanced-block pre { max-height: 16em; overflow: auto; diff --git a/planning/BROWSER-UX-SPEC.md b/planning/BROWSER-UX-SPEC.md index 5bbf3f4..81ade25 100644 --- a/planning/BROWSER-UX-SPEC.md +++ b/planning/BROWSER-UX-SPEC.md @@ -128,6 +128,35 @@ The current endpoint should be clear. The input box always continues from the currently active story head. +### 8A. The reader must be able to tell where they are (recorded 2026-09-07) + +A hands-on session against accepted M8 found the sentence above **too weak to +hold the behaviour it names**. Undo worked correctly, the input box did continue +from the active head — so both statements above were satisfied — and the reader +still could not tell which point in the story they had moved to. + +The requirement, stated so that a working implementation cannot satisfy it while +a reader is lost: + +> After Undo, Redo, a Save Point restore, an edit to an earlier turn, or any +> other movement of the active story position, the reader should be able to +> identify **where they now are in the visible story** — and, where it matters, +> whether later story remains available ahead of them. + +Two constraints on any solution: + +- **No implementation terminology.** `branch`, `fork`, `node`, `head` and + `depth` stay off the reader-facing surface (§38), which is what makes this a + presentation problem rather than a labelling one. +- **It must be observable**, not merely inferable from the transcript scrolling, + so that a release test can decide it. + +**No wording is prescribed here, and none is ratified.** A lightweight named +position — `Moment 8` becoming `Moment 7` after an Undo, optionally noting that +later story is available — is one candidate among others. Ownership is M11 +release polish; `V1-ACCEPTANCE-TESTS.md` §P1 records what must be settled before +this can become an acceptance test. + ## 9. User Turn Presentation User messages should support: diff --git a/planning/BUILD-MILESTONES.md b/planning/BUILD-MILESTONES.md index dcaddd0..0e80daa 100644 --- a/planning/BUILD-MILESTONES.md +++ b/planning/BUILD-MILESTONES.md @@ -1118,6 +1118,187 @@ Make campaigns portable and recoverable without losing lineage, state, knowledge A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact. +## Status: COMPLETE — 2026-09-07, pending independent review + +Implemented on `m9-recovery` from the signed M8 commit `1ce9972`, measured +before and after against the same fixture, and verified in a real browser +against a real narrator. `planning/reports/M9-IMPLEMENTATION-REPORT.md` is the +implementer's account, written for a reviewer. + +**What it delivered, beyond the scope list above:** + +- **The bundle became a version, and the reason is a rule.** `ai-dnd-adventure-v3`. + Everything M9 adds could have been an optional key, the way four earlier + additions were — and that mechanism fails exactly here, because a v2 file with + no prompt provenance is ambiguous between "written before M9" and "written by + M9 from a campaign with none". A version is how a recovery file states what it + was capable of recording. v1 and v2 are still read, and every seam from + pre-active-head onward is tested. +- **A third category of data.** "Chosen travels, derived is recomputed" was + enough until stored prompts had to be decided. They are derived and must + travel, so the rule is now chosen / **evidence** / rebuildable, and the test + separating the last two is not "could this be recomputed" but "would a + recomputation answer the same question". +- **The M8 handoff on prompt provenance is closed.** An old turn in a restored + campaign shows what it was actually given, after the source has been deleted, + the canon edited and the state moved on. +- **State events and proposals travel**, so a moved campaign can still say why + its state is what it is — and a manual correction is still identifiable as + one, which it was not before. +- **Summaries travel with their coordinates**, so an abandoned line's summary is + still ineligible after the move, and a moved campaign resumes with its + long-story continuity instead of behaving like a new one. +- **A verified SQLite backup**, using the online backup API rather than a file + copy, taken while the application is running, with a browser control in + Settings. +- **Take parentage**, memory authority and parser/chunking versions, each + closing a smaller fidelity loss. + +**Three defects found by running the milestone's own tests, and fixed here:** + +1. **Deleting a campaign leaked its FTS index rows**, and SQLite then handed the + freed ids to the *next* source imported into *any* campaign, which failed + with an integrity error. Reindex could not repair it either. Pre-existing + since M7; both ends are now closed and an already-damaged database repairs + itself with no migration. +2. **An imported node with no state snapshot was stamped with the campaign's + head state**, so an Undo to turn 2 showed what the story knew at turn 20 — + M5's review finding 3, arriving through the import. +3. **A snapshot's `source_id` was not being translated on import** because the + relink mutated a dict in place, which a non-`MutableDict` column does not + notice. Found by a test asserting the outcome rather than the call. + +**One decision the brief asked for, made and recorded:** + +**Story cards are compatibility-only legacy data, and no longer enter the +narrator's prompt.** They still travel in both directions, the rows and the API +stay, and `memorybank.cast_brief` still reads them as the summariser's character +roster. What stops is the injection: a keyword-matched card arrived in front of +the narrator as `World Lore: …` with no class, no visibility, no source, no way +to switch it off since M8 removed the editor, and no row in the context +inspector — which is `IMPORTED-KNOWLEDGE-DESIGN.md` §73's "alternate untracked +path around the new knowledge authority/provenance rules" in as many words. + +**Debt carried forward, deliberately:** + +- **A long campaign's bundle has a measured ceiling: ~279 turns** against the + 20 MB import limit. A per-turn prompt contains the story so far, so carrying + one per turn is O(turns²); compressing them inside the file cut that to about + an eighth of what it would have been. **M9's additions account for only 12% of + the ceiling.** The other 88% is the per-position narrative state document, + which is 74% of a bundle and which v2 already carried — so lifting the ceiling + means addressing that, not the evidence. Far beyond M11's 100-turn + certification (14% of the cap), and stated with its measurement rather than + hidden. A streaming or chunked import is the fix if a later milestone needs + one; the asymmetry to know about is that such a campaign can still be exported + and would be refused on import. +- **No discarded-history recovery screen** (§63). M9's job was that retained + history survives correctly so a later screen can use it; it does. +- **No whole-transcript copy and no story search** (§77, §78) — still M11 or + later. + +--- + +# Post-M8 hands-on playtest findings — recorded 2026-09-07, owned by M11 + +**Not M9 defects, not caused by M9, and they did not block M9 acceptance.** +Recorded here rather than only in the M9 report because a milestone report is +archived when the next one replaces it, and these must not go with it. + +They come from a real play session against **accepted, signed M8**: real +browser, real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, disposable +isolated campaign database. **That database was deliberately destroyed +afterwards**, so the stored context snapshot for finding D is gone and no root +cause is claimed for it. Full write-up, with the verification behind each +mechanism, is in the M9 report's **§Y** (in `archive/milestone-reports/` once +M10's report replaces it). + +**These are M11's, and explicitly not M10's.** M10 is media-readiness +architecture and stays bounded; it inherits them as known carry-forward only. + +### A. The browser still calls the product "AI D&D" — M11 release polish + +`frontend/index.html` still carries the inherited `AI D&D`. +M8 changed the navigation, the inspector and the screens, and never claimed the +title — so this is an uncovered gap rather than a false claim. + +**Do not fix it with a find-and-replace to "Adventure Storyteller".** +`SPECIFICATION.md` requires a genre-agnostic engine, and *Adventure* is narrower +than the product. The naming decision is the repository owner's; a neutral +working name such as **Interactive Story** and a tab form such as +` — Interactive Story` are candidates, not decisions. + +### B. After Undo, the reader cannot tell where they are — M11 UX polish + +Undo behaved correctly (M3 semantics; re-verified throughout M9). The reader +could not tell **which point in the story** they had moved to. + +`BROWSER-UX-SPEC.md` §8 said only *"The current endpoint should be clear"*, and +its companion sentence was already true while the reader was lost — so the +requirement could not hold the behaviour. §8 has been strengthened to state the +orientation requirement; **no UI text is prescribed**, because none is ratified. +A `Moment 8` → `Moment 7` style indicator is a candidate. Needs a browser +regression scenario. + +### C. Narration length has no measurable effect — M11 realistic-model behaviour + +The setup choice becomes **one English sentence** in the campaign's +`ai_instructions` and changes **no generation setting**. Independently, +`length_hint()` derives a numeric word range from the **global** +`Settings.max_output_tokens` and places it after the history — and at the default +800 it reads *"must not exceed 506 words, and it should not stop short of about +177"* **identically for brief, medium and long**. + +That is a mechanism, verified by reading and running the code — **not a proven +cause** of what the reader saw; the narrator's instruction following is also in +play. A reproduction must measure what enters the stored prompt, whether the +setting moves any generation budget, and actual word/paragraph counts across +repeated turns, on the reference 3B narrator **and** a stronger local one. +Clearer numeric targets (`Brief ~100-200 words` and so on) are a design +candidate, not ratified. **Do not hard-truncate prose** — the state block is +emitted last and truncation removes it. + +### D. Character identity / coreference confusion — M11 diagnostic + +Four people in one scene — Bill (protagonist), Roger, John, Alice — and later +narration treated Alice as two different Alices. + +**Root cause UNKNOWN and no longer establishable.** Candidates: a model +coreference failure on a correct prompt; duplicate/conflicting state; a +context/summary/memory assembly failure; or a context that is not contradictory +but too implicit for a small model. + +**One structural fact to check first**, verified by reading the code: the +narrative state **permits two entities to share a display name and reports +nothing**. Entities are keyed by the model-supplied id; `DUPLICATE_ENTITY` +rejects only a repeated *key*; no check exists on `name`. That is one of this +finding's failure modes, and establishes nothing about what happened. + +**M11 must run an explicit diagnostic** with a protagonist and three same-scene +supporting characters, stressing pronouns, dialogue attribution, entrances and +exits, reference by name and by role, and one character speaking about another. +It must detect duplicate creation, same-name duplication, protagonist drift, +misattributed dialogue, self-as-other reference, and state/context disagreement +— and on any failure preserve the pre-generation state, the exact stored prompt +snapshot, history, summaries, memories, imported knowledge, narrator output and +model settings, then classify: + +```text +STATE DEFECT / CONTEXT ASSEMBLY DEFECT / DERIVED MEMORY-SUMMARY DEFECT / +MODEL FAILURE WITH CORRECT CONTEXT / AMBIGUOUS +``` + +**Do not "fix" a model failure by changing authoritative state, and do not blame +the model if the prompt already contained the error.** M9 made all of that +evidence portable, so a failing campaign can be exported and handed over intact. + +**The standard fixture does not cover this class.** `TEST-CAMPAIGN-FIXTURE.md`'s +seven traps are knowledge, authority, branch leakage and possession; there is no +identity trap and its on-stage cast is effectively two people. The established +fixture was **not modified** — it is the deterministic baseline earlier results +are compared against. A companion fixture, `Multi-Character Identity Test`, is +proposed in an appendix to that document. + --- # M10 — Future Media Extension Hooks Only @@ -1126,6 +1307,13 @@ A campaign can be safely exported, imported into a clean data directory, and reo Preserve the approved future media interfaces without adding a media-generation dependency to v1. +## Note — the post-M8 playtest findings are **not** M10 scope + +The four findings recorded above are owned by M11. M10 inherits them as known +carry-forward items only: it should neither implement nor test them, and its +scope below is unchanged by them. They are listed before this milestone rather +than after it only because they were recorded during M9's closeout. + ## Scope - scene snapshots/packets suitable for future providers, @@ -1178,7 +1366,11 @@ Validate the full product against the release contract after all functional mile - fantasy and science-fiction fixtures, - export/import/recovery tests, - migration tests, -- documentation and packaging. +- documentation and packaging, +- **the four post-M8 hands-on playtest findings above**: the browser product + name, reader orientation after history movement, a narration-length setting + with a measurable effect, and the multi-character identity diagnostic with its + companion fixture. ## Tests / Acceptance diff --git a/planning/DATA-MODEL.md b/planning/DATA-MODEL.md index 7fe3bba..5217392 100644 --- a/planning/DATA-MODEL.md +++ b/planning/DATA-MODEL.md @@ -887,6 +887,101 @@ has stopped being told the rules, with nothing to notice. A bundle written before M7 has no knowledge section and imports with an empty library, which is what such a campaign had. +### M9: the format became a version, and the evidence started travelling + +**M9 bumped the format to `ai-dnd-adventure-v3`**, and the reason is a rule +rather than a preference. Everything M9 added *could* have been an optional key +read with `.get`, the way `persona`, `checkpoints`, `narrativeState` and +`knowledge` each were. That mechanism stops working at exactly this addition: +a v2 file carrying no prompt provenance is **ambiguous** — written before M9, +when no file could carry one, or by M9 from a campaign whose turns predate the +column? Those are different facts about the campaign and a reader has to be able +to tell them apart. It is the same distinction the head rule above draws when it +says a pre-M3 file opens at its tip *because that is the position such a file +recorded*. A version number is how a recovery file states what it was capable of +recording. The reader keeps every older version; only the writer moved. + +What each version can be trusted to say: + +```text +v1 a linear story, its turns, and its retries as a repeating group +v2 + the tree, the live flags, the after-snapshots, the chosen head, + Save Points, the narrative state document, imported knowledge +v3 + state events and proposals, historical prompt/context provenance, + lineage-anchored summaries, take parentage, memory authority +``` + +**Three categories, not two.** §31 below distinguishes authoritative from +derived, which was sufficient until M9 had to decide about stored prompts. They +are derived — a machine assembled them — and they must travel anyway, so the +rule the bundle applies has a middle category: + +```text +chosen what a person decided: the story, the head, the takes, the + Save Points, the classifications, the canon. travels +evidence what happened, and what the application was told at the time: + the state events and proposals, the per-turn prompt and the + passages it was shown, the model and generation settings that + turn ran under. travels +rebuildable a deterministic function of what travels: knowledge passages, + the FTS index, embeddings, the branch lineage cache. + rebuilt on import +``` + +The test that separates evidence from rebuildable is **not** "could this be +recomputed" but "would a recomputation answer the same question". Rebuilding the +FTS index answers the same question it answered before. Rebuilding an old turn's +prompt does not — it would say what that turn *would be told now*, from today's +canon, today's sources and today's state, which is the opposite of what the +context inspector is for. Historical evidence is not a cache. + +So v3 additionally carries, all restored verbatim: + +- **`stateEvents` and `stateProposals`.** §17's hybrid keeps the events for + audit and the snapshots for restore; v2 carried only the snapshots, so a moved + campaign could be read at any position and could no longer say what changed + there, who asserted it, or what the value was before. A **manual correction** + was the worst case: the one state change no narration explains, and with the + events gone nothing distinguished it from something the story established. + Both tables travel, because the inspector reads both — the event says what was + accepted and the proposal says what the model asked for and what was refused. +- **A per-node context snapshot**, which is `SPECIFICATION.md` §6.1's exact + prompt/context record. It carries the assembled prompt section by section, the + passages retrieved with the text each supplied, which summary was eligible, + and the model and generation settings the call ran under — so an old turn can + still say what it was told after the source was deleted, the canon edited, the + chunker changed and the state moved on. Stored once per turn on the live + attempt, so a retried turn is not a multiplier. +- **`summaries`**, with the coordinate that decides eligibility. v2 carried only + the `storySummary` mirror, which has no lineage of its own, so a restored + campaign resumed with no usable long-story continuity — and a summary + belonging to an abandoned line stays ineligible after the move for the same + reason it was before it: eligibility is the coordinate lying on the active + capped lineage, not a stored flag. +- **Take parentage**, so attempts under two different takes of one turn stay two + pagers rather than merging into one. +- **Memory `authority`**, so a heuristic memory is not promoted to accepted + story by being moved (F07). +- **Per-source `parserVersion`/`chunkingVersion`**, recording what produced the + passages a historical retrieval record describes. + +**Encoding.** A per-turn prompt contains the story so far, so one per turn is +O(turns²) in campaign length — measured at 20,797 bytes per turn at turn 20 and +55,291 at turn 120, 68% of a 9.7 MB file. The snapshot therefore travels as +`contextSnapshotZ`: the same JSON, zlib-compressed and base64-encoded, using the +same pack/unpack the database column already uses. Nothing is dropped or +summarised; the file is still JSON, and every other section of it is still plain +text. The plain `contextSnapshot` key is still read and takes precedence, so a +hand-edited file keeps importing. + +**Two pointers are translated on import, and nothing else is.** Branch numbers +already were. M9 adds the `source_id` inside a restored retrieval record: it +names a row on the machine that wrote the file, so left alone it would point the +inspector's "open this source" at whatever holds that id here. Where the file's +own knowledge section contains the source it is repointed; where it does not — a +source deleted before the export — it becomes `null`, and the record keeps its +text and filename. The evidence is never rewritten; only the pointer is. + ## 30. Deletion vs Archival The system must distinguish: @@ -916,6 +1011,23 @@ Retry, Undo, Restore, and branch switching must not silently perform permanent d Derived data should be rebuildable where practical. +**"Where practical" does real work in that sentence, and M9 had to split this +list to act on it** (§29). Embeddings and the lexical and semantic indexes are +deterministic functions of content that travels, so a rebuild answers the same +question and they are not exported. A **summary** is not: it took a model call, +it describes a stretch of story that may since have been abandoned, and +regenerating one on another machine produces different prose about a different +reading — so it is derived, not practically rebuildable, and it travels with the +coordinate that decides whether it still applies. + +The same reasoning puts **stored prompt/context snapshots** on the travelling +side, and they are not in either list above because they are neither: they are +not authoritative — nothing decides anything from them — and calling them +derived would invite a rebuild. They are *evidence*: a record of what the +application was told at the time, which a regeneration would not reproduce +because it would use today's canon, today's sources and today's state. §29 states +the three-way rule the export applies. + ## 32. Provenance Important information should answer: diff --git a/planning/IMPORTED-KNOWLEDGE-DESIGN.md b/planning/IMPORTED-KNOWLEDGE-DESIGN.md index 2617af1..d344213 100644 --- a/planning/IMPORTED-KNOWLEDGE-DESIGN.md +++ b/planning/IMPORTED-KNOWLEDGE-DESIGN.md @@ -1110,6 +1110,47 @@ Normal imported files are campaign-level source material and need not inherit st Story Cards may remain as an inherited authored-rule/lore primitive during migration if useful, but they must not become an alternate untracked path around the new knowledge authority/provenance rules. +### Settled in M9 (2026-09-07): compatibility-only, and out of the prompt + +M8 removed the Story Card browser editor and left the question open; the M9 +brief asked for it to be decided. The finding was that story cards **were** the +alternate untracked path the paragraph above forbids, and not in principle: a +keyword-matched card was injected into the narrator's prompt as +`World Lore: `, taking up to 40% of what was left after the imported +knowledge had been placed, with + +- no class, so nothing framed how far the narrator could rely on it; +- no visibility, so no narrator-only distinction existed; +- no source, no hash and no lifecycle, so there was nothing to disable; +- no browser surface after M8, so a reader could neither see nor switch it off; +- no row in the context inspector, which renders `knowledge` and never rendered + `cards`; + +and competing with imported Canon for one budget, which is the arrangement M7 +spent a milestone separating. + +**The decision, and it is the smallest change that closes it:** + +| | | +| --- | --- | +| New export | carries them, unchanged, under `storyCards` | +| Legacy import | accepted, unchanged, from every format version | +| Normal narration | **no longer reached.** The `world_lore` section is gone | +| Re-export | carries them again, so a round trip destroys nothing | + +Nothing is deleted. The rows stay, the `/api/story-cards` endpoints stay, and +`memorybank.cast_brief` still reads them as the **summariser's character +roster** — that names who is on stage so a memory says "Aldric" rather than +"he", never reaches the narrator, and every memory written from it is +authority-classified by the application afterwards. The `cards` key stays in the +context report and is now always empty for a new turn, because M9 made +historical snapshots portable and an old turn's record must go on saying that +story cards were included. + +A campaign that wants the narrator to know something imports it as Canon, +Reference or Inspiration, where it is classified, inspectable, disableable and +attributable — which is what §73 asks for. + ### Retrieval implementation direction Use: diff --git a/planning/README.md b/planning/README.md index c2f207c..b0c6693 100644 --- a/planning/README.md +++ b/planning/README.md @@ -3,7 +3,8 @@ **This file is the index. Start here.** **Current state:** Phase 0 complete; AI-DnD forked as the production base; -milestones **M1 through M6 implemented and accepted** (M3 and M4: 2026-09-03; +milestones **M1 through M8 implemented and accepted**, and **M9 implemented and +awaiting review**. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03; M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an independent review found a real defect and a corrective pass fixed it. @@ -19,8 +20,15 @@ carries the closeout: the build-evidence classification in its §P, finding 14's operational resolution, and the acceptance record in its §V. The M8 tree is staged and awaits the repository owner's signed commit. -**Next: M9 — Export, Backup, Recovery, and Migration Hardening.** It has not -been started. +**M9 — Export, Backup, Recovery, and Migration Hardening — is implemented and +awaiting independent review** (2026-09-07). +`reports/M9-IMPLEMENTATION-REPORT.md` is the implementer's account, written for +a reviewer: a set of claims with the measurements attached, not yet a record of +acceptance. M8's report has moved to `archive/milestone-reports/`, which is +where a milestone report goes once the next milestone's report replaces it. + +**Next: M10 — Future Media Extension Hooks Only.** It has not been started, and +no brief for it exists. **Package version:** see `VERSION.md`, which records what each revision changed and why. @@ -82,8 +90,8 @@ Two standing qualifications: | Document | What it is for | | --- | --- | | `SPECIFICATION.md` | What the product must do. The top of the authority order. | -| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M7 built, recorded as fact. | -| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the export shape. | +| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M9 built, recorded as fact. | +| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the v3 export contract. | | `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. | | `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. | | `IMPORTED-KNOWLEDGE-DESIGN.md` | Canon / Reference / Inspiration knowledge as a first-class subsystem. | @@ -110,7 +118,7 @@ Two standing qualifications: 10. `BROWSER-UX-SPEC.md` 11. `V1-ACCEPTANCE-TESTS.md` 12. `DECISIONS/` — all of them; they are short. -13. `reports/M8-IMPLEMENTATION-REPORT.md`, for what the most recent milestone +13. `reports/M9-IMPLEMENTATION-REPORT.md`, for what the most recent milestone actually left behind — reading it as a claim to check, not a record, until it is reviewed. Nothing in `planning/archive/` unless sent there. @@ -143,20 +151,18 @@ work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in `reports/` holds the report for the milestone most recently completed, because that is the one the next milestone's planning has to consult: -- `reports/M8-IMPLEMENTATION-REPORT.md` — the M8 implementation, its baseline - UX measurement, and the browser evidence for every acceptance test it claims. - Written by the implementer for an independent reviewer, and completed at - closeout after that review accepted the milestone: it is a set of claims with - the measurements attached **and** the record of the acceptance. Its §U carries - the M9 handoff — the four questions the next brief has to decide. +- `reports/M9-IMPLEMENTATION-REPORT.md` — the M9 implementation: the measured M8 + portability baseline it started from, the final bundle contract, and the + evidence for every acceptance test it claims. Written by the implementer for + an independent reviewer, so it is a set of claims with the measurements + attached and **not** a record of acceptance. Its §W carries the M10-M11 + handoff. - **It stays here until M9's report replaces it.** A milestone report is useful - during the immediately following milestone; M8's is not archived merely - because M8 is accepted. + **It stays here until M10's report replaces it.** -Completed earlier milestones are in `archive/milestone-reports/`, which M7's -report joined when M8's was written: a milestone report is useful during the -immediate next milestone and historical afterwards. M1-M7 are all there, +Completed earlier milestones are in `archive/milestone-reports/`, which M8's +report joined when M9's was written: a milestone report is useful during the +immediate next milestone and historical afterwards. M1-M8 are all there, unedited. ## The decision this package rests on @@ -282,33 +288,39 @@ Milestone M8 COMPLETE / ACCEPTED (2026-09-06) story operations review + closeout, in sequence | v -Milestone M9 NEXT — not started - export, backup, recovery, see BUILD-MILESTONES.md - migration hardening +Milestone M9 COMPLETE — awaiting review (2026-09-07) + export, backup, recovery, reports/M9-IMPLEMENTATION-REPORT.md + migration hardening bundle format v3; SQLite online backup | v -M10-M11, one at a time see BUILD-MILESTONES.md +Milestone M10 NEXT — not started + future media extension hooks see BUILD-MILESTONES.md + | + v +Milestone M11 see BUILD-MILESTONES.md ``` ## Stop Rule **One milestone at a time. Do not begin a milestone before its brief exists.** -**No M9 brief has been prepared.** Writing one is the current action, informed -by the M8 report and by the debt `BUILD-MILESTONES.md` records against M8 — in -particular that the campaign bundle still carries no context snapshots, so an -imported campaign has no historical prompt provenance; that story cards survive -in the backend and the bundle with no browser surface, and M9 should decide -deliberately whether the bundle keeps carrying them; and that a deployment whose -Ollama enforces a small context window truncates an imported long campaign -immediately unless the `DEVELOPMENT.md` procedure or a matching -`context_token_budget` is applied. +**No M10 brief has been prepared**, and M9 is not accepted — it is implemented +and awaiting an independent review. Writing the M10 brief is the action after +that review closes, informed by the M9 report's §W. -**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`: story -cards have no browser editor, the RPG world state is read-only, copy is -per-message only, there is no discarded-history recovery screen, and the tablet -layout is usable but untuned. Each names the milestone that owns it; none is an -open M8 condition. +All three questions the M8 debt raised against M9 are settled and recorded: +the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story +cards are compatibility-only legacy data and no longer reach the narrator +(`IMPORTED-KNOWLEDGE-DESIGN.md` §73); and the deployment context ceiling is +documented in `DEVELOPMENT.md` with the note that an imported long campaign +meets it on its first turn rather than gradually. The window itself stays M11's. + +**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`. Two of +its five items are now closed by M9 — story cards have a settled policy, and the +bundle carries the provenance. The RPG world state is still read-only, copy is +still per-message only, there is still no discarded-history recovery screen, and +the tablet layout is still untuned. Each names the milestone that owns it; none +is an open M8 condition. The M6 retrieval debt this milestone was warned about is partly addressed and partly still open. Imported material does **not** compete with story memory for diff --git a/planning/TECHNICAL-DESIGN.md b/planning/TECHNICAL-DESIGN.md index a0d9a0b..2d1d926 100644 --- a/planning/TECHNICAL-DESIGN.md +++ b/planning/TECHNICAL-DESIGN.md @@ -608,6 +608,89 @@ stated reason. A misplaced head affects every read in the file; a bookmark pointing outside the story affects only itself, and rejecting a whole campaign to protect one bookmark would lose the story to save the pointer. +### 9.3 As implemented in M9 — the version, and the third category + +**The format is now `ai-dnd-adventure-v3`, and the bump is the design.** §9.1 and +§9.2 each declined one, correctly: an absent `headDepth` or `checkpoints` key is +unambiguous, because a file either states a position or it does not. That +property fails for what M9 adds. A v2 file with no prompt provenance may have +been written before M9, when no file could carry any, or by M9 from a campaign +whose turns predate the column — different facts about the campaign, and a +reader has to be able to tell them apart. A version number is how a recovery +file states what it was capable of recording, which is exactly the reasoning +§9.1 uses to justify opening a pre-M3 file at its tip. The reader keeps every +version; only the writer moved. + +**§9.1's two categories became three.** "Chosen travels, derived is recomputed" +was sufficient until M9 had to decide about stored prompts, which are derived — +a machine assembled them — and must travel anyway: + +```text +chosen the story, the head, the takes, the Save Points, the + classifications, the canon travels +evidence the state events and proposals, the per-turn prompt and the + passages it was shown, the model and generation settings that + turn ran under travels +rebuildable knowledge passages, the FTS index, embeddings, the branch + lineage cache rebuilt on import +``` + +The test separating the last two is not "could this be recomputed" but "would a +recomputation answer the same question". A rebuilt FTS index answers the same +question. A rebuilt prompt does not — it says what the turn *would be told now*, +from today's canon, today's sources and today's state, which is the opposite of +what the inspector is for. Historical evidence is not a cache, so M9 does not +regenerate one on import at any point. + +`DATA-MODEL.md` §29 lists what v3 carries. Three implementation facts belong +here rather than there: + +1. **The snapshots are encoded, not summarised.** A per-turn prompt contains the + story so far, so one per turn is O(turns²) — measured at 68% of a 9.7 MB file + at 120 turns, against a 20 MB import ceiling. The snapshot therefore travels + as `contextSnapshotZ`, zlib-compressed and base64-encoded through the same + `compression.pack`/`unpack` the database column already uses. The file is + still JSON and every other section of it is still plain text. The readable + `contextSnapshot` key is still accepted and wins when both are present, so a + hand-edited file keeps importing. A residual ceiling remains and is stated in + the M9 report rather than hidden. +2. **One more pointer is translated, and only pointers ever are.** Branch + numbers already were. A restored retrieval record's `source_id` names a row + on the machine that wrote the file, so it is repointed at the source that + landed here, or set to `null` when the file carries no such source. The text + the record holds — the evidence — is never rewritten. +3. **Import stays two-phase inside one transaction.** `plan` refuses everything + a hand-edited file can get wrong before a row exists; `materialize` writes, + and the endpoint commits once and rolls back explicitly otherwise. A + *rebuildable* index failing after that does not roll the campaign back: it is + reported on the response as a warning, shown per source in the Knowledge + panel, and repaired by Reindex. So a caller sees either "the campaign is not + there" or "the campaign is complete", never a third thing. + +### 9.4 The database backup, as implemented in M9 + +A second recovery tool, deliberately not merged with the first. The bundle is a +logical, portable, human-readable copy of **one campaign** and is the supported +way to move a campaign between installations; the backup is a physical copy of +**this machine's whole database** and is what you take before an upgrade. + +`backend/app/backup.py` uses SQLite's online backup API rather than a file copy, +because a copy taken while the application runs can read one page before a +transaction and another after it and produce a file that opens, reports a schema +and is quietly missing rows. It writes to a temporary name beside the +destination, runs `PRAGMA quick_check` against the finished file, and only then +renames it into place; it opens the source read-only, never overwrites an +existing backup, and leaves nothing behind on failure. + +No path comes from a caller: the destination is derived from the database the +application already has open and the filename from the clock, so the endpoints +accept no body at all (H08). + +**There is no restore endpoint, and that is a decision.** Restoring means +replacing the file the running process has open, which is how both copies are +lost at once. The procedure is in `DEVELOPMENT.md` and is a procedure precisely +because each step needs the application stopped. + ## 10. Authoritative Narrative State ### 10.1 Do not retain the RPG state protocol as the product model diff --git a/planning/TEST-CAMPAIGN-FIXTURE.md b/planning/TEST-CAMPAIGN-FIXTURE.md index bb5f86a..f013551 100644 --- a/planning/TEST-CAMPAIGN-FIXTURE.md +++ b/planning/TEST-CAMPAIGN-FIXTURE.md @@ -996,3 +996,117 @@ actual persistence / memory / canon bugs The prose may vary. The expected state, authority, and lineage rules should not. + +--- + +# Appendix A — Proposed companion fixture: Multi-Character Identity Test + +**Status: proposed, not built. This appendix changes nothing above it.** + +## Why a companion rather than an extension + +The Continuity Test above is the deterministic baseline that several milestones' +results are compared against. Adding characters or turns to it would invalidate +those comparisons, so **it is deliberately left exactly as it is**. + +## The gap this fills + +Reviewed on 2026-09-07 against a hands-on finding. The Continuity Test's seven +deliberate traps (§13) are: + +```text +1-3 knowledge boundaries and secrets +4 reference authority +5 canon precedence +6 branch leakage +7 possession +``` + +There is **no identity trap**, the word *coreference* does not appear, and the +on-stage cast is effectively two people — Aldric and Mara, with Edrin +established as missing rather than present. So **same-scene multi-character +identity continuity is not exercised anywhere in the standard fixture.** + +A play session against accepted M8 produced exactly that failure: four people in +one office, and narration that treated one of them as two different people +sharing a name. Root cause is unknown and no longer establishable — the playtest +database was destroyed — which is itself part of why a *deterministic* fixture +for this class is worth having. + +## Shape + +Four people, all present in one ordinary scene, with no fantasy vocabulary — the +point is identity, not genre: + +```yaml +bill: { type: character, role: protagonist, controlled_by: reader } +alice: { type: character, role: coworker } +roger: { type: character, role: coworker } +john: { type: character, role: coworker } +location: { type: location, name: the office } +``` + +Identities and roles established unambiguously before the first test turn, so +that any later ambiguity is the system's and not the setup's. + +## What the sequence must stress + +- pronouns with more than one plausible referent in scene; +- dialogue attribution across three speakers; +- characters entering and leaving; +- reference by name **and** by role, for the same person; +- one character speaking *about* another; +- one character speaking about **themself in the third person**, which is the + shape the observed failure took. + +## Traps + +| | | +| --- | --- | +| **I1 — one Alice** | No turn may produce a second entity whose display name is `Alice`. The state model currently permits this and reports nothing, so the trap is real rather than theoretical | +| **I2 — the protagonist stays the protagonist** | Bill must not drift into being narrated as a third party, or acquire a second entity | +| **I3 — attribution** | A line spoken by Roger must not be attributed to John | +| **I4 — self-reference** | No character may refer to themself as a separate same-named person | +| **I5 — state and context agree** | The authoritative state and the assembled prompt must not disagree about who is present or who anyone is | +| **I6 — exit and return** | A character who leaves and returns is the same entity, not a new one | + +## Required evidence on failure + +Unlike the fixture above, this one exists to **classify** a failure rather than +only to detect it, because its failure modes are split between the application +and the model. Any failing turn must capture: + +```text +authoritative state immediately before generation +the exact stored context/prompt snapshot +the recent-history section +summaries +retrieved memories +imported knowledge, if any +narrator output +model identifier and generation settings +``` + +then classify: + +```text +STATE DEFECT +CONTEXT ASSEMBLY DEFECT +DERIVED MEMORY/SUMMARY DEFECT +MODEL FAILURE WITH CORRECT CONTEXT +AMBIGUOUS / MULTIPLE CONTRIBUTORS +``` + +Two rules for whoever runs it: **do not "fix" a model failure by editing +authoritative state**, and **do not blame the model when the prompt already +contained the identity error.** + +Since M9 the whole of that evidence is portable in one campaign bundle, so a +failing run can be exported intact and investigated elsewhere. + +## Ownership + +M11, alongside the realistic-model review. `V1-ACCEPTANCE-TESTS.md` §P3 records +which parts of this are candidate **acceptance** criteria — the state-level +traps, which are decidable — and which are model-quality observations that +belong in a recorded review rather than in the pass/fail contract. diff --git a/planning/V1-ACCEPTANCE-TESTS.md b/planning/V1-ACCEPTANCE-TESTS.md index 5088018..00b3c7d 100644 --- a/planning/V1-ACCEPTANCE-TESTS.md +++ b/planning/V1-ACCEPTANCE-TESTS.md @@ -1800,6 +1800,16 @@ Export standard campaign. ### Pass Export completes locally and contains enough data to restore story. +### Result — PASS (M9, 2026-09-07) +`ai-dnd-adventure-v3`, checked section by section rather than by file size: +the story and its whole retained tree, the branches and their disposition, the +chosen head, Save Points, the authoritative state and its per-position +snapshots, the state events and proposals, the imported library, the +lineage-anchored summaries, the memories, and a stored prompt for every narrator +turn that has one. `test_m9_portability.py::test_i01_*`, and reproducible with +`python -m tools.m9_portability_report`, which classifies every data family as +PRESERVED, OMITTED or DERIVED/REBUILDABLE. + --- ## I02 — Import Exported Campaign @@ -1814,6 +1824,14 @@ Export completes locally and contains enough data to restore story. ### Pass Active transcript and state are restored. +### Result — PASS (M9, 2026-09-07) +Across a **genuine clean data directory**: a second server process, in a second +directory, against a database file that has never existed, with the exporting +process stopped. Transcript, authoritative state, canon, Save Points, knowledge, +state events and the size of the retained tree all match the source campaign +(`test_m9_clean_import.py`). The same round trip inside one process is in +`test_m9_portability.py`, and is labelled there as the weaker of the two. + --- ## I03 — Branch/Disposable History Export @@ -1823,6 +1841,19 @@ Active transcript and state are restored. ### Pass Export preserves retained alternate/disposable history needed for recovery, unless user explicitly chooses a trimmed export. +### Result — PASS (M9, 2026-09-07) +Both futures come back and stay distinguishable: the abandoned line's turns are +present and readable, the branch that was left carries the depth it was left at, +and a superseded take is still at its coordinate with the selected take still +selected. There is no trimmed export — M9 offers no option to drop history, so +the clause about one does not arise. + +**M9 also closed a fidelity gap here.** Take parentage was not exported, so every +imported node landed parentless and the pager grouped on the coordinate instead. +That is right for a plain retry and wrong once two takes of one turn each have +takes of their own beneath them; the copy read `5/5` where the source read `2/2` +and `3/3`. v3 carries the parentage. + --- ## I04 — Checkpoint Export @@ -1841,6 +1872,15 @@ campaign. Importing Save Points does **not** move the active head — the head still comes from the bundle's `headDepth`. Bundles written before M4 carry no `checkpoints` key, import cleanly, and create none. +### Re-verified — PASS (M9, 2026-09-07) +Unchanged by the format bump, and extended in two directions. Every restored +Save Point resolves, restores to the position it names through M3's head +movement, leaves the retained history it moved back over intact, and the two +in the M9 fixture restore to *different* states. A Save Point whose coordinate +is not in the file is **dropped with the rest of the campaign kept**, never +retargeted to a nearby turn: the reader named a position, and if that position +is not in the file then no other position is the one they named. + --- ## I05 — Knowledge Provenance Export @@ -1870,13 +1910,29 @@ hand-edited knowledge block with an unknown classification or empty content refuses the import rather than half-landing in it; and an edited content hash is recomputed from what actually arrived and the discrepancy recorded on the source. -**Limit, unchanged from before M7 and owned by M9.** The bundle carries no -context snapshots at all, so an imported campaign has no historical prompt -provenance — for imported knowledge or for any other component. Nothing M7 -creates is turned into a dangling id by a round trip, because no ids are -exported; the evidence simply is not in the file. -`test_historical_prompt_evidence_survives_an_export_round_trip` pins that -behaviour so it cannot regress silently. +**That limit is closed (M9, 2026-09-07).** The paragraph below is kept as +written because it records what was true through M7 and M8, and the M9 decision +is only legible against it. + +> *The bundle carries no context snapshots at all, so an imported campaign has no +> historical prompt provenance — for imported knowledge or for any other +> component. Nothing M7 creates is turned into a dangling id by a round trip, +> because no ids are exported; the evidence simply is not in the file.* + +M9 carries the snapshots. An old turn in a restored campaign shows the prompt it +was actually assembled from, the passages it was shown and the text each +supplied — after the source has been deleted, the canon edited and the state +moved on. `test_historical_prompt_evidence_survives_an_export_round_trip` was +inverted rather than deleted: it now pins the thing M7 was worried about and +could not check, which is that the provenance arriving on the other side names +*this* campaign's sources rather than the ids they had where the file was +written. Only that pointer is translated; the evidence is restored verbatim, and +a source the file does not carry becomes `null` rather than pointing at a +different file. + +Also added in M9: `parserVersion` and `chunkingVersion` per source, recording +what produced the passages a historical retrieval record describes, and a +`sourceId` that exists only so the translation above can be made. --- @@ -1887,6 +1943,28 @@ behaviour so it cannot regress silently. ### Pass No external API credentials are embedded in campaign export. +### Result — PASS (M9, 2026-09-07) +Tested rather than assumed, and tested against the file's **text** rather than +against a list of columns — a field added to a model the exporter walks would +otherwise reach the bundle with no test of a column noticing. The inert +`api_key` column is written with a recognisable value first, so its absence is +evidence rather than a tautology. + +Nothing matching `api_key`, `apiKey`, the written secret, the inference endpoint +or its port, or an absolute filesystem path appears anywhere in an export — +including inside the compressed snapshots, which the test decodes rather than +skipping. Checked in both suites, so the clean-directory run covers the same +ground across a real process boundary +(`test_m9_portability.py::test_i06_*`, `test_m9_clean_import.py`). + +**What deliberately does not travel**, and why it is not an omission: the +inference endpoint, the model name, the context budget and every other row of +`settings`. Those describe the machine, not the campaign, and importing a +campaign must not silently repoint the destination's inference at the source's. +Per-turn model and generation settings *do* travel, inside the historical +snapshot, because there they are a record of what happened rather than a +configuration to apply. + --- ## I07 — Export/Import Preserves an Undone Active Head @@ -1921,6 +1999,28 @@ An export whose stated head lies beyond the story it contains is a file disagreeing with itself and must be refused rather than opened at a guessed position. +### Result — PASS (M9, 2026-09-07) +All three clauses, and the main one across a genuine machine boundary. + +- **The exact head.** The M9 fixture ends two Undos behind its own branch's + retained tip and behind the abandoned line's, and the last thing it does is an + Undo — so the head is not the newest row written, not the deepest row, not the + tip, and not on the branch holding the most story. An importer guessing any one + of those lands somewhere else. The copy opens exactly where the source was, + the later turns are still in the database as retained future, and Redo is + offered rather than the story having silently been redone. Redo then walks to + the same next turn in both. +- **The legacy clause.** A bundle with its `headDepth` removed opens at the tip + of its head branch, offers no Redo, and offers Undo — which is the position + such a file recorded, because at the time it was written the head could not be + anywhere else. Checked at every seam in `test_m9_legacy_bundles.py`. +- **The self-disagreeing file.** A head past the retained story is refused with + a message naming where the branch actually ends, and nothing is written. + +Evidence: `test_m9_clean_import.py::test_it_opens_at_the_exact_head_it_was_exported_at` +(second process, empty directory), plus `test_m9_portability.py::test_i07_*` and +the head cases in `test_m9_corrupt_bundles.py`. + --- # J. Genre Independence @@ -2082,6 +2182,17 @@ turn. ### Pass State at each position matches original accepted state. +### Result — PASS (M9, 2026-09-07) +Measured **after a round trip**, which is the M9 form of it: the copy and the +source are walked back three turns and forward three turns in step, and the +authoritative document is compared at every position. They agree throughout. + +That this stays a snapshot read rather than a replay is the point. Undo, Redo +and Save Point restore all resolve a coordinate and read the state recorded +there (`TECHNICAL-DESIGN.md` §10.4), so an import that carried the events and +dropped the per-position snapshots would have made every one of them +proportional to campaign length. Both halves of §17's hybrid travel. + --- ## L03 — Checkpoint Reconstruction After Restart @@ -2105,6 +2216,15 @@ exactly that value after the campaign had been advanced past it `TestClient` restart, which could not distinguish durable state from a live object. +### Re-verified after a move — PASS (M9, 2026-09-07) +The same claim with a machine boundary in front of it. A campaign is exported +from one server process, imported into a **second process against a database +file that has never existed**, a Save Point is restored there, that process is +killed, and a **third** process against the same file is asked again. The +transcript, the authoritative state and the size of the retained tree all match +what the second process had after restoring +(`test_m9_clean_import.py::test_l03_*`). + --- ## L04 — Derived Data Can Be Rebuilt @@ -2121,6 +2241,34 @@ using a safe test copy. ### Pass Authoritative campaign history remains intact and derived structures can be recreated. +### Result — PASS (M9, 2026-09-07), and one defect found by running it +On a safe copy — an imported campaign, not the original. Every physically +derived structure is destroyed and rebuilt from the source content the bundle +carried: passages, the FTS rows and the vectors. Afterwards every source is +`ready` with passages again, retrieval works, and the transcript, the +authoritative state, the classifications and the lifecycle flags are identical +either side. Deleting the vectors alone leaves lexical retrieval working, which +is M7's rule that the lexical half is a production path and not a fallback. A +rebuild does not make an abandoned line's summary eligible. + +**Running it found a real defect, which is fixed here.** The FTS5 index is a +virtual table, so no foreign key reaches it and no `ON DELETE CASCADE` covers +it: deleting a campaign dropped its passages and left one index row per passage +behind. Nothing read them — every search joins through `knowledge_chunks` — so +the leak was invisible until SQLite handed the freed primary key out again, at +which point the **next source imported into any campaign** failed with an +integrity error. Reindex could not repair it either, because `clear_index` finds +index rows *through* the chunks, and there were none. Both ends are closed: a +campaign's index rows are removed before it is deleted, and the index insert now +replaces a stale row rather than colliding with it — so a database already +carrying the leak repairs itself and needs no migration. See the M9 report's +findings. + +The rebuild path M9 depends on is therefore implemented and tested rather than +assumed, which the SHOULD priority above did not require but the milestone did: +knowledge passages and indexes are omitted from the bundle precisely because +they can be rebuilt. + --- # M. Long-Run Test @@ -2327,3 +2475,92 @@ L01-L03 ``` This keeps implementation work tied to observable behavior rather than repository-specific architecture. + +--- + +# P. M11 Test-Design Tasks — Not Yet Acceptance Tests + +**Nothing in this section is part of the pass/fail contract.** These are three +behaviours a real play session against accepted M8 showed are worth testing, for +which **the pass criterion is not yet settled**. They are recorded here so the +release tester finds them where they will look, and they are deliberately not +written as A-through-O items: giving them IDs would imply a criterion has been +ratified when it has not, and would destabilise a document whose value is that +every item in it is decidable. + +Each names **what must be settled** before it can become an acceptance test. +Full context is in the M9 report's §Y and, durably, in `BUILD-MILESTONES.md` +under the post-M8 playtest findings. + +## P1 — The reader can tell where they are after history movement + +**From:** a real session in which Undo worked correctly and the reader could not +tell which point in the story they had reached. + +**Behaviour to test:** after Undo, Redo, a Save Point restore, or an edit to an +earlier turn, the reader can identify their current position in the visible +story without implementation terminology (`branch`, `head`, `node`, `depth` +remain forbidden at the surface). + +**Settle first:** what the indicator *is*. "The reader can tell" is not +decidable as written — it needs an observable artifact, such as a named position +that changes with movement and is present in the DOM. `BROWSER-UX-SPEC.md` §8 +now carries the requirement; **no wording is ratified**, and this cannot become +an acceptance test before one is. + +## P2 — Narration length has a measurable directional effect + +**From:** a reader who chose *2-4 paragraphs* and received substantially longer +replies. + +**Behaviour to test:** the narration-length setting produces a **measurable +directional difference** in output length across repeated realistic turns — +brief shorter than standard, standard shorter than detailed — on at least the +reference 3B narrator and one stronger local narrator. + +**Settle first:** the numbers, and the mechanism. Today the setting adds one +English sentence to the campaign instructions and changes **no** generation +budget, while a separate numeric hint derived from the global +`max_output_tokens` is identical for every setting (M9 report §Y). Until the +product decides what each setting *means* — and whether it moves the budget — +there is no threshold to test against. A directional test is stateable; an +absolute one is not, and this should not become an acceptance test that asserts +word counts nobody has ratified. + +**Do not** turn this into a truncation test: the state block is emitted last and +hard truncation removes it. + +## P3 — Multi-character identity continuity + +**From:** four people in one scene, and narration that treated one of them as +two different people of the same name. **Root cause unknown** — the playtest +database was destroyed, so no evidence survives. + +**Behaviour to test:** across a multi-turn scene with a protagonist and three +supporting characters, the story does not create duplicate characters, does not +duplicate a display name across two entities, does not drift the protagonist's +identity, does not misattribute dialogue, and does not have a character refer to +themself as a separate same-named character — and the authoritative state and +the assembled context do not disagree about who anyone is. + +**Settle first:** which of those are **product** guarantees and which are +**model-quality** observations. They are not the same kind of claim and must not +share one verdict: + +- *"The state never holds two entities with the same display name"* is + decidable and enforceable, and is a candidate acceptance test today. The + implementation currently permits it and reports nothing (M9 report §Y). +- *"The narrator never confuses two same-named characters"* is not a pass/fail + property of this application — it depends on the model — and belongs in M11's + realistic-model review with a recorded classification, not in this contract. + +**A run of this must capture**, on any failure: pre-generation state, the exact +stored prompt snapshot, history, summaries, retrieved memories, imported +knowledge, narrator output, and model settings — then classify as a state, +context-assembly, derived-data, or model failure. M9 made all of that portable, +so a failing campaign can be exported whole and investigated elsewhere. + +**Fixture:** the standard Continuity Test does not exercise this — its traps are +knowledge, authority, branch leakage and possession, and its on-stage cast is +effectively two people. A companion fixture is proposed in +`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged. diff --git a/planning/VERSION.md b/planning/VERSION.md index 29d4a02..6d5729d 100644 --- a/planning/VERSION.md +++ b/planning/VERSION.md @@ -1,8 +1,101 @@ # Planning Package Version -- **Package:** Adventure Storyteller Planning Package v3.3 -- **Revision date:** 2026-09-06 -- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted** (M8 closed out 2026-09-06). M9 is next and has not been started. +- **Package:** Adventure Storyteller Planning Package v3.5 +- **Revision date:** 2026-09-07 +- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 implemented and awaiting independent review** (2026-09-07). M10 has not been started. + +## v3.5 — Post-M8 hands-on playtest findings recorded (2026-09-07) + +**Documentation only. No application code changed, and M9's verified result is +untouched** — see the note at the end of this entry. + +A real play session against **accepted, signed M8** (real browser, trusted-LAN +Ollama, `qwen2.5:3b-instruct-16k`, disposable database since destroyed) surfaced +four product-quality observations. **None is an M9 defect, none was caused by +M9, and none blocks M9 acceptance.** They are recorded so they cannot be lost +when the M9 report is archived. + +| Document | Change | Kind | +| --- | --- | --- | +| `reports/M9-IMPLEMENTATION-REPORT.md` | **New §Y**, the full write-up with the verification behind each mechanism, plus a pointer from §W. Explicitly labelled as not-M9. | observation record | +| `BUILD-MILESTONES.md` | **New section before M10**, the durable copy owned by M11; a note bounding M10 out of it; the four items added to M11's scope list. | milestone sequencing | +| `V1-ACCEPTANCE-TESTS.md` | **New §P**, three items as **test-design tasks, explicitly not acceptance tests**, each naming what must be settled before it could become one. No existing test changed or weakened. | test design | +| `BROWSER-UX-SPEC.md` | **New §8A.** §8's *"The current endpoint should be clear"* was satisfied while a real reader was lost, so it could not hold the behaviour. States the orientation requirement; **prescribes no wording**. | requirement clarification | +| `TEST-CAMPAIGN-FIXTURE.md` | **New Appendix A** proposing a companion `Multi-Character Identity Test`. The established deterministic fixture is **unchanged** — altering it would invalidate earlier milestones' comparisons. | test design | + +**What was verified rather than assumed**, since the playtest campaign no longer +exists and root causes largely cannot be proven: + +- The browser title genuinely is `AI D&D` (`frontend/index.html`), and **no + accepted document ever claimed otherwise** — so this is an uncovered gap, not + documentation needing correction. Nothing was corrected and nothing was + renamed: *Adventure Storyteller* is itself narrower than the genre-agnostic + engine `SPECIFICATION.md` requires, and the naming decision is the owner's. +- The narration-length setting adds **one English sentence** and changes **no + generation budget**, while the numeric hint derived from the global + `max_output_tokens` is **identical for every setting** — measured at the + default as *"must not exceed 506 words, and it should not stop short of about + 177."* Recorded as a mechanism to check first, **not** as the proven cause. +- The narrative state **permits two entities to share a display name and + reports nothing** — `DUPLICATE_ENTITY` rejects a repeated key only. That is + one of the identity finding's failure modes; it establishes nothing about what + actually happened. + +**M9 is unaffected.** No application file changed in this pass, so M9's final +verified result stands exactly as recorded: **1,102 backend passed / 14 skipped +/ 0 failed**, 145 frontend, 36/36 browser. The expensive M9 suites were +deliberately **not** re-run, because only Markdown changed. + +## v3.4 — M9 Implementation (2026-09-07) + +M9 — Export, Backup, Recovery, and Migration Hardening — is implemented on +`m9-recovery` from the signed M8 commit `1ce9972`. This revision records what the +implementation settled. It is **not** an acceptance: the milestone report is +written for a reviewer and the tree is staged for the repository owner's signed +commit. + +**Requirement corrections:** none. M9 altered no product requirement. +`SPECIFICATION.md` and `SECURITY-THREAT-MODEL.md` are unchanged — §16 already +required the export to preserve the exact active position, §6.1 already required +an exact prompt/context snapshot per turn, and M9 implements both rather than +redefining either. + +| Document | Change | Kind | +| --- | --- | --- | +| `DATA-MODEL.md` §29 | The v3 format, the three-category rule (chosen / evidence / rebuildable), what each version can be trusted to say, the encoding of the snapshots, and the two pointers the import translates. | implementation fact | +| `TECHNICAL-DESIGN.md` §9.3 | Why the version was bumped when §9.1 and §9.2 each correctly declined one; the third data category; the two-phase transaction and the warning path for a failed derived rebuild. | implementation fact | +| `TECHNICAL-DESIGN.md` §9.4 | **New.** The SQLite backup: the online backup API rather than a file copy, the verify-then-rename order, and why there is no restore endpoint. | implementation fact | +| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **New subsection.** Story Cards settled as compatibility-only legacy data and removed from the narrator's prompt, with the evidence that they were the "alternate untracked path" §73 already forbade. | newly settled design decision | +| `V1-ACCEPTANCE-TESTS.md` I01-I07, L02-L04 | Results recorded. I05's M7-era limit is marked closed with the original paragraph kept, because the M9 decision is only legible against it. L04 records the defect running it found. | implementation fact | +| `BUILD-MILESTONES.md` M9 | Marked complete, with what it delivered, the three defects it found, the story-card decision, and the debt carried forward. | implementation fact | +| `README.md`, `VERSION.md` | Status. | implementation fact | +| `DEVELOPMENT.md` | **New section**: the two recovery tools and when each applies, taking a backup, and the stop-move-start restore procedure. Plus a note that an imported long campaign meets a small context ceiling on its first turn rather than gradually. | implementation fact | + +**The four M8 handoff questions, answered** + +| | Answer | +| --- | --- | +| **A. Complete campaign portability** | Every family travels and is measured family by family, before and after, by a tool a reviewer can rerun. | +| **B. Historical prompt provenance** | **It belongs in the bundle, and it is in it.** An old turn in a restored campaign shows what it was actually given, after the source has been deleted and the canon edited. | +| **C. Legacy story cards** | Compatibility-only. Carried in both directions; removed from the narrator's prompt; still the summariser's character roster. | +| **D. Context-window portability** | The campaign travels; the machine's model configuration does not. Importing changes no setting of the destination's, and a campaign imports whether or not any model is installed. The window itself remains M11's. | + +**What the implementation found rather than assumed** + +Three defects, all found by running the milestone's own tests rather than by +reading: an FTS index leak that made an ordinary import fail in an unrelated +campaign and that Reindex could not repair; an imported node with no state +snapshot being stamped with the campaign's *head* state; and a snapshot pointer +that was not being translated because the code mutated a dict in place. The +first two predate M9. + +One measurement changed a plan, and then corrected the conclusion drawn from it. +Carrying per-turn prompts looked like it would halve the length of campaign that +can be restored. Measured, and after compressing them inside the file, everything +M9 added costs **12%** of reachable campaign length — the import ceiling moves +from about 318 turns to about 279, against a 100-turn certification target. The +dominant cost is not M9's at all: the **per-position narrative state document is +74% of a bundle**, and v2 already carried it. ## v3.3 — M8 Closeout (2026-09-06) diff --git a/planning/reports/M8-IMPLEMENTATION-REPORT.md b/planning/archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md similarity index 100% rename from planning/reports/M8-IMPLEMENTATION-REPORT.md rename to planning/archive/milestone-reports/M8-IMPLEMENTATION-REPORT.md diff --git a/planning/reports/M9-IMPLEMENTATION-REPORT.md b/planning/reports/M9-IMPLEMENTATION-REPORT.md new file mode 100644 index 0000000..bc91864 --- /dev/null +++ b/planning/reports/M9-IMPLEMENTATION-REPORT.md @@ -0,0 +1,1622 @@ +# M9 — Export, Backup, Recovery, and Migration Hardening + +**Implementation report, written for an independent reviewer.** + +This is a set of claims with the measurements attached. It is not a record of +acceptance: M9 is implemented and verified, and acceptance comes after review. + +Two conventions carried from M8, because they are what make a report checkable: + +- **Where a claim names a boundary, the boundary was crossed.** "A clean data + directory" is a second process against a database file that has never existed. + "A restart" is a different PID. "A backup" is opened as its own database and + read. "Migration from M8" is a database built by a server running the signed M8 + commit. Where a substitute was used, it is labelled. +- **Every defect found is numbered in §U, including the ones fixed before the + final tree.** Seven are recorded. Three were found by *running* M9's own tests + rather than by reading them, two more by building the browser harness, and + **four of the seven predate M9** — they were reachable in M8 and nothing had + gone looking. + +--- + +## A. Executive result + +**PASS.** + +Every required M9 condition is verified, by the boundary it names. The Definition +of Done — *a campaign can be safely exported, imported into a clean data +directory, and reopened at the exact intended active position with authoritative +history/state intact* — is satisfied and is measured three separate ways: in the +suite, across two real server processes with two data directories, and in a real +browser against a real narrator. + +| | | +| --- | --- | +| Bundle format | **`ai-dnd-adventure-v3`.** v1 and v2 still import; every seam from pre-active-head onward is tested | +| Historical prompt provenance | **carried.** The M8 handoff is closed | +| Story cards | **compatibility-only**, and out of the narrator's prompt | +| Summaries / memories | **carried**, with the coordinates that decide eligibility | +| Derived indexes | **not carried, and rebuilt** — with the rebuild path fixed | +| SQLite backup | **SQLite online backup API**, verified before it is kept | +| Schema change | **none**, proved against a database M8's own code wrote | +| Backend | **1,102 passed, 14 skipped, 0 failed** in 14:54 (M8 baseline: 950) | +| Frontend | 145 passed (was 132), lint and production build clean | +| Browser | **36/36 checks**, one frozen build, real narrator, two installations (§Q) | +| Docker | production image builds clean; the backup works inside it, on the mounted volume | + +Nothing is claimed here that was only compiled. §S maps every acceptance item to +the evidence that decides it, and names what is asserted from a synthetic +substitute. + +--- + +## B. Repository and provenance state + +| | | +| --- | --- | +| Branch | `m9-recovery` | +| Base | `1ce99727606124b0f8b3dc3b88842410c874ee14` — *M8: the browser becomes the storyteller* | +| Base signature | **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569` | +| Working tree at start | clean | +| `HEAD` at writing | still the M8 base — M9 is staged, not committed (§32) | +| Upstream ancestry | `d72f7c1bda0f34fccd84afb7a25c34eb01c901de` **is** an ancestor | +| `LICENSE` | unchanged; `git diff 1ce9972 -- LICENSE` is 0 lines | +| Schema version | **92 before, 92 after** | +| Migrations | **none added.** `git diff 1ce9972 -- backend/app/migrations.py` is 0 lines | +| Columns | **none added or removed.** 0 `mapped_column` changes in `models.py` | + +### Change inventory + +``` +backend/app/bundle.py 928 +- the v3 format, both directions +backend/app/backup.py 278 NEW — the online backup +backend/app/routers/backups.py 75 NEW — two endpoints +backend/app/routers/adventures/bundle_io.py 59 +- the two-phase import +backend/app/context/builder.py 63 +- story cards out of the prompt +backend/app/knowledge/fts.py 44 +- the index leak (finding 1) +backend/app/knowledge/importer.py 22 + clear_campaign_index +backend/app/schemas.py 19 + ImportedAdventureOut +backend/app/routers/adventures/crud.py 8 + clear the index before delete +backend/app/main.py 6 +- register the router +backend/app/starter.py 4 +- materialize returns a pair + +frontend/src/pages/Settings.jsx + the backup control +frontend/src/components.jsx +- the file picker (finding 4) +frontend/src/api.js, styles/library.css + + +backend/tests/m9_fixture.py 406 NEW — the portability fixture +backend/tests/test_m9_portability.py 1,186 NEW +backend/tests/test_m9_corrupt_bundles.py 689 NEW +backend/tests/test_m9_clean_import.py 460 NEW — two processes +backend/tests/test_m9_backup.py 451 NEW +backend/tests/test_m9_legacy_bundles.py 365 NEW +backend/tools/m9_portability_report.py 354 NEW — the baseline tool +backend/tools/m9_migration_proof.py 331 NEW +backend/tools/m9_scale_report.py 221 NEW +frontend/src/pages/backup.test.jsx 116 NEW +frontend/src/filePicker.test.jsx 85 NEW +``` + +Five existing test files changed. Each had pinned a gap M9 closes, and each was +**reworked to assert the new guarantee rather than deleted** — the rule +`BUILD-MILESTONES.md` records from M2 about moving instrumentation rather than +removing it. §U-6 lists them. + +--- + +## C. The M8 portability baseline, measured + +The brief required the baseline to be **measured, not assumed**. It was, with a +tool a reviewer can rerun on either commit: + +```bash +cd backend && python -m tools.m9_portability_report # or --json +``` + +It builds the M9 fixture (§E) in a throwaway database, exports it, imports the +result, and classifies every data family. Run on the M8 commit and on the M9 +tree, the difference is the milestone's claim in reproducible form. + +### The fixture it measures + +Deliberately not the standard Continuity Test: that one is shaped to read like a +story, and this one is shaped to break a round trip. 22 actions, 2 branches, 2 +Save Points on different branches, 1 superseded take, 12 state events, 12 +proposals, 11 stored prompts, 2 memories, 2 summaries (one on the abandoned +line, one at the head), 5 imported sources across all three classes including +one disabled and one narrator-only, and a manual state correction. + +The head ends **behind the retained tip of its own branch and behind the +abandoned line's**, and the last thing the fixture does is an Undo — so the head +is not the newest row written, not the deepest row, not the tip, and not on the +branch holding the most story. An importer guessing any one of those lands +somewhere else. + +### M8 — `ai-dnd-adventure-v2`, 29,630 bytes + +| Verdict | Family | +| --- | --- | +| PRESERVED | campaign identity, transcript, branches, branch disposition, active head, alternate takes, Save Points, narrative state (current), narrative state (per position), imported knowledge, memories, scene metadata, story cards | +| **OMITTED** | **take grouping** — SP9 parentage, so imported nodes landed parentless | +| **OMITTED** | **state events** — the audit half of §17's hybrid | +| **OMITTED** | **state proposals** — what the model asked for, and what was refused | +| **OMITTED** | **manual corrections** — indistinguishable from what the story established | +| **OMITTED** | **historical prompt/context** — the M8 handoff | +| **OMITTED** | **retrieval provenance** — which passages an old turn was shown | +| **OMITTED** | **per-turn model settings** — what a historical turn ran under | +| **OMITTED** | **knowledge parser versions** | +| **OMITTED** | **summaries** — only the lineage-less mirror column travelled | +| **OMITTED** | **memory authority** — a heuristic memory imported as accepted story | + +Two of those were visible from outside, through the API a reader reads: after a +round trip the copy's **state events** and **summaries** differed from the +source's. The rest were invisible until something needed them. + +Derived and correctly not carried: knowledge passages, the FTS index, knowledge +and memory embeddings, the branch lineage cache, derived status. + +### M9 — `ai-dnd-adventure-v3`, 90,117 bytes on the same fixture + +**Every family PRESERVED. No disagreement on any reader-visible family.** + +The file is 3.0× larger on this fixture; §R has the measurement on a long +campaign, where the ratio is very different and the conclusion is not the +obvious one. + +--- + +## D. The final bundle contract + +Documented in `DATA-MODEL.md` §29 and `TECHNICAL-DESIGN.md` §9.3; the reasoning +is repeated in `bundle.py`'s own docstring, which is where the next maintainer +will be standing. + +### Why the version was bumped + +Everything M9 adds *could* have been an optional key read with `.get`, the way +`persona`, `checkpoints`, `narrativeState` and `knowledge` each were — §9.1 and +§9.2 each considered a bump and correctly declined one. + +That mechanism stops working here, and the reason is the rule the format already +lives by. A v2 file with no prompt provenance is **ambiguous**: written before +M9, when no file could carry one, or by M9 from a campaign whose turns predate +the column? Those are different facts and a reader has to be able to tell them +apart — the same distinction I07 draws when it says a file written before the +head was carried opens at the tip *because tip was the only position that format +could represent*. **A version number is how a recovery file states what it was +capable of recording.** + +The family name is unchanged, deliberately: renaming it would break every reader +for no gain, and `PROVENANCE.md` is where the fork is recorded. + +```text +v1 a linear story, its turns, and its retries as a repeating group +v2 + the tree, the live flags, the after-snapshots, the chosen head, + Save Points, the narrative state document, imported knowledge +v3 + state events and proposals, historical prompt/context provenance, + lineage-anchored summaries, take parentage, memory authority +``` + +The importer reads all three. A version it has never heard of is **refused with +its own name in the message**, rather than read as the newest it knows. + +### Three categories, not two + +§31 of `DATA-MODEL.md` distinguishes authoritative from derived, which was +sufficient until M9 had to decide about stored prompts. They *are* derived — a +machine assembled them — and they must travel anyway: + +```text +chosen the story, the head, the takes, the Save Points, the + classifications, the canon travels +evidence the state events and proposals, the per-turn prompt and the + passages it was shown, the model and generation settings + that turn ran under travels +rebuildable knowledge passages, the FTS index, embeddings, the branch + lineage cache rebuilt on import +``` + +The test separating the last two is **not** "could this be recomputed" but +"would a recomputation answer the same question". A rebuilt FTS index answers +the same question. A rebuilt prompt does not — it says what the turn *would be +told now*, from today's canon, today's sources and today's state, which is the +opposite of what the inspector is for. **Historical evidence is not a cache**, +and M9 regenerates no prompt at any point in the import. + +### Required, optional, and what an absent key means + +| Key | Required | Absent means | +| --- | --- | --- | +| `format` | **yes** | refused | +| `actions` | in practice | an empty campaign | +| `branches`, `headBranch` | no | the root branch | +| `headDepth` | no | **the format could not say — open at the tip** (I07) | +| `checkpoints` | no | the campaign had none, or predates M4 | +| `narrativeState`, `narrativeStateAfter` | no | predates M5; empty document | +| `knowledge` | no | predates M7; empty library | +| `stateEvents`, `stateProposals` | no | **predates v3 — no audit trail is invented** | +| `contextSnapshotZ` / `contextSnapshot` | no | predates v3, or the turn has none | +| `summaries` | no | predates v3 | +| `id` / `parentId` on a node | no | pre-SP9 rule: `attempts.group` falls back to the coordinate | +| `authority` on a memory | no | the column default | + +### Integrity and validation + +One derived value is in the file, and only because its purpose is to be checked: +a knowledge source's `contentHash`. The import **recomputes** it from what +arrived, stores the recomputed value, and records the discrepancy on the source +where a reader can find it. The stated hash is never trusted and never silently +discarded. + +### The two pointers that are translated + +Branch numbers already were. M9 adds `source_id` inside a restored retrieval +record: it names a row on the machine that wrote the file, so left alone it +points the inspector's "open this source" at whatever holds that id here. It is +repointed at the source that landed, or set to `null` where the file carries no +such source — at which point the record still holds the passage's title, +filename and text. **The evidence is never rewritten; only the pointer is.** +`chunk_id` is deliberately untouched: passages are rebuilt and get new ids, so +no translation exists, and `chunk_index` and `heading_path` still say which +passage it was. + +--- + +## E. Story graph, head, and takes + +**Round-tripped and compared through the API a reader reads**, never by row +count. `test_m9_portability.py`, and again across two processes in +`test_m9_clean_import.py`. + +| Property | Evidence | +| --- | --- | +| Every accepted action, live and superseded | the whole exported tree is the same size either side | +| Branch and fork relationships | both branches, the fork depth, and the names | +| Active branch and head | §I07 below | +| Retained tip | the future past the head is in the database and Redo reaches it | +| Branch disposition | the superseded branch still carries the depth it was left at (6) | +| Retries / alternate takes | the superseded take is still at its coordinate; the selected one still selected | +| Which take is active | exactly one live node per coordinate, decided by the writer, not read from the file | +| Timestamps | `createdAt` restored per node, per branch, per Save Point, per event | + +### The undo-head case (I07) + +Exported with `active head < retained tip`, imported into a genuinely clean data +directory, and the copy opens **at exactly the exported head**. The later turns +are present as retained future, Redo is offered rather than the story having +silently been redone, and Redo then walks to the same next turn in both +campaigns. + +### The diverged case (I03) + +Both futures survive and stay distinguishable — the abandoned line's turns are +readable on their branch, and the branch that was left still records that it was +left, at the depth it was left at. There is no trimmed-export option, so I03's +clause about one does not arise. + +### The retry case, and a fidelity gap M9 closed + +v2 exported no take parentage, so every imported node landed parentless and +`attempts.group` fell back to the coordinate. That is right for a plain retry +and **wrong as soon as two takes of one turn each have takes of their own +beneath them**: those share a branch and a depth, so the copy read `5/5` where +the source read `2/2` and `3/3`. v3 carries the parentage, and the pager now +reads identically in the copy and the original. + +`test_retry_variants.py` had recorded the old behaviour as "a gap in the import"; +that test now asserts `2/2` (§U-6). + +--- + +## F. Save Points + +| | | +| --- | --- | +| Name, note, coordinate, `createdAt` | all restored; compared through the panel's own API | +| Resolution | every restored Save Point reports `resolved: true` | +| Restore | reaches the position it names, through **M3's head movement** — no second restore architecture | +| Later history | intact after restoring; the retained tree is the same size | +| Two Save Points, two branches | restore to **different** states, so they are not accidentally the same pointer | +| The head | unmoved by importing them: a campaign exported at turn 30 with a Save Point at turn 12 opens at turn 30 | + +### A Save Point the file cannot satisfy + +**Dropped, with the rest of the campaign kept — and never retargeted.** The +decision and its reasoning are in `bundle.py` and `TECHNICAL-DESIGN.md` §9.2: + +- not a refusal, because a bookmark costs a bookmark, where refusing the + campaign would lose the story to save the bookmark; +- not a repair, because the reader named a position, and if that position is not + in the file then **no other position is the one they named**. + +Three cases are covered: a coordinate beyond the retained story, a branch the +file does not list, and a name that is blank once trimmed. In each, the *other* +Save Point in the fixture survives — asserted, because "dropped" must not mean +"dropped them all". + +--- + +## G. Narrative state and the audit trail + +### What travels, and why each + +The audit inventory was made from the contract rather than by serialising every +table: + +| Record | Travels | Why | +| --- | --- | --- | +| `narrative_state` (campaign) | yes | the authoritative document at the exported head | +| `narrative_state_after` (per node) | yes | **the restore path.** Without it, head movement becomes replay | +| `state_events` | **yes (new)** | §17's audit half: what changed, where, who asserted it, what it was before | +| `state_proposals` | **yes (new)** | the inspector reads both — the event says what was accepted, the proposal says what was asked for and refused | +| rejected/unparseable proposals | yes | the record that explains why the state does not say what the narration seems to say | +| `derived_status` | no | describes the last run of a background pass, not the story | + +Exporting only the accepted half would have kept the answers and lost every +question, which is why both tables travel. + +### Verified after the move + +- Current authoritative state at the exported head: **identical**, compared as + the document. +- Historical positions: the copy and the source are walked back three turns and + forward three turns **in step**, and the document is compared at every + position. They agree throughout (L02). +- Restoration stays **O(1)-style**: every position's snapshot travels, so Undo, + Redo and Save Point restore remain one row read rather than a replay + (`TECHNICAL-DESIGN.md` §10.4, and the M4 note that made it load-bearing). +- **A manual correction is still identifiable as one.** This is the case the + omission hurt most: it is the one state change no narration explains, and with + the events gone nothing distinguished it from something the story established. +- The two tables arrive **linked**: events point at proposals that are in this + campaign, and proposals point at this campaign's own turns. + +### Refuse, repair, drop — decided per field + +| | Rule | +| --- | --- | +| **Refuse** | an audit record naming a turn the file does not contain. Unlike a Save Point, an audit record that quietly did not arrive leaves a campaign whose state cannot be explained — and the explanation is what a reader goes looking for exactly when something looks wrong | +| **Accept** | a proposal with *no* action: that is a manual correction, which has a coordinate and no narration behind it | +| **Keep the coordinate, drop the pointer** | an event whose *proposal* is missing. That is what the schema's `ON DELETE SET NULL` already says happens | +| **Normalise** | a malformed state document — M5's rule, unchanged: the story is the valuable thing, and a malformed section should cost the section | + +--- + +## H. Historical prompt and retrieval provenance + +**The M8 handoff (question B) is closed: the evidence travels.** + +`SPECIFICATION.md` §6.1 requires every accepted turn to record an *exact +prompt/context snapshot or reproducible equivalent*. The product already stored +one; the bundle did not carry it, so that guarantee was local to the machine +that played the campaign. It is now portable. + +Each node carries `contextSnapshotZ`: the assembled prompt section by section, +the passages retrieved **with the text each supplied**, which summary was +eligible, the token accounting, and the model and generation settings the call +ran under. Restored **verbatim** — nothing is regenerated, and nothing is +validated beyond its being an object, because imposing today's expectations on a +record an older build wrote would be the retroactive reading the inspector +exists to rule out. + +### Verified to survive each thing that could destroy it + +| | | +| --- | --- | +| A source disabled later | the record is unchanged | +| **A source deleted after import** | asserted directly: the sources an old turn was shown are deleted in the copy, and the turn still shows the same passages and the same text. The record holds the text, not a pointer to a row that can go away (`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50) | +| A different chunking on the importing machine | the record is not rebuilt from chunks, so it cannot move | +| Current state moved on | the snapshot is a record, not a re-render | +| Canon edited later | same | +| Summaries rebuilt | same | + +And across a real machine boundary: on machine B, which never assembled that +prompt and could not reassemble it, Inspect Context on an old narrator turn +returns a prompt with its sections intact. + +### Retrieval provenance + +Each `used` record keeps source title, filename, classification, visibility, +chunk index, heading path, the scores each path gave it, how it was admitted, +and the rendered text. `source_id` is translated (§D); everything else is +evidence and is untouched. + +### Cost control + +The prompt is stored **once per turn on the live attempt** (`attempts.py`), so a +superseded take carries only its own slices — its reply, its state proposal, its +token accounting. Asserted: no superseded take in the bundle carries `sections` +or `prompt`. A campaign retried twenty times does not carry twenty prompts. + +--- + +## I. Imported knowledge + +| | | +| --- | --- | +| Preserved | title, original filename (metadata only), content, classification, enabled, visibility, always-include, notes, media type, import timestamp, content hash, **parser and chunking version**, and a `sourceId` that exists only so historical retrieval records can be relinked | +| Rebuilt | passages, FTS rows, vectors — before the import returns for the first two, so the campaign is searchable immediately | +| Needs the original path | **no.** Verified with the exporting machine's process stopped | +| Disabled source | still disabled, and still out of retrieval | +| Narrator-only source | still narrator-only | +| Canon with always-include | still both | +| Content | byte-identical, by SHA-256, checked in the browser run too | + +### L04 — the rebuild, and the defect running it found + +On a **safe copy** (the imported campaign, not the original), every physically +derived structure is destroyed — passages, FTS rows, vectors — and rebuilt from +the content the bundle carried. Afterwards: every source `ready` with passages, +retrieval working, and the transcript, authoritative state, classifications and +lifecycle flags identical either side. Deleting the vectors alone leaves lexical +retrieval working, which is M7's rule that lexical is a production path and not +a fallback. A rebuild does not make an abandoned line's summary eligible. + +**Running it found finding 1** (§U), a pre-existing defect that made Reindex — +the documented repair — unable to repair the state it most needed to. + +### The M7 calibration rule + +Untouched. `nomic-embed-text` keeps its measured calibration and an uncalibrated +embedding model still degrades to lexical-only rather than borrowing the 0.58 +threshold; no M9 code path reads or writes the threshold, and the M7 calibration +tests pass unchanged. + +--- + +## J. Summaries, memories, and derived data + +### Summaries — carried, with their lineage + +v2 carried only `adventures.story_summary`, the **mirror column with no lineage +of its own**, so a restored campaign resumed with no usable long-story +continuity. v3 carries the rows: text, coordinate, source range, trigger, model, +timestamp. + +Eligibility is not a stored flag — it is whether the coordinate lies on the +active capped lineage — so restoring the coordinate restores the answer, +**including the "no"**. The fixture has one generated summary on the line it +abandons and one typed at the head it keeps, and after the move: + +- exactly one is eligible, and it is the one at the head; +- the abandoned line's text does not appear in the prompt the copy would send + next; +- the same holds after a full index rebuild, so **M6's leak does not return by + the rebuild path either**; +- `story_summary` is re-derived from the lineage after import rather than + trusted from the file, because a mirror disagreeing with the lineage is + exactly M6's finding M6-F1. + +### Memories — and an authority that was being promoted + +Text, pinned, forgotten, source range, use count, coordinate, and now +**`authority`**. Without it every imported memory landed on the column default, +which silently promoted a `heuristic` memory to `accepted_story` — the one +direction F07 forbids. An unreadable value is read as `heuristic`, which is the +safe direction: a demoted memory costs a little ranking weight, a promoted one +puts a guess in front of the narrator as fact. + +### Vectors + +Not carried, in either subsystem, and rebuilt against whatever embedding model +*this* machine has. Exporting them would tie a campaign to one machine's model, +which is the reason M7 gave and M9 keeps. + +--- + +## K. Story cards — the compatibility decision + +**Settled: compatibility-only legacy data, and no longer in the narrator's +prompt.** Recorded in `IMPORTED-KNOWLEDGE-DESIGN.md` §73. + +### What was found + +Story cards **were** the "alternate untracked path around the new knowledge +authority/provenance rules" that §73 already forbade, and not in principle. A +keyword-matched card was injected into the narrator's prompt as +`World Lore: `, taking up to 40% of what remained after the imported +knowledge had been placed, with: + +- no class, so nothing framed how far the narrator could rely on it; +- no visibility, so no narrator-only distinction existed; +- no source, no hash, no lifecycle — nothing to disable; +- **no browser surface at all since M8 removed the editor**; +- and no row in the context inspector, which renders `knowledge` and never + rendered `cards`. + +It also competed with imported Canon for one budget, which is the arrangement M7 +spent a milestone separating. + +### What happens now + +| | | +| --- | --- | +| **New export** | carries them, unchanged, under `storyCards` | +| **Legacy import** | accepted unchanged, from every format version | +| **Normal narration** | **not reached.** The `world_lore` section is gone | +| **Re-export** | carries them again — a round trip destroys nothing | + +Nothing is deleted: the rows stay, the `/api/story-cards` endpoints stay, and +`memorybank.cast_brief` still reads them as the **summariser's character +roster** — that names who is on stage so a memory says "Aldric" rather than +"he", never reaches the narrator, and every memory written from it is +authority-classified by the application afterwards. The `cards` key stays in the +context report and is now always empty for a new turn, because M9 has just made +historical snapshots portable and an old turn's record must go on saying that +story cards were included. That is asserted. + +This is the **smallest safe change**: it closes the one path that asserted +campaign facts to the narrator without any of §73's controls, and touches +nothing else. + +--- + +## L. Scene and media metadata + +**No media tables exist in this build, and none were invented.** Two active +documents say that is the right answer rather than a gap: + +- **K04's second branch** — "if media tables are deferred, architecture/types + should demonstrate equivalent extension point". +- **`MEDIA-EXTENSION-CONTRACT.md` §58**, on export, in as many words: *"For v1, + if media is not implemented, preserve schema compatibility."* + +M9 preserves it in the strongest available form: the format is now **versioned**, +so M10 can add media sections as v4 and every v3 reader keeps working, and every +v3 file keeps importing. Inventing an empty `media_assets` table to satisfy the +word "metadata" would have been schema M10 then had to live with. + +What does exist is the `scene` section of the authoritative narrative state +document, which is `SPECIFICATION.md` §14's scene snapshot as this build +represents it. It is campaign-scoped and per-position, and it round-trips with +the document: asserted both on the campaign-level state and on the per-node +snapshots. + +Nothing in M9 adds image, video, TTS, STT, a media provider, or a job +architecture. + +--- + +## M. Model and context-window portability + +The M8 handoff's question D, answered by keeping two things apart. + +### What travels + +**Per-turn model and generation settings**, inside the historical snapshot — +model name, API mode, temperature, max output tokens. There they are a *record +of what happened*, which is exactly §6.1's requirement. + +**Campaign-owned configuration**: title, narrator instructions, persona, canon, +memory/summary switches, the campaign's own plot fields. + +### What does not, and why that is the point + +The `settings` row — endpoint, model, context budget, timeouts, embedding model. +Those describe the **machine**, not the campaign. Asserted directly on the +importing machine: after importing a campaign played against a configured +endpoint, machine B's `endpoint_url` is still its own default and its +`context_token_budget` is still 16,384. **Importing a campaign is not a way to +reconfigure the destination's inference.** + +### Import does not depend on a model + +Asserted: machine B has **no model configured at all**, and the campaign still +imports, opens, and shows its state and transcript. Playing on would fail; +recovering does not. *Campaign restored successfully* is not *the destination +narrator has the same capacity*, and M9 does not treat it as such. + +### The ceiling, and what M9 does about it + +No assumption that 16,384 tokens are available is hard-coded anywhere M9 +touched, and no native context-window detection architecture was built — that is +M11's, with the 100-turn certification. + +What M9 adds is the sentence a deployer needs, in `DEVELOPMENT.md`: a long +imported campaign fills a prompt on its **first** turn, so a machine that has +applied neither the derived-model procedure nor a matching budget meets its +ceiling immediately rather than gradually. Check `/api/ps` on the destination +*before* playing on an imported campaign. + +--- + +## N. SQLite backup + +`backend/app/backup.py`, `backend/app/routers/backups.py`, and a control in +Settings → Advanced. + +### Why not a file copy + +A copy taken while the application runs is a copy of a moving target: SQLite +writes in pages, and a plain copy can read page 5 before a transaction and page +900 after it, producing a file that opens, reports a schema, and is quietly +missing rows. Nothing warns anyone, and it is found later. + +So it uses **SQLite's online backup API** (`sqlite3.Connection.backup`), which +copies under the right locks and yields a transactionally consistent snapshot of +a committed point. **The application keeps running**; no session is closed and +no turn is blocked. + +### The procedure's guarantees, each tested + +| | | +| --- | --- | +| Source opened **read-only** | via a `mode=ro` URI, so it is a guarantee and not an observation. The source file is byte-identical after a backup | +| Temporary destination, finalised by rename | the `.partial` file is gone and the final name never wore it | +| `PRAGMA quick_check` before the rename | a failed verification leaves **nothing** behind | +| Never overwrites | two backups get two names; two in the same second get two names | +| Failure reported | a clear error with the reason; the source untouched | +| No caller-supplied path | the endpoint accepts **no request body at all** — asserted against the OpenAPI schema, and by sending one anyway and checking where the file landed (H08) | + +### Opened independently, and read + +Every backup test **opens the copy as its own database** and asks what is in it — +a test that only checked a file appeared would pass against a `cp`. Verified +present in the copy: the schema, campaigns with their titles, story rows, the +head branch and depth **naming a turn that exists in the copy**, Save Points, +knowledge sources, state events, and the authoritative state document decoded and +compared against the live campaign's. + +### Taken under load + +The test that separates this from a file copy: turns are played from a second +thread **throughout** the copy. The resulting file passes `quick_check`, has no +dangling foreign key, has a transcript with **no gap in its depths**, and a head +that names a turn present in it. Which committed point it caught is not asserted +— that would be asserting on a race. + +### No restore endpoint, and no download + +Both are decisions, stated in the code and in the panel. Restore means replacing +the file the running process has open, which is how both copies are lost at +once; the stop-move-start procedure is in `DEVELOPMENT.md`. Download would put a +copy of every campaign into the browser's download directory and cache, which +for an application whose premise is that the story does not leave the machine is +a worse default than a path the reader can copy. + +--- + +## O. Import validation and corrupt bundles + +`test_m9_corrupt_bundles.py`. Every case starts from a **real export of the M9 +fixture** and breaks exactly one thing, so what each test measures is that break +rather than a hand-written shape nothing ever wrote. + +### Two rules, checked rather than assumed + +**Nothing lands.** A refused import is checked by counting rows in **eleven +tables** before and after, not by trusting the status code. No campaign, no +branch, no orphan action, no Save Point pointing at nothing, no knowledge owned +by a campaign that does not exist. + +**Nothing is fetched, read or run.** A URL in a bundle stays text; a path stays +text. No test needs a network guard to pass — there is no code path that would +use one. + +### Coverage + +| Class | Cases | +| --- | --- | +| Format / version | not an object; empty object; missing `format`; `format` of four wrong types; an unsupported future version, refused **with its own name in the message** | +| History graph | action on an unlisted branch; branch forking from one listed after it (which is also **how a cycle is made impossible rather than detected**); a branch forking from itself; a fork with no depth; an action with no depth; a negative depth; **two actions claiming one identity**; a turn with no live take; a parent naming a node the file lacks; **a node that is its own parent** | +| Head | past the retained story (refused, message names where the branch ends); four wrong types; a head branch that is not listed (repaired to the root, and the reasoning for why *that* one is repaired) | +| Save Points | coordinate beyond the story; branch not listed; blank name; the whole section not a list | +| State | sections not lists; an event that is not an object; an event with no type; **an event naming a turn the file lacks (refused)**; a proposal naming a turn the file lacks; an event whose proposal is gone (keeps its coordinate); a malformed state document; a malformed per-position snapshot | +| Knowledge | section not a list; no content; unknown classification; a source that is not an object; unreadable visibility; **content hash mismatch, recomputed and reported**; over the source cap; over the size cap | +| Provenance | a snapshot that is not an object (dropped, never guessed at); a nonsense knowledge block; a snapshot naming an impossible source (relinked to `null`, evidence intact) | +| Summaries | no coordinate (dropped, **never placed at a guess** — that is E03's leak by a new route); section not a list | +| Caps | actions, branches, and a body past the import ceiling (413 from the middleware, on the declared length, before any parse) | +| Inert data | URLs, `file://`, remote images; four traversal filenames; a title that looks like a command; an over-long title | + +### The authoritative/derived distinction on failure + +Two outcomes a caller can see, and they are distinguishable: + +```text +4xx the authoritative import failed and no campaign exists +201 the authoritative import succeeded and the campaign is complete +``` + +A **rebuildable** index failing is neither, and does not roll the campaign back: +it is reported on the 201 as `import_warnings`, visible per source in the +Knowledge panel, and repaired by Reindex. The response model is a subclass +(`ImportedAdventureOut`) rather than a field on `AdventureOut`, because "which of +your indexes failed to rebuild" is a fact about one import, not a property of a +campaign. + +The rollback is **explicit** rather than left to the session closing, so a test +can assert on it rather than on teardown. + +--- + +## P. Migration and backward compatibility + +### No schema change, proved + +``` +git diff 1ce9972 -- backend/app/migrations.py → 0 lines +mapped_column additions/removals in models.py → 0 +PRAGMA user_version → 92 before, 92 after +``` + +`git diff` is necessary and not sufficient — a migration can also be *missing*, +and the failure then is a database that opens and quietly answers wrongly. So +the claim is proved the way M8 proved its own: + +**A campaign is built by a server running the signed M8 commit**, from the git +worktree at `1ce9972`, using M8's own interpreter and M8's own code: story, an +alternate take, a Save Point, a manual state correction, memories, a summary, +three imported sources including a disabled and a narrator-only one, per-turn +context snapshots, and two Undos so the head is behind the tip. M9 then opens +that file. + +``` +$ python -m tools.m9_migration_proof open --expected +{ "problems": [], + "schema_version_before": 92, "schema_version_after": 92, + "schema_version_second_open": 92, + "checked": { "families": 12, "snapshot_rows": 8, + "counts": { "actions": 16, "checkpoints": 1, + "knowledge_chunks": 3, "knowledge_sources": 3, + "memories": 2, "state_events": 9, + "state_proposals": 9, "summaries": 1 } } } +``` + +Twelve families compared and identical; **a second open migrates nothing** +(idempotence); and the M8-built campaign then exports as v3 and round-trips, and +Redo still works on it. Neither half of the tool imports the other's code — what +crosses is the database file and one JSON report. + +### Older bundle seams + +`test_m9_legacy_bundles.py` ages a real v3 export backwards, removing only what +each era's format genuinely could not carry — a checked-in fixture drifts, and a +hand-written one tests a shape nothing ever wrote. + +| Era | No story lost | Nothing invented | Head | +| --- | --- | --- | --- | +| pre-M9 (v2) | ✓ | no events, no proposals, no summaries, no prompts | honoured, as v2 carried it | +| pre-M7 | ✓ | + no knowledge sources, no chunks | honoured | +| pre-M5 | ✓ | + empty state, no canon | honoured | +| pre-Save-Point | ✓ | + no Save Points | honoured | +| pre-active-head | ✓ | + no Save Points | **opens at the tip**, offers no Redo | +| v1 (flat) | ✓ | both attempts arrive, one live | tip | +| pre-M2 (with `scripts`) | ✓ | scripting keys ignored, not rejected | as its era | + +And each era can still be *used*, not merely read: a pre-M5 campaign gains state +from its next turn, a pre-M7 campaign imports a source and it indexes, a +pre-Save-Point campaign can be given one. + +**A pre-M9 campaign re-exports as v3 with its evidence sections empty**, because +the campaign genuinely has none — which is what lets a later reader trust a v3 +file's empty `stateEvents` to mean "this campaign has no audit trail" rather +than "the file could not say". + +--- + +## Q. Browser recovery workflow + +**36/36 checks, zero failures**, on one frozen production build against a real +narrator across two installations. + +### Environment + +| | | +| --- | --- | +| Browser | Firefox 154.0.1, headless, driven over W3C WebDriver (geckodriver 0.37.1) | +| Application | the **frozen production build** `dist/assets/index-DcHbz7ga.js`, served by the real backend as static files — not a dev server | +| Backend | two `uvicorn app.main:app` processes on loopback, one per machine, each with its **own data directory** | +| Databases | machine A: a real campaign. Machine B: **a file that had never existed**, created by the migrations on first start | +| Narrator | `qwen2.5:3b-instruct-16k` on the trusted-LAN Ollama, over HTTPS with a privately issued certificate — **real inference for every turn** | +| Campaign | played over real HTTP: five turns, a retry, a Save Point, three imported sources (one disabled, one narrator-only), a manual state correction, and two Undos so the export is taken behind the retained tip | + +### The §25 workflow, item by item + +| | | +| --- | --- | +| **1. Export campaign** | **PASS.** Clicked in the browser; the page produced a **60,160-byte `ai-dnd-adventure-v3`** bundle | +| **2. Return to Campaigns** | **PASS** | +| **3. Import into a clean target** | **PASS.** Machine B had never had a database, its library was empty, and it had **no narrator configured at all**. The import reported no incomplete rebuild | +| **4. Open imported campaign** | **PASS** | +| **5. Opens at the exact active position** | **PASS.** The transcript matches the source exactly; Redo is offered; the retained future is in the database (**11 rows behind 7 on screen**); and all 4 narrated turns are found in the rendered page | +| **6. State is correct** | **PASS.** The authoritative document matches, the manual correction is in the restored audit trail, with the fact the reader asserted, and the whole trail arrived | +| **7. Save Points present** | **PASS**, in the browser's own panel, with the note | +| **8. Knowledge present, with classifications and status** | **PASS.** All three sources in the panel; the disabled one still disabled, the narrator-only one still narrator-only, every source reindexed and searchable, and the content **byte-identical by hash** | +| **9. Historical Inspect Context works** | **PASS.** 3 of 3 restored narrator turns show the prompt they were given; 3 name the passages they were shown; and the inspector opens through the browser's own control | +| **10. Undo/Redo coherent** | **PASS.** Both offered at the restored head; Redo moves forward into the retained future | +| **Backup control** | **PASS.** Taken from Settings; the panel names the file; the file is on disk where it says; it **passes its own integrity check opened alone**, and holds the restored campaign with all 11 of its rows | +| **I06 on the downloaded file** | **PASS.** No credential, no endpoint, no path from the exporting machine | + +### Two substitutions, and both are on the browser's side of the line + +Named precisely, because §30 of the brief is about not letting a substitute +stand in for the boundary it bypasses. + +1. **The export's write to disk.** `downloadJSON` builds a Blob, makes a `blob:` + URL and clicks an anchor. **Headless Firefox does not complete that download + in this configuration** — the click fires, the page reports "Campaign + exported.", and no file appears in the download directory, in `~/Downloads`, + or anywhere under the snap's own tree. So `URL.createObjectURL` was wrapped + to capture the exact bytes the page handed the browser, and those bytes were + written to disk and used for everything downstream. *Proved:* the browser + assembled the correct v3 file and offered it. *Not proved:* Firefox's own + download writer, which is browser behaviour. +2. **The import's read of the file.** Firefox here is a **snap**, and its + sandbox will not let the content process read a path outside its + confinement. WebDriver hands the input the file successfully — the `File` + arrives with the right name and **the right byte count**, and `change` fires + — and `FileReader` then fails with `NotFoundError`. The click and the + handover are exercised; the last hop, the POST the page would make with the + bytes it read, was made against the same endpoint the page calls. + +Neither is papered over, and neither leaves the end-to-end path unproven: +`test_m9_clean_import.py` does the whole of it — export, a real file on disk, +a real HTTP import into a second process against a database that never existed +— with no browser in the way, 11/11. + +There is no unconfined Firefox and no Xvfb on this machine, so a windowed run +was not available. That is a property of the machine and is recorded as such. + +### One observation worth a reviewer's attention + +**The real narrator emitted no typed state blocks across five turns.** +`qwen2.5:3b-instruct-16k` narrated well and never produced a `state` block, so +the campaign's only state event is the **manual correction** — which is why the +audit-trail checks above read "1 events, source had 1". + +That is not an M9 defect and not a regression: it is the behaviour ADR 003 and +ADR 010 exist for. The application owns state and the model does not, C04's +manual correction path is how a reader asserts what the model did not, and the +per-position snapshots keep the head, the transcript and the state in agreement +regardless. It does mean the browser run's state evidence is thinner than the +suites', where a scripted narrator produces twelve events and two proposals — +so the rich state case is proved there and the *real-model* case is proved here, +and this report does not let the second stand in for the first. + +--- + +## R. Performance and size + +### The M9 fixture + +| | M8 (v2) | M9 (v3) | +| --- | --- | --- | +| Bundle | 29,630 B | 90,117 B | +| Actions | 22 | 22 | +| Source database | 258 kB | 319 kB | +| Export | 0.028 s | 0.044 s | +| Import (write + rebuild) | 0.063 s | 0.119 s | + +### A long campaign — and the result that changed the plan + +`python -m tools.m9_scale_report --turns 150 --budget 16384`. A real prompt +builder, a 16,384-token budget, and prose long enough that a turn is a turn. + +``` + turns actions bundle B snapshots B B/turn export s import s + 25 51 311,346 130,766 5,230 0.066 0.171 + 50 101 883,616 279,208 5,584 0.129 0.335 + 75 151 1,715,306 439,634 5,861 0.269 0.232 + 100 201 2,870,121 675,880 6,758 0.572 0.371 + 125 251 4,313,937 951,870 7,614 0.518 0.603 + 150 301 6,022,430 1,242,324 8,282 1.012 1.109 +``` + +A per-turn prompt contains the story so far, so carrying one per turn is +**O(turns²)**. That was the risk in the decision, and measuring it changed what +was done about it and what is claimed: + +- **Uncompressed, the snapshots were 68% of a 9.7 MB file at 120 turns.** They + now travel compressed inside the bundle — the same JSON through the same + `compression.pack`/`unpack` the database column already uses, base64-encoded + so the file is still JSON. **7.9× smaller**, and every other section of the + file is still plain readable text. +- At M11's 100-turn certification target: **2.87 MB, 14% of the cap.** +- Export and import stay under 1.1 s at 150 turns. + +### Where the bytes are, and what M9 actually cost + +The same tool's per-section breakdown, at 150 turns. It exists because "the +prompts are 21% of the file" does not answer "what did M9 add" — the events, the +proposals and the summaries are v3 additions too, and a cost claim counting only +the prompts would understate it. + +``` + per-position state (v2 already) 4,474,294 74.3% + prompts (contextSnapshotZ) 1,245,144 20.7% + state proposals 78,938 1.3% + state events 41,870 0.7% + node ids + parentage 6,993 0.1% + summaries 3,958 0.1% + --- everything v3 added 1,376,903 22.9% + a v2 file of the same campaign 4,645,347 + ceiling with v3 additions: ~279 turns + ceiling without them: ~318 turns +``` + +**The result is not the one the decision was worried about.** The largest thing +in a campaign bundle is not the prompts M9 added — it is the **per-position +narrative state document, at 74% of the file, which v2 already carried**. It +grows because a state document accumulates facts, so it is quadratic for the +same reason the prompts are, and it has been since M5. + +Everything M9 added comes to **22.9%** of the file, and moves the import ceiling +from about **318 turns to about 279** — a cost of roughly **12% of reachable +campaign length** for closing the provenance gap, not the halving it looked like +before it was measured, and not the 11% a prompts-only reading would have +claimed. + +Export and import stay under 1.2 s at 150 turns. + +### N+1 and duplication + +Checked for, and none found: + +- The export undefers every deferred per-node column in **one** query, including + `context_snapshot`, and the proposals' `detail` once for the whole list. +- Nodes are referenced by their own exported ids, so the audit records need no + per-row lookup. +- Knowledge content appears **once** per source; the prompt appears **once** per + turn, on the live attempt, with superseded takes carrying only their own + slices (asserted). +- The import's flushes are bounded by the number of **knowledge sources**, not + by the size of the story: one after the nodes, one after the proposals, two in + `materialize`, and one per source — the last because `build_index` needs the + source's id, which is the same one-flush-per-source the live upload path + already does. Nothing flushes per action, per branch, per event or per + Save Point, which is where a long campaign would have felt it. + +### Residual + +The ceiling above is a **residual limit, stated with its measurement**. It is far +beyond M11's certification target, and the fix if a later milestone needs one is +a streaming or chunked import, which is architecture beyond M9. The asymmetry +worth naming for a reviewer: a campaign past that length can still be +**exported** and would be refused on **import** with a 413. + +--- + +## S. Acceptance matrix + +Every M9-owned item, with the evidence that decides it. Where a claim names a +boundary, the row says which boundary was actually crossed. + +### I-series + +| | Result | Evidence, and the boundary | +| --- | --- | --- | +| **I01** Export campaign | **PASS** | Every section present and non-empty on the M9 fixture, checked family by family rather than by file size. `test_i01_*`; reproducible with `tools/m9_portability_report` | +| **I02** Import exported campaign | **PASS** | **Second process, second directory, database file that never existed, exporting process stopped.** Transcript, state, canon, Save Points, knowledge, events and retained-tree size all match. `test_m9_clean_import.py`. Same round trip in-process in `test_m9_portability.py`, labelled there as the weaker of the two | +| **I03** Branch / disposable history | **PASS** | Both futures present and readable; the abandoned branch still records the depth it was left at; the superseded take still at its coordinate. No trimmed-export option exists, so that clause does not arise | +| **I04** Checkpoint export | **PASS** | Names, notes, coordinates; every one resolves; each restores to its own position and its own state; later history intact. A Save Point the file cannot satisfy is dropped, never retargeted | +| **I05** Knowledge provenance | **PASS** | Content byte-identical by SHA-256, classification, enabled, visibility, always-include, filename, parser/chunking versions. Verified **with the exporting machine stopped**, and in the browser | +| **I06** No API secrets | **PASS** | Searched against the file's **text**, with the inert `api_key` column written first so absence is evidence. Compressed snapshots decoded, not skipped. Checked in the unit suite, the two-process run, and the browser run | +| **I07** Undone active head | **PASS** | All three clauses. The exact head **across a machine boundary**; a head-less file opens at its tip with no Redo; a head past the story is refused with a message naming where the branch ends | + +### L-series + +| | Result | Evidence | +| --- | --- | --- | +| **L01** Atomic turn commit | **PASS (not regressed)** | M9 changed no turn-commit path; the inherited tests pass unchanged. M9 adds the import's own version, in two forms: a **refused** import writes zero rows across eleven tables, checked by counting; and a failure **deep inside the write phase** — past every planner check, with the branches, nodes, parentage, memories, head and Save Points already in the session — also leaves zero rows, which proves the transaction rather than the planner. The rollback is explicit, so the next import succeeds rather than inheriting a poisoned session | +| **L02** State reconstruction | **PASS** | Measured **after a round trip**: copy and source walked back three turns and forward three turns in step, document compared at every position. Restoration stays a snapshot read, not a replay | +| **L03** Save Point after restart | **PASS** | **Three processes.** Export from A, import into a clean B, restore a Save Point in B, kill B, start a third against the same file, ask again. Transcript, state and retained-tree size all match | +| **L04** Derived data rebuilt | **PASS** | Passages, FTS rows and vectors destroyed and rebuilt on a copy; authoritative campaign identical either side; retrieval works again; vectors alone gone still leaves lexical working. **Running it found finding 1** | + +### M9-specific round trips + +| | Result | +| --- | --- | +| Undone-head round trip | **PASS** (I07) | +| Branched campaign round trip | **PASS** (I03) | +| Checkpoint round trip | **PASS** (I04) | +| Knowledge provenance round trip | **PASS** (I05) | +| Derived indexes can be rebuilt | **PASS** (L04) | +| **Historical prompt/context round trip** | **PASS** — survives the source being deleted, in the copy | +| **State audit round trip** | **PASS** — events, proposals, and the manual correction still identifiable | +| **Summary lineage round trip** | **PASS** — the abandoned line's summary is still ineligible, before and after a rebuild | +| **Take parentage round trip** | **PASS** — the pager reads identically in copy and source | +| **Memory authority round trip** | **PASS** — a heuristic memory is not promoted | +| **Legacy round trips** | **PASS** — v1, v2, and four earlier seams; nothing invented at any of them | +| **Migration from an M8-built database** | **PASS** — 12 families, 0 problems, version 92→92→92 | + +### Adjacent items re-verified rather than assumed + +| | Result | +| --- | --- | +| **E03** abandoned summary cannot leak | **PASS** after a move, and after an index rebuild | +| **F07** heuristic memory is not canon | **PASS** — the authority survives the move | +| **C04** manual correction | **PASS** — still identifiable as one in the copy | +| **H08** path traversal | **PASS** — four shapes sanitised; the backup endpoints accept no path | +| **H09** ZIP slip | **N/A** — no archive format was introduced; the bundle is one JSON document | +| **G04** disabled source stays out of retrieval | **PASS** after the move | + +### Synthetic substitutes, labelled + +| Where | Substitute | Why it is sound | +| --- | --- | --- | +| Every backend suite | `ScriptedProvider` narrator, stub embedder/summariser | The property under test is the round trip, not the model. Both derived factories are stubbed, per M6's finding M6-F3 | +| `test_m9_clean_import.py` | the spawned server's deterministic narrator | Real processes, real HTTP, real files; only the model text is canned | +| Legacy-bundle seams | v3 exports aged backwards | A checked-in fixture drifts and a hand-written one tests a shape nothing wrote | +| **Browser run: the export's write to disk** (§Q) | the bytes the page produced, captured from `URL.createObjectURL` | Headless Firefox does not complete a `blob:` download here. Proves the browser assembled and offered the right file; does **not** prove Firefox's download writer | +| **Browser run: the import's read of the file** (§Q) | the POST the page would make, made against the same endpoint | Firefox is snap-confined and cannot read a WebDriver-uploaded path — the `File` arrives with the right size and `change` fires, then `FileReader` raises `NotFoundError`. The click and handover **are** exercised | +| **Not substituted** | the browser run's narrator and its two installations, the migration proof's M8 build, the backup's SQLite, every process boundary, the file-on-disk import in `test_m9_clean_import.py` | Each is the boundary the claim names. The end-to-end file path the browser could not complete is proved there, with no browser in the way | + +--- + +## T. Security and local-only + +| | | +| --- | --- | +| Internet | none. No M9 code path opens a socket. Export, import and backup are file and database operations | +| Embedded URLs | inert. A URL in a bundle is stored and displayed as text | +| Remote images | unchanged; CSP still `img-src 'self' data:` | +| Shell | none. No `subprocess` in any M9 application file | +| Imported content executed | never. It is stored and rendered through M8's safe Markdown path | +| Loopback | unchanged. The backup endpoints are on the same loopback API | +| Inference endpoint policy | untouched. `endpoints.py` unchanged | +| Cloud storage / remote backup | none. The backup writes to a directory beside the database | + +### I06 + +Tested against the file's **text**, not against a list of columns — a field +added to a model the exporter walks would otherwise reach the bundle with no +column test noticing. The inert `api_key` column is written with a recognisable +value first, so its absence is evidence rather than a tautology. Searched for +and absent: the written secret, `api_key`, `apiKey`, the endpoint and its port, +and absolute filesystem paths. The compressed snapshots are **decoded** before +searching rather than skipped. + +Checked in three places: the unit suite, the two-process run (against a real +endpoint URL), and the browser run. Measured on the M9 fixture, over 90,074 +bytes of bundle **plus 154,279 bytes of decoded snapshot**: + +``` + 'SECRET-MUST-NOT-TRAVEL' present: False (written into api_key first) + 'api_key' / 'apiKey' present: False + the endpoint host and port present: False + 'endpoint_url' present: False + '/home/' , '/tmp/' present: False + 'secret.key' present: False +``` + +What **is** present, and correctly: the per-turn model name inside each +historical snapshot. That is a record of what the turn ran under, which §6.1 +requires — not a configuration the import applies. §M has the distinction. + +### File, path and archive safety + +**No archive format was introduced**, so H09 does not arise: the bundle is still +one JSON document, and the compression in §R is an encoding of one field inside +it, not a container. + +- No bundle-supplied value becomes a read or write path. `originalFilename` is + metadata; four traversal shapes are sanitised on import and asserted absent. +- Knowledge restores from the content in the bundle, never by reopening the + source machine's path — verified with the exporting process stopped. +- The backup endpoints accept no path at all. + +--- + +## U. Findings + +Numbered, including those fixed before the final tree. + +### 1. Deleting a campaign leaked its lexical index, and broke the next import — PRE-EXISTING (M7), FIXED + +**Severity: high.** An ordinary upload in an unrelated campaign failed with a +500. + +The FTS5 index is a virtual table, so no foreign key reaches it and no +`ON DELETE CASCADE` covers it. Deleting a campaign cascaded +`knowledge_sources` → `knowledge_chunks` and stopped there, leaving one index row +per passage pointing at a chunk that no longer existed. Nothing read them — every +search joins through `knowledge_chunks` — so the leak was **invisible** until +SQLite handed the freed primary key out again, at which point the next source +imported into **any** campaign collided on `INSERT` and raised. + +Worse, `clear_index` finds index rows *through* the chunks, so with the chunks +gone the orphans were unreachable: **Reindex, the documented repair, could not +repair it.** + +Found by **running** L04 — it failed on the rebuild with an integrity error, +which is the only way this shows itself. Fixed at both ends: `importer.clear_campaign_index` removes +a campaign's index rows before it is deleted, and `fts.add` uses `INSERT OR +REPLACE` — the rowid is a chunk's primary key, so a row already there is by +definition stale. The second half means **a database already carrying the leak +repairs itself, with no migration**. Three regression tests. + +### 2. An imported node with no state snapshot was given the head's state — PRE-EXISTING, FIXED + +**Severity: medium.** M5's review finding 3, arriving through the import. + +`tree.stamp_outcome` runs on every flush and fills a missing +`narrative_state_after` from *the campaign's current state*. On an import that is +the state at the exported head — so every position whose snapshot the file did +not carry came back holding the newest position's state, and an Undo to turn 2 +showed what the story knew at turn 20. + +`_write_nodes` now writes the empty document explicitly when the file carries +none. Legacy behaviour is unchanged: a pre-M5 bundle has no `narrativeState` +either, so the fallback was already writing `empty()` for those files. + +### 3. The snapshot relink did not persist — FIXED + +**Severity: high, and it would have shipped looking correct.** + +`_relink_snapshots` edited the dict the attribute already held and assigned it +back. `context_snapshot` is a plain `CompressedJSON` column, not a +`MutableDict`, so SQLAlchemy tracks it by assignment: at flush the loaded value +and the current value were the same object, the history reported no change, and +no `UPDATE` was emitted. The import looked right in memory and wrote the +untranslated ids to disk. + +Found because the test asserted the **outcome** — that the restored records name +this campaign's sources — rather than that the function was called. It now builds +a new document. + +### 4. Cancelling the import file dialog hung the Import button — PRE-EXISTING, FIXED + +**Severity: low, and user-visible.** `pickJSONFile` resolved on `change` only, so +closing the picker without choosing anything never settled the promise: the +Campaigns screen's `await` never returned, its `finally` never ran, and the +Import button stayed disabled reading "Importing…" until the page was reloaded. + +Found while making the input reachable for the browser suite. The screen's own +comment already said "a cancelled file picker is not a failure worth a message" — +it had simply never received one. Now `oncancel` rejects with an empty message, +which is exactly what that comment describes. + +### 5. The file input was detached from the document — FIXED + +**Severity: none for a user; structural for testing.** `pickJSONFile` created an +input and clicked it without appending it, so no element existed for a test or +for WebDriver to hand a path to — the import workflow could only ever be checked +by calling the API underneath it, which is not the workflow. It is now in the +document and removed however the promise settles. Six tests. + +**And a second lesson from fixing it.** The first attempt marked the input +`hidden`, which reads correctly and is wrong: a `hidden` element is +non-interactable, and WebDriver will set `files` on one **without dispatching +`change`** — the file lands and nothing happens, which is a worse failure than +the detached input because it looks like it worked. It now uses the ordinary +visually-hidden pattern — off-screen, zero-sized, `aria-hidden`, out of the tab +order — so no reader meets a stray "Choose file" control while the browser's own +dialog is what they are looking at. The test asserts `hidden === false` and says +why, so the next person does not "tidy" it back. + +### 6. Five existing tests pinned the gaps M9 closes — REWORKED, NOT DELETED + +Each asserted the M8 behaviour that was the debt. Following the rule +`BUILD-MILESTONES.md` records from M2 — move the instrumentation, do not delete +the test — each now asserts the new guarantee, and each carries a note saying +what it used to assert and why it changed. + +| Test | Was | Now | +| --- | --- | --- | +| `test_historical_prompt_evidence_survives_an_export_round_trip` | the copy's turn has **no** snapshot (404) | it has the same snapshot, and its `source_id`s name this campaign | +| `test_export_and_import_round_trips_variants` | the pager reads `1/1` on the copy, "a gap in the import" | it reads `2/2`, as the source does | +| `test_an_unknown_format_is_refused` | used `v3` as its "from the future" placeholder — which M9 made real, so it started importing the file it meant to reject | uses `v99`, and asserts against `bundle.FORMAT` rather than a literal | +| `test_export_carries_the_whole_story` | `format == "ai-dnd-adventure-v2"` | `format == bundle.FORMAT` | +| `test_the_shipped_file_is_a_bundle_this_build_can_import` | `version == bundle.FORMAT` | `version in bundle.READABLE` — the shipped starter is a v2 file and was **not** regenerated, because rewriting a shipped asset to keep a test's equality holding is changing the evidence to fit the test | + +### 7. The scale tool measured the wrong thing after compression — FIXED IN THE TOOL + +It looked for the plain `contextSnapshot` key and reported 0% once the export +started writing `contextSnapshotZ` — i.e. it would have reported the evidence as +*omitted* when it was merely encoded, which is precisely the mistake the +portability report exists to avoid making about anything. Both tools now decode. +Recorded because a reviewer reading an intermediate number in this report's +history would otherwise be misled. + +--- + +## V. Planning changes + +Distinguished by kind, as the brief asks. + +### Requirement corrections: none + +M9 altered no product requirement. `SPECIFICATION.md` and +`SECURITY-THREAT-MODEL.md` are **unchanged**: §16 already required the export to +preserve the exact active position, §6.1 already required an exact prompt/context +snapshot per turn, and M9 implements both rather than redefining either. No +acceptance condition was weakened. + +### Newly settled design decisions + +| Document | Decision | +| --- | --- | +| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **Story Cards are compatibility-only legacy data and no longer enter the narrator's prompt.** The brief asked for this to be decided; §K has the evidence that they were the untracked path §73 already forbade | +| `DATA-MODEL.md` §29, `TECHNICAL-DESIGN.md` §9.3 | **The bundle carries historical evidence**, and the two-category rule becomes three | +| `DATA-MODEL.md` §29, `TECHNICAL-DESIGN.md` §9.3 | **The format is versioned rather than extended**, and why §9.1's and §9.2's reasoning does not stretch to cover this addition | + +### Implementation facts recorded + +`TECHNICAL-DESIGN.md` §9.3 (the v3 contract, the encoding, the two-phase +transaction) and **new §9.4** (the backup); `DATA-MODEL.md` §29 (the format, the +three categories, required vs optional, the pointers translated) and **§31**, +which listed summaries as derived-and-rebuildable and now says what "where +practical" excludes, and where stored prompts sit — a reader landing on §31 +alone would otherwise conclude the export should be regenerating both; +`V1-ACCEPTANCE-TESTS.md` I01-I07 and L02-L04 results — with I05's M7-era limit +marked closed and **the original paragraph kept**, because the M9 decision is +only legible against it; `BUILD-MILESTONES.md` M9 status; `README.md` and +`VERSION.md`; `DEVELOPMENT.md` (a new section on the two recovery tools, taking a +backup, and the stop-move-start restore, plus the imported-campaign context note). + +M8's report moved to `archive/milestone-reports/`, by the rotation convention +`planning/README.md` states. + +--- + +## W. Residual risks and the M10-M11 handoff + +### Residual risks + +1. **A long campaign's bundle has a measured ceiling.** **~279 turns** against + the 20 MB import limit; a longer campaign can still be exported and would be + refused on import, which is the asymmetry worth naming. Far beyond M11's + 100-turn target, which is 14% of the cap. M9's own additions account for + **12%** of that ceiling — the other 88% is the per-position state document v2 + already carried, so lifting the ceiling means addressing *that*, not the + evidence. *Risk: low for v1; the fix is a streaming or chunked import, which + is architecture beyond M9.* +2. **`quick_check` rather than `integrity_check`** on a backup. It does the + structural work without the full index cross-check; a corrupt index that + `quick_check` misses would ride into the copy. *Risk: low, and the trade is + stated in the code — a backup verified too slowly to be taken is worse.* +3. **The backup is not offered on a schedule, and nothing prompts for one.** + *Risk: low, and outside M9's scope. `DEVELOPMENT.md` shows the `curl` for + `cron`.* +4. **`chunk_id` in a restored snapshot is stale.** It is a React key and a data + attribute, not a live pointer, and `chunk_index` plus `heading_path` still + identify the passage. *Risk: very low; documented in the code.* +5. **The importing machine's context window may differ from the source's.** + Documented rather than detected; detection is M11's. +6. **This machine cannot drive a file into or out of the browser.** Firefox here + is a snap, so its sandbox blocks reading a WebDriver-uploaded path, and its + headless download manager does not complete a `blob:` download. §Q labels + both substitutions and the end-to-end file path is proved without a browser + in `test_m9_clean_import.py`. *Risk: low for the product, real for future + evidence — M11 will meet the same wall on any file-based browser claim, and + the cheapest fix is an unconfined Firefox or Xvfb on the certifying machine.* + +### Carried to M10 + +- Media tables do not exist. §L records that as **not applicable** rather than + inventing schema, and M10 owns the coordinator, the providers and the jobs. +- The `scene` section of the state document is the extension point that exists + today, and it round-trips. + +### Carried to M11 + +- **The context window**, end to end: detection, the 100-turn certification, and + whether a Settings warning that reads the real window belongs in v1. +- **The streaming import**, if the bundle ceiling above is judged too low. +- Contrast and focus measurement, tablet tuning, and the broader WCAG audit + (inherited from M8's §U, untouched by M9). +- **An unconfined browser on the certifying machine**, so that a file-based + browser claim can be made end to end rather than in two labelled halves + (residual 6). + +### Recorded here but not M9's: four hands-on playtest findings + +A play session against **accepted, signed M8** — after M8 was accepted and +before M9 closeout — surfaced four product-quality observations. **None is an M9 +defect, none was caused by M9, and none blocks M9 acceptance.** They are written +up in full in **§Y**, and durably in `BUILD-MILESTONES.md` under M11 so they +survive this report's archival: + +| | Owner | +| --- | --- | +| A. The browser tab still reads `AI D&D` | M11 release polish | +| B. After Undo, the reader cannot tell where they are | M11 UX/release polish | +| C. The narration-length setting has no measurable effect | M11 realistic-model behaviour | +| D. Character identity / coreference confusion (root cause **unknown**) | M11 realistic-model / context diagnostic | + +D is the significant one, and its evidence is gone — the disposable playtest +database was destroyed, so no root cause is claimed. §Y records the reproduction +and classification plan, and one structural fact worth checking first: the +narrative state permits two entities to share a display name and reports +nothing. + +### Still open from earlier milestones, and not M9's + +Cross-layer duplication (`CONTEXT-AND-MEMORY.md` §22); the discarded-history +recovery screen (§63); whole-transcript copy and story search (§77, §78); the +read-only RPG world state; the inert legacy tables and the dual-dialect +migration code awaiting their cleanup migration. + +--- + +## X. Final milestone assessment + +### Is M9's Definition of Done satisfied? + +`BUILD-MILESTONES.md` states it as: *"A campaign can be safely exported, +imported into a clean data directory, and reopened at the exact intended active +position with authoritative history/state intact."* + +**Yes.** Measured three ways — in the suite, across two server processes with +two data directories and the exporting process stopped, and in a real browser +against a real narrator. + +### The brief's closing questions, answered directly + +| Question | Answer | +| --- | --- | +| **Can a campaign be moved into a clean data directory and reopened at the exact head?** | **Yes.** Into a database file that had never existed, from a second process, with the first stopped. The fixture's head is behind its own branch's tip *and* the abandoned line's, and is not the newest row written — so an importer guessing the tip, the newest row or the deepest row lands elsewhere. It opens where it was left, the retained future is there, and Redo reaches the same next turn in copy and source | +| **Is all authoritative state intact?** | **Yes.** The document at the exported head matches exactly; every historical position matches, walked in step; restoration stays a snapshot read rather than a replay; and the audit that explains the state travels with it — including the manual correction, which is the one change no narration explains | +| **Is historical prompt/context evidence intact?** | **Yes**, and this is the M8 handoff closed. An old turn in a restored campaign shows the prompt it was actually given and the passages it was shown, **after the source has been deleted in the copy** — asserted directly, not argued | +| **Are Save Points intact?** | **Yes.** Names, notes and coordinates; each resolves; each restores to its own position and its own state through M3's head movement; later history survives. A Save Point the file cannot satisfy is dropped, never retargeted | +| **Is knowledge intact without original filesystem paths?** | **Yes.** Content identical by SHA-256, classifications and lifecycle intact, searchable immediately — verified with the exporting machine's process dead and its directory holding a database the importer never opened | +| **Can derived data be rebuilt?** | **Yes**, and running that test found a defect that had made Reindex — the documented repair — unable to repair the state it most needed to. Both ends are fixed, and a database already carrying the damage repairs itself | +| **Can a consistent SQLite backup be created?** | **Yes**, through SQLite's online backup API, **while the application is being written to**, verified with `quick_check` before it is kept, and opened independently and read in every test. It works in the Docker image and lands on the mounted volume | +| **Are legacy bundles still supported?** | **Yes.** v1, v2, and every seam from pre-active-head onward — with nothing invented at any of them. A pre-M9 file gets no manufactured audit trail; a pre-M3 file opens at its tip *because that is the position it recorded* | +| **Are any required M9 conditions unverified?** | **No.** Every required condition has evidence at the boundary it names. What remains is stated as residual risk in §W with its measurement, not as an unverified claim | +| **Is it safe to proceed to M10 after review?** | **Yes**, on this evidence. M9 changed no schema, altered no product requirement, and added no media surface. It leaves M10 the `scene` section of the state document as the extension point that exists today, and §L records the absence of media tables as *not applicable* rather than inventing schema to satisfy the word "metadata" | + +### What a reviewer should look at hardest + +Named deliberately, because a report that only presents its strengths is harder +to review than one that points at its own seams: + +1. **The version bump (§D).** It is the one place M9 chose differently from two + earlier milestones that faced a similar decision. The argument is that an + absent key is unambiguous for a *position* and ambiguous for *evidence*; a + reviewer who disagrees should say so, because it is a contract decision and + not an implementation detail. +2. **The story-card decision (§K).** It removes something from the narrator's + prompt, which is a behaviour change in a portability milestone. The brief + authorised it and §73 required it; the judgement to check is whether stopping + at the prompt — and leaving the summariser's roster alone — is the right + line. +3. **The bundle ceiling (§R).** Real, measured, and stated rather than hidden. + Worth checking that ~279 turns is acceptable for v1 given a 100-turn + certification target, and that the export-succeeds/import-refuses asymmetry + is tolerable until a streaming import exists. +4. **Finding 3.** It would have shipped looking correct in memory and wrong on + disk. It was caught only because the test asserted the outcome rather than + the call, which is worth generalising. + +### Not done, and deliberately + +No image, video, TTS, STT, media provider or job architecture; no M10 media +coordinator; no 100-turn certification; no WCAG audit; no new context-window or +provider architecture; no discarded-history recovery browser. Retained history +survives correctly so a later screen can use it, which was M9's part of that. + +--- + +## Y. Post-M8 hands-on playtest findings + +**These are not M9 defects, were not caused by M9, and do not block M9 +acceptance.** They are recorded here because this is the report a reviewer is +holding, and because the M9 report will be archived when M10's replaces it — +`BUILD-MILESTONES.md` carries the durable copy under M11 for that reason. + +They come from a real play session against **accepted, signed M8** — a real +browser, a real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, and a +disposable isolated campaign database. **That database was deliberately +destroyed afterwards**, so the stored context snapshot for the turn in finding +D no longer exists. Everything below is therefore recorded as an *observed +symptom*, and no root cause is claimed that cannot now be proven. + +**Nothing in this section changed any application code.** Where the text states +how the product behaves today, it is from reading the code and running it, and +it is labelled as a mechanism rather than as a proven cause of what was seen. + +### A. The browser still calls the product "AI D&D" + +**Observed:** the browser tab/title reads `AI D&D`. + +**Verified:** `frontend/index.html` line 18 is `AI D&D`. It +is the inherited upstream title and M8 did not change it. + +**Not a false claim by any accepted document.** M8's report never asserted the +title was changed — its terminology audit covered `branch`, `fork`, `node`, +`head` and `depth`, and its "does it feel like the intended storyteller" +argument lists the navigation, the inspector, the glyph buttons and the removed +screens. So this is an **uncovered gap**, not documentation that needs +correcting. No documentation was corrected, and no code was changed. + +**The naming question is genuinely open, and should not be closed by a +find-and-replace.** The planning package and `README.md` call the product +*Adventure Storyteller*, but `SPECIFICATION.md` is explicit that the engine must +stay genre-agnostic — science fiction, mystery, horror, historical, westerns — +and *Adventure* is narrower than the product it names. A repository-wide rename +was deliberately **not** performed in this pass. For planning purposes a neutral +working name such as **Interactive Story** is used, and a browser-tab form such +as ` — Interactive Story` is a candidate, not a decision. + +**Owner: M11 release polish.** No earlier milestone touches the shell metadata. +It is a small, self-contained change whose only hard part is the naming +decision, which is the repository owner's. + +### B. Undo/Redo loses the reader's orientation + +**Observed:** Undo worked, and it was hard to tell which point in the story the +reader had moved to. + +**Not a correctness problem.** M3's active-head semantics behaved correctly and +M9 re-verified them at every boundary (§E, §S). This is presentation. + +**What the spec said, and why that was not enough.** `BROWSER-UX-SPEC.md` §8 — +*"The current endpoint should be clear."* — is five words, and the second +sentence beside it ("the input box always continues from the currently active +story head") was **already true** while the reader was lost. A requirement that +a working implementation satisfies while a real user cannot answer the question +is too vague to hold the behaviour. §8 has been strengthened in this pass; no UI +text is prescribed, because none is ratified. + +**The requirement, stated without wording:** after Undo, Redo, a Save Point +restore, an edit to an earlier turn, or any other movement of the active +position, a reader should be able to tell where they now are in the visible +story **without needing implementation terminology** — `branch`, `head`, `node` +and `depth` remain forbidden at the surface (§38 of the spec, and M8's audit). +A lightweight indicator such as `Moment 8` → `Moment 7`, optionally noting that +later story is still available, is a candidate. **Exact wording deliberately not +settled here.** + +**Owner: M11 UX/release polish**, with a browser regression scenario. + +### C. Narration length did not feel effective + +**Observed:** setup offered roughly *one paragraph* / *2-4 paragraphs* / *longer +exposition*; the reader chose **2-4 paragraphs** and felt replies were +substantially longer than that. + +**Deliberately not recorded as "the model ignored instructions."** The evidence +does not establish that, and there is a mechanism in the code worth checking +first. + +**The mechanism, verified by reading and running the code.** There are **two +independent length controls, and they do not reference each other**: + +1. The setup choice becomes **one English sentence** appended to the campaign's + `ai_instructions` — `"Keep responses to roughly two to four paragraphs."` + (`frontend/src/pages/NewCampaign.jsx`, `LENGTH_SENTENCE`). It changes **no + generation setting**. It sits in the **cached system block**, near the top of + the prompt. +2. `context/builder.py`'s `length_hint()` derives a numeric word range from + `Settings.max_output_tokens` — a **global application setting**, not the + campaign's choice — and emits it as a `[Hard limit: …]` line placed **after + the history**, near the end of the prompt. + +At the default `max_output_tokens = 800`, that second line reads, measured: + +``` +[Hard limit: this turn must not exceed 506 words, and it should not stop short + of about 177. Prefer the lower end of that range unless the scene genuinely + needs more. Finish the narration and append the state block well inside the + limit.] +``` + +**and it is byte-identical whether the reader chose brief, medium or long.** A +reader asking for two to four paragraphs is simultaneously told, in the more +recent and more numerically explicit of the two instructions, not to stop short +of about 177 words and that up to 506 are permitted. + +**This is a mechanism, not a proven cause.** Whether it produced what this +reader saw needs the reproduction below; the narrator's own instruction +following is also in play, and both could contribute. + +**What a reproduction must measure**, rather than judge by eye: + +1. exactly what length instruction(s) enter the **stored** prompt — both of the + above, with their positions; +2. whether the setup choice changes `max_output_tokens` or any generation + setting (**today: it does not**); +3. actual words, tokens and paragraph counts across repeated realistic turns per + setting, so a *directional* effect can be shown or disproved; +4. at least the reference 3B narrator **and** a stronger local narrator, since + instruction-following differs. + +**A design candidate, explicitly not ratified:** clearer approximate targets — +`Brief ~100-200 words`, `Standard ~200-400`, `Detailed ~400-700` — with the +setting actually moving the numeric budget. **Do not hard-truncate prose**: the +state block is emitted last and truncation removes it, which is the failure +`length_hint`'s own comments exist to avoid. + +**Owner: M11 realistic-model behaviour validation**, with a realistic-model test +rather than only a scripted-provider one. + +### D. Character identity / coreference confusion — the most important one + +**Observed:** a story established four people in an office — **Bill** +(protagonist), **Roger**, **John**, **Alice**. Later narration treated Alice as +though there were two different Alices, in language equivalent to *"Alice +wondered what Alice was doing."* + +**Root cause: UNKNOWN, and it cannot now be established.** The disposable +playtest database was deliberately destroyed, so the stored context snapshot for +that turn is gone. This report claims neither an inference-model defect nor an +application defect. The candidate classes are: + +| | | +| --- | --- | +| **Model failure** | the stored prompt correctly identifies one Alice and the 3B model still makes a coreference error | +| **State failure** | the narrative state holds duplicate or conflicting Alice records | +| **Context/derived failure** | state is correct, but assembled context, a summary, a retrieved memory or the history rendering presents two identities | +| **Combined weakness** | the context is not contradictory but is insufficiently explicit for a small model, producing an avoidable failure | + +**One structural fact a reproduction should check first**, verified by reading +the code and stated as a fact about the implementation rather than as a cause: + +> **The narrative state permits two distinct entities to share one display +> name, and nothing reports it.** Entities are keyed by the id the model +> supplies (`state["entities"][event["entity"]]`, `narrative/apply.py`), with +> `name` a separate display field. `validate.py`'s `DUPLICATE_ENTITY` rejects +> re-creating an entity **with the same key** — `key in known` — and there is no +> check anywhere in the narrative layer on the display name. So +> `create_entity(entity="alice", …, name="Alice")` followed by +> `create_entity(entity="alice_2", …, name="Alice")` both succeed, and the state +> then holds two entities that both render as "Alice". + +That is exactly one of the failure modes this finding describes. It does **not** +establish that it happened here — no evidence survives — and a reproduction may +well land on one of the other three classes instead. + +**Owner: M11 realistic-model / context diagnostic.** The test plan is below and +in `BUILD-MILESTONES.md` under M11. + +### The M11 character-identity diagnostic, as a test plan + +Deterministic setup with at least a protagonist and three same-scene +supporting characters — **Bill** (protagonist), **Alice**, **Roger**, **John** +— with unambiguous identities and roles established up front. Then a +multi-character interaction over enough turns to stress: pronouns; dialogue +attribution; people entering and leaving; reference by name; reference by role; +and one character speaking *about* another. + +**Detect and report, at minimum:** duplicate character creation; same-name +entity duplication; protagonist identity drift; dialogue attributed to the wrong +person; a character referring to themself as a separate same-named character; +and state/context disagreement about identity. + +**On any failure, preserve and report all of:** the authoritative state +immediately before generation; the exact stored context/prompt snapshot; the +recent-history section; summaries; retrieved memories; imported knowledge if +any; the narrator output; and the model identifier and settings. **M9 makes all +of that portable** (§H), so a failing campaign can now be exported and handed to +whoever investigates it — which is the practical reason this diagnostic is +newly worth writing. + +Then classify from the evidence: + +```text +STATE DEFECT +CONTEXT ASSEMBLY DEFECT +DERIVED MEMORY/SUMMARY DEFECT +MODEL FAILURE WITH CORRECT CONTEXT +AMBIGUOUS / MULTIPLE CONTRIBUTORS +``` + +Two rules for whoever runs it: **do not "fix" a model failure by changing +authoritative story state**, and **do not blame the model if the prompt already +contained the identity error.** That distinction should become part of M11's +realistic-model review methodology rather than a one-off judgement. + +### The standard fixture does not cover this failure class + +Reviewed, as asked. `TEST-CAMPAIGN-FIXTURE.md` stresses knowledge boundaries, +secrets, authority precedence, branch leakage and possession — its seven +deliberate traps are all of those — and it contains **no identity trap at all**; +the word *coreference* does not appear in it. Its on-stage cast is effectively +two people, Aldric and Mara, with Edrin established as missing rather than +present. + +So the user's suspicion is correct: **same-scene multi-character identity +continuity is not exercised by the standard fixture.** + +**The established fixture was deliberately not modified.** It is the +deterministic baseline several milestones' results are compared against, and +changing it would invalidate those comparisons. A **companion** fixture is +proposed instead, in a new appendix to that document, named +`Multi-Character Identity Test` and explicitly additive. + +--- + +*Written at the end of implementation, before review. The tree is staged; §32 of +the brief governs what happens next.*