Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1013c94eb1 | ||
|
|
44edece67e |
@@ -9,6 +9,11 @@ __pycache__/
|
||||
|
||||
# Database
|
||||
*.db
|
||||
# M9: verified database backups land beside the database. `*.db` already covers
|
||||
# the files; this names the directory so its purpose is obvious in a listing and
|
||||
# so nothing else that ends up there is committed by accident.
|
||||
backend/backups/
|
||||
data/backups/
|
||||
|
||||
# Node
|
||||
node_modules/
|
||||
|
||||
+118
@@ -143,6 +143,18 @@ outbound request, so a database edited by hand or a hostname that starts
|
||||
resolving somewhere new cannot turn a local install into an exfiltration path.
|
||||
There is no setting to relax it.
|
||||
|
||||
### A future media provider would be held to a stricter rule
|
||||
|
||||
The same file decides, plus one extra condition. A media endpoint — a local image
|
||||
or speech generator, when one is eventually supported — must be **loopback**, not
|
||||
merely on your LAN (`backend/app/media/providers.py`,
|
||||
`endpoint_rejection_reason`). A picture of a scene carries the scene with it, and
|
||||
a GPU that renders your campaign is a machine you are sitting at.
|
||||
|
||||
Nothing to configure today: no media provider ships, the registry is empty, and
|
||||
there is deliberately no media endpoint setting to fill in. The rule exists so
|
||||
that whoever adds the first provider finds it already there.
|
||||
|
||||
### Same host (the default)
|
||||
|
||||
```text
|
||||
@@ -223,6 +235,10 @@ Fourteen backend tests skip without something the machine may not have: seven
|
||||
need a second machine or an environment the suite cannot create, and the rest
|
||||
are the real-model tests below.
|
||||
|
||||
The suite takes about fifteen minutes. Several files spawn genuine server
|
||||
processes — a restart is only evidence if the process really went away — and
|
||||
those dominate the wall clock.
|
||||
|
||||
### The frontend component suite
|
||||
|
||||
M8 added one, because until M8 there was none — the browser was covered by real
|
||||
@@ -328,6 +344,100 @@ subsystem comes back as a route, if an API key becomes settable again, if the
|
||||
model timeout stops being configurable or becomes unbounded, or if a supported
|
||||
start path stops binding loopback.
|
||||
|
||||
## Backing up, and getting a campaign back
|
||||
|
||||
There are two recovery tools and they answer different questions. Using the
|
||||
wrong one is the most common way to be surprised later, so they are described
|
||||
together.
|
||||
|
||||
| | Campaign export | Database backup |
|
||||
| --- | --- | --- |
|
||||
| Covers | one campaign | every campaign, and your settings |
|
||||
| Shape | a JSON file you can read | a copy of the SQLite database |
|
||||
| Moves between machines | **yes** — this is the supported way | no; it is this machine's database |
|
||||
| Taken from | Export, on a campaign | Settings → *Back up everything on this machine* |
|
||||
| Restored by | Import campaign, on the library screen | replacing the database file, below |
|
||||
|
||||
### Exporting and importing a campaign
|
||||
|
||||
Export is on each campaign in the library, and in the campaign's own Settings
|
||||
panel. It writes one `.json` file holding the whole campaign: the story and its
|
||||
entire retained tree, the branch you are on and **the exact position you are
|
||||
reading at** — including one you undid back to — every alternate take, your Save
|
||||
Points, the authoritative state and its per-position snapshots, the state
|
||||
history that explains it, your imported knowledge with its classifications, the
|
||||
summaries and memories, and the prompt each turn was actually given.
|
||||
|
||||
Import is on the library screen and takes that file back, into this or any other
|
||||
installation. Nothing about the file refers to the machine that wrote it: the
|
||||
imported files come back from their content, not from a path, and no setting of
|
||||
yours is changed by importing somebody's campaign.
|
||||
|
||||
Two things it deliberately does **not** carry: your inference endpoint and model
|
||||
settings, which describe your machine rather than the campaign, and the
|
||||
rebuildable search indexes, which are rebuilt from the imported content before
|
||||
the import returns.
|
||||
|
||||
**A campaign imports whether or not the model that wrote it is installed here.**
|
||||
Recovering a campaign and being able to play it on are separate questions; the
|
||||
first never depends on the second.
|
||||
|
||||
### Backing up the whole database
|
||||
|
||||
Settings → Advanced → *Back up everything on this machine*. It writes a verified
|
||||
copy into a `backups/` directory beside the database itself, and tells you where.
|
||||
|
||||
It is a real backup rather than a file copy. It uses SQLite's online backup API,
|
||||
so it is safe to take **while you are playing** — a `cp` of a live database can
|
||||
read one page before a transaction and another after it, producing a file that
|
||||
opens, reports a schema, and is quietly missing rows. The copy is checked with
|
||||
`PRAGMA quick_check` before it is kept, an existing backup is never overwritten,
|
||||
and a failure leaves nothing behind.
|
||||
|
||||
You can also take one from the command line, or from `cron`:
|
||||
|
||||
```bash
|
||||
curl -s -X POST http://127.0.0.1:8000/api/backups | python3 -m json.tool
|
||||
```
|
||||
|
||||
### Restoring a whole database
|
||||
|
||||
There is deliberately no restore button, because restoring means replacing the
|
||||
file the running application has open — which is how you lose both copies at
|
||||
once. It is a three-step procedure and each step needs the application stopped:
|
||||
|
||||
```bash
|
||||
# 1. Stop the application. Nothing below is safe while it is running.
|
||||
# (Ctrl-C the server, or `docker compose down`.)
|
||||
|
||||
# 2. Keep what is there now, whatever state it is in. You may want it back.
|
||||
mv backend/data.db backend/data.db.before-restore
|
||||
|
||||
# 3. Put the backup in its place, and start the application again.
|
||||
cp backend/backups/adventure-storyteller-20260907-043000.db backend/data.db
|
||||
```
|
||||
|
||||
Check the file before you trust it, and check it again after starting:
|
||||
|
||||
```bash
|
||||
sqlite3 backend/backups/adventure-storyteller-20260907-043000.db 'PRAGMA quick_check;'
|
||||
# -> ok
|
||||
```
|
||||
|
||||
The database path is `backend/data.db` by default, and whatever `AIDND_DB_PATH`
|
||||
names otherwise — in Docker that is the mounted volume.
|
||||
|
||||
There is one file to move and no others: this build leaves SQLite in its default
|
||||
rollback-journal mode, so there are no `-wal` or `-shm` companions beside the
|
||||
database (`PRAGMA journal_mode` reports `delete`). A build that switched to WAL
|
||||
would have to move those too, and leaving them behind would pair a new database
|
||||
with an old write-ahead log.
|
||||
|
||||
**Prefer the campaign export for anything smaller than "everything".** Restoring
|
||||
a whole database rolls every campaign back to the moment the backup was taken,
|
||||
including the ones you did not mean to touch. To recover one campaign, export it
|
||||
and import it.
|
||||
|
||||
## The context window your Ollama actually enforces
|
||||
|
||||
**Check this before a long campaign.** The application budgets a prompt up to
|
||||
@@ -380,6 +490,14 @@ If you would rather not raise it at all, set **How much story to send** in
|
||||
Settings to the number `/api/ps` reports, and the prompt will be assembled to
|
||||
fit.
|
||||
|
||||
**This matters most on the machine you import to.** A campaign carries its
|
||||
history, not the window the machine that wrote it had, and a long imported
|
||||
campaign fills a prompt on its very first turn — so a deployment that has applied
|
||||
neither the derived model above nor a matching budget meets its ceiling
|
||||
immediately rather than gradually. Importing succeeds either way; it is the first
|
||||
turn afterwards that truncates. Check `/api/ps` on the destination before playing
|
||||
on an imported campaign, not after.
|
||||
|
||||
## What was made offline-safe, and how to check
|
||||
|
||||
Two runtime downloads were removed in Milestone M1. Both were invisible on a
|
||||
|
||||
@@ -65,11 +65,16 @@ that isn't the live one starts a new branch.
|
||||
change is recorded with what it was before and which turn caused it, so the Story State panel
|
||||
can show what changed and why. You can correct it by hand, and your correction outranks the
|
||||
story.
|
||||
- **AI Dungeon-compatible context engine.** Memory, author's note, and story cards (world
|
||||
info) are triggered by keywords in recent story text, then assembled under a token budget
|
||||
(`backend/app/context/builder.py`). Story cards are the inherited authored-lore primitive and
|
||||
are kept; they are **not** the knowledge library below, which is a first-class subsystem with
|
||||
its own classification, provenance, chunking and index.
|
||||
- **A context engine you can account for.** Memory, the author's note, the campaign's own
|
||||
rules, the authoritative state, the summary that applies here, and the retrieved imported
|
||||
passages are assembled under one token budget, in an order chosen so that a section which
|
||||
changes does not re-price the cached prefix above it (`backend/app/context/builder.py`).
|
||||
|
||||
**Story cards** — AI Dungeon's world-info primitive, inherited with the fork — are kept as
|
||||
legacy data and travel with an export, but they no longer reach the narrator. A keyword-matched
|
||||
card used to arrive in front of it as a world fact with no class, no visibility, no source and
|
||||
nothing to switch it off, competing with your imported Canon for the same budget; the knowledge
|
||||
library below replaces it, and does all of that explicitly.
|
||||
- **Total prompt transparency.** Every turn stores the exact prompt sent to the model.
|
||||
**Inspect context** on any narrator turn opens a readable account of what it was given —
|
||||
what it remembered, what it read, what it believes, and what each part cost — with the
|
||||
@@ -122,13 +127,33 @@ that isn't the live one starts a new branch.
|
||||
no story, and deleting a branch a Save Point is kept on is refused until you
|
||||
remove the Save Point yourself, so nothing takes a named moment away behind
|
||||
your back.
|
||||
- **Import and export.** AI Dungeon-compatible scenario format; JSON for everything else. An adventure exports as `ai-dnd-adventure-v2`, which carries the whole tree:
|
||||
every branch, every take, the fork points, which branches the story has left behind, the Save
|
||||
Points and the position it is being read at — all of them chosen rather than computed, which is
|
||||
the rule for what a bundle carries. A campaign opens where its head says, never at a Save Point
|
||||
merely because it has one. A campaign exported after two Undos imports still undone, with its
|
||||
retained future intact, instead of silently reopening at its newest turn. Files that predate
|
||||
the head position, and files saved in the old single-line format, still import.
|
||||
- **Import and export, as a recovery contract.** A campaign exports as one JSON file,
|
||||
`ai-dnd-adventure-v3`, and imports into a clean install on another machine. It carries the whole
|
||||
tree — every branch, every take, the fork points, which branches the story has left behind, the
|
||||
Save Points and the position it is being read at — and, since it is meant to be *recovery*
|
||||
rather than a copy of the text, everything that explains that story: the authoritative state
|
||||
and the typed events behind it, **the exact prompt each turn was given and the passages it was
|
||||
shown**, the summaries with the coordinates that decide whether they still apply, and your
|
||||
imported files with their classifications. A restored campaign can still answer "why does the
|
||||
state say this?" and "what was the narrator actually told?" — after the source file has been
|
||||
deleted and the canon edited since.
|
||||
|
||||
A campaign opens where its head says, never at a Save Point merely because it has one. Exported
|
||||
after two Undos, it imports still undone, with its retained future intact. Search indexes are
|
||||
not carried: they are rebuilt from the content, before the import returns. Nothing about your
|
||||
machine travels — no endpoint, no model, no path — so importing somebody's campaign never
|
||||
reconfigures your inference, and a campaign imports whether or not you have the model that
|
||||
wrote it. Older files still import: the flat single-line format, files that predate the head
|
||||
position, and files that predate everything above. AI Dungeon-compatible scenario format is
|
||||
still read and written for scenarios and story cards.
|
||||
- **A verified backup of everything, taken while you play.** Settings → *Back up everything on
|
||||
this machine* writes a copy of the whole database through SQLite's online backup API — not a
|
||||
file copy, which of a live database can read one page before a transaction and another after it
|
||||
and produce a file that opens and is quietly missing rows. It is checked with `PRAGMA
|
||||
quick_check` before it is kept, and an existing backup is never overwritten
|
||||
(`backend/app/backup.py`). Restoring one is a documented stop-move-start procedure in
|
||||
`DEVELOPMENT.md`, deliberately not a button: replacing the file the running application has
|
||||
open is how you lose both copies.
|
||||
- **Single user, no accounts.** There is no sign-up, no login, no session and no API key
|
||||
anywhere in the product. The storyteller API binds to loopback and is unauthenticated by
|
||||
design, because the only person who can reach it is the person running it. A new install
|
||||
@@ -234,8 +259,8 @@ leave it there.
|
||||
player input
|
||||
→ assemble context: [narrator prompt] + [world state + stat guide] + [AI instructions]
|
||||
+ [plot essentials] + [story summary] + [retrieved memories]
|
||||
+ [triggered story cards] + [retrieved imported knowledge,
|
||||
framed by class and bounded by its own budget]
|
||||
+ [retrieved imported knowledge, framed by class and
|
||||
bounded by its own budget]
|
||||
+ [history along this branch, token-budgeted]
|
||||
+ [author's note] + [player action]
|
||||
→ snapshot context (Insights)
|
||||
@@ -249,7 +274,7 @@ player input
|
||||
```
|
||||
frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||||
├─ routers/ scenarios, adventures, knowledge, story cards, chat, settings, debug
|
||||
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource
|
||||
├─ models.py SQLAlchemy: Scenario, Adventure, Branch, Action, StoryCard, Settings, Memory, KnowledgeSource, VisualProfile
|
||||
├─ migrations.py hand-rolled, versioned via PRAGMA user_version (92 and counting)
|
||||
├─ endpoints.py the inference-endpoint address policy
|
||||
├─ tlstrust.py one TLS context: the OS trust store unioned with certifi's
|
||||
@@ -262,7 +287,9 @@ frontend/ React + Vite SPA ──HTTP/SSE──► backend/ FastAPI
|
||||
├─ worldstate/ the inherited RPG stat engine — legacy, no longer authoritative
|
||||
├─ memorybank.py auto-summarization + embedding retrieval
|
||||
├─ knowledge/ the imported library: import, chunk, FTS5, embed, rank, inject
|
||||
├─ bundle.py the export/import formats, v2 (tree) and a v1 reader
|
||||
├─ bundle.py the export/import formats: v3, and readers for v2 and v1
|
||||
├─ media/ the future-media seam: scene packets, visual profiles, provider contracts
|
||||
├─ backup.py a verified whole-database copy, via SQLite's backup API
|
||||
├─ providers/ OpenAI-compatible adapter, streaming
|
||||
└─ data.db SQLite (path overridable via AIDND_DB_PATH)
|
||||
```
|
||||
@@ -272,7 +299,7 @@ development, Vite proxies `/api` to FastAPI.
|
||||
|
||||
## Tests
|
||||
|
||||
920 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||||
1,191 backend tests: unit tests plus full HTTP integration through the real turn engine, with
|
||||
the model provider mocked. They run with no route to the Internet, which is a requirement
|
||||
rather than a convenience — an offline claim proved on a machine that has been online once
|
||||
proves nothing. A further handful need a real local model and skip without one; they exist
|
||||
|
||||
@@ -0,0 +1,278 @@
|
||||
"""M9: a consistent copy of the whole database, taken while the app is running.
|
||||
|
||||
This is **not** the campaign bundle, and the two are not alternatives. They are
|
||||
different recovery tools and M9 keeps them apart deliberately:
|
||||
|
||||
campaign bundle one campaign, logical, portable between installations,
|
||||
importable into a clean data directory on another
|
||||
machine, readable by a human and by a later build
|
||||
database backup every campaign, every setting, physical, this machine,
|
||||
restored by putting the file back
|
||||
|
||||
The bundle is the primary cross-install recovery path and is what the acceptance
|
||||
tests measure. This exists for the other question: the reader has one database
|
||||
holding everything they have ever played, and wants a copy of it before they
|
||||
upgrade, move a disk, or try something they might regret.
|
||||
|
||||
## Why not `cp data.db backup.db`
|
||||
|
||||
Because a copy taken with the application running is a copy of a moving target.
|
||||
SQLite writes a database in pages, and a plain file copy can read page 5 before
|
||||
a transaction and page 900 after it — the result is a file that opens, reports a
|
||||
schema, and is silently missing or duplicating rows. In WAL mode it is worse: the
|
||||
committed data may be in a `-wal` file the copy never touched. Nothing warns
|
||||
anyone. The corruption is found later, by which time the original may be gone.
|
||||
|
||||
So this uses SQLite's own **online backup API** (`sqlite3.Connection.backup`),
|
||||
which is the supported mechanism for exactly this: it copies page by page while
|
||||
holding the right locks, restarts if a write moves the source underneath it, and
|
||||
produces a file that is a transactionally consistent snapshot of some committed
|
||||
point. The application keeps running throughout; no session is closed and no
|
||||
turn is blocked.
|
||||
|
||||
## What the procedure guarantees
|
||||
|
||||
1. The source database is opened **read-only** and is never written to. A backup
|
||||
that could damage what it is backing up would be worse than no backup.
|
||||
2. The copy is written to a temporary file beside the destination and renamed
|
||||
into place only after it has been verified, so an interrupted or failed run
|
||||
never leaves a half-written file wearing a backup's name. `os.replace` is
|
||||
atomic on the same filesystem, which is why the temporary sits in the
|
||||
destination's own directory rather than in `/tmp`.
|
||||
3. `PRAGMA quick_check` runs against the finished copy, opened as its own
|
||||
database, before it is renamed. A backup nobody verified is a belief.
|
||||
4. An existing file is never overwritten. Each run writes a new name stamped
|
||||
with the time, so yesterday's backup survives today's mistake — which is most
|
||||
of what a backup is for.
|
||||
5. Failure is reported and leaves nothing behind but the log line.
|
||||
|
||||
## What it does not do
|
||||
|
||||
There is no restore endpoint. Restoring a whole database means replacing the
|
||||
file the running application has open, and doing that from inside that
|
||||
application is a way to lose both copies. The procedure is in `DEVELOPMENT.md`:
|
||||
stop the app, move the file into place, start it. Campaign-level recovery — the
|
||||
common case, and the one that crosses machines — is the bundle.
|
||||
|
||||
No path comes from a caller. The destination directory is derived from the
|
||||
database the application is already using and the filename is generated here, so
|
||||
there is no request that can direct a write anywhere else (H08).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import os
|
||||
import sqlite3
|
||||
from dataclasses import dataclass
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
|
||||
from .database import DB_PATH
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
#: Where backups go: a directory beside the database itself. Beside, rather than
|
||||
#: inside a configurable location, because the one thing this must not do is
|
||||
#: write somewhere a request can name.
|
||||
DIRECTORY_NAME = "backups"
|
||||
|
||||
#: The stem every backup file carries, so a directory listing sorts by date and
|
||||
#: says what these files are without being opened.
|
||||
PREFIX = "adventure-storyteller"
|
||||
|
||||
|
||||
class BackupError(RuntimeError):
|
||||
"""A backup did not complete. The source database is untouched."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Backup:
|
||||
"""One finished, verified backup file."""
|
||||
|
||||
path: Path
|
||||
bytes: int
|
||||
pages: int
|
||||
seconds: float
|
||||
integrity: str
|
||||
|
||||
def as_dict(self) -> dict:
|
||||
return {
|
||||
# The name alone, not the path. The full path is a fact about this
|
||||
# machine's filesystem, and the reader is told the directory once by
|
||||
# the endpoint that lists them.
|
||||
"filename": self.path.name,
|
||||
"bytes": self.bytes,
|
||||
"pages": self.pages,
|
||||
"seconds": round(self.seconds, 3),
|
||||
"integrity": self.integrity,
|
||||
}
|
||||
|
||||
|
||||
def directory(db_path: Path | None = None) -> Path:
|
||||
"""The backup directory for a database, created if it does not exist."""
|
||||
root = (db_path or DB_PATH).parent / DIRECTORY_NAME
|
||||
root.mkdir(parents=True, exist_ok=True)
|
||||
return root
|
||||
|
||||
|
||||
def create(db_path: Path | None = None, *, now: datetime | None = None) -> Backup:
|
||||
"""Takes one verified backup of the live database, and returns it.
|
||||
|
||||
Raises `BackupError` on any failure, having removed whatever it had written.
|
||||
The source database is opened read-only and is never modified, so a failure
|
||||
here costs the backup and nothing else.
|
||||
"""
|
||||
source_path = db_path or DB_PATH
|
||||
if not source_path.exists():
|
||||
raise BackupError(f"There is no database at {source_path}.")
|
||||
stamp = (now or datetime.now()).strftime("%Y%m%d-%H%M%S")
|
||||
target = _unused_name(directory(source_path), stamp)
|
||||
# The temporary sits in the destination directory so the rename below is a
|
||||
# rename rather than a copy across filesystems, which would not be atomic.
|
||||
working = target.with_name(target.name + ".partial")
|
||||
started = datetime.now()
|
||||
try:
|
||||
pages = _copy(source_path, working)
|
||||
integrity = _verify(working)
|
||||
except BackupError:
|
||||
_discard(working)
|
||||
raise
|
||||
except Exception as exc: # noqa: BLE001 - reported, never raised raw
|
||||
_discard(working)
|
||||
log.exception("Backup of %s failed", source_path)
|
||||
raise BackupError(f"{type(exc).__name__}: {exc}") from exc
|
||||
size = working.stat().st_size
|
||||
# Only now does the file get the name a reader would trust.
|
||||
os.replace(working, target)
|
||||
return Backup(
|
||||
path=target,
|
||||
bytes=size,
|
||||
pages=pages,
|
||||
seconds=(datetime.now() - started).total_seconds(),
|
||||
integrity=integrity,
|
||||
)
|
||||
|
||||
|
||||
def _copy(source_path: Path, working: Path) -> int:
|
||||
"""Runs SQLite's online backup from `source_path` into a new file.
|
||||
|
||||
The source is opened through a URI with `mode=ro`, so this connection cannot
|
||||
write to it even by accident. The destination is a fresh database that this
|
||||
function creates; `backup()` overwrites whatever is in it, and the caller has
|
||||
guaranteed the name is unused.
|
||||
|
||||
Returns the number of pages copied, which is the one honest measure of how
|
||||
much was actually written — the file size counts pages the source had
|
||||
already allocated.
|
||||
"""
|
||||
source = sqlite3.connect(f"file:{source_path}?mode=ro", uri=True)
|
||||
try:
|
||||
destination = sqlite3.connect(working)
|
||||
try:
|
||||
copied = 0
|
||||
|
||||
def progress(_status, remaining, total):
|
||||
nonlocal copied
|
||||
copied = total - remaining
|
||||
|
||||
# `pages=-1` copies the whole database in one step while holding the
|
||||
# source's read lock, which is the right trade for a local
|
||||
# single-user database: it is the fastest option, it cannot restart
|
||||
# partway, and the lock it holds does not block readers.
|
||||
source.backup(destination, pages=-1, progress=progress)
|
||||
return copied
|
||||
finally:
|
||||
destination.close()
|
||||
finally:
|
||||
source.close()
|
||||
|
||||
|
||||
def _verify(working: Path) -> str:
|
||||
"""Runs `PRAGMA quick_check` against the finished copy.
|
||||
|
||||
Opened as its own connection, so what is checked is the file on disk rather
|
||||
than any page cache the copy left behind. `quick_check` rather than
|
||||
`integrity_check` because it does the structural work — every page reachable,
|
||||
every record readable — without the full index cross-check, which on a large
|
||||
database is minutes rather than moments. A backup nobody verified is a
|
||||
belief; a backup verified slowly enough that nobody takes one is worse.
|
||||
"""
|
||||
connection = sqlite3.connect(f"file:{working}?mode=ro", uri=True)
|
||||
try:
|
||||
rows = connection.execute("PRAGMA quick_check").fetchall()
|
||||
finally:
|
||||
connection.close()
|
||||
result = ", ".join(str(row[0]) for row in rows) if rows else "no result"
|
||||
if result != "ok":
|
||||
raise BackupError(
|
||||
f"The backup was written but did not verify: {result}. It has been "
|
||||
f"discarded; the original database is untouched."
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def _unused_name(root: Path, stamp: str) -> Path:
|
||||
"""A name in `root` that nothing is using.
|
||||
|
||||
An existing backup is never overwritten. Two backups taken inside one second
|
||||
are the only way to collide, and the counter settles that rather than one of
|
||||
them silently replacing the other.
|
||||
"""
|
||||
candidate = root / f"{PREFIX}-{stamp}.db"
|
||||
counter = 2
|
||||
while candidate.exists() or candidate.with_name(candidate.name + ".partial").exists():
|
||||
candidate = root / f"{PREFIX}-{stamp}-{counter}.db"
|
||||
counter += 1
|
||||
return candidate
|
||||
|
||||
|
||||
def _discard(working: Path) -> None:
|
||||
"""Removes a partial file, ignoring a file that is already gone."""
|
||||
try:
|
||||
working.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
def existing(db_path: Path | None = None) -> list[dict]:
|
||||
"""Every backup in the directory, newest first.
|
||||
|
||||
Names and sizes only. Reading one to report what is inside it would mean
|
||||
opening a database on every page load for a screen that is a list.
|
||||
|
||||
`taken_at` is read out of the **filename**, which is the stamp `create`
|
||||
wrote when it took the backup, and falls back to the file's modification
|
||||
time only for a name that does not parse. The two usually agree, and where
|
||||
they disagree the name is the one telling the truth: copying a backup to
|
||||
another disk, restoring it from an archive, or touching it all move the
|
||||
mtime, and a list that then reordered itself would report when the file was
|
||||
last handled rather than when the backup was taken.
|
||||
"""
|
||||
root = directory(db_path)
|
||||
rows = []
|
||||
for path in root.glob(f"{PREFIX}-*.db"):
|
||||
try:
|
||||
stat = path.stat()
|
||||
except OSError:
|
||||
continue
|
||||
rows.append({
|
||||
"filename": path.name,
|
||||
"bytes": stat.st_size,
|
||||
"taken_at": (
|
||||
_stamp_in(path.name) or datetime.fromtimestamp(stat.st_mtime)
|
||||
).isoformat(timespec="seconds"),
|
||||
})
|
||||
rows.sort(key=lambda row: (row["taken_at"], row["filename"]), reverse=True)
|
||||
return rows
|
||||
|
||||
|
||||
def _stamp_in(filename: str) -> datetime | None:
|
||||
"""The time in a backup's name, or `None` if it does not carry one."""
|
||||
rest = filename[len(PREFIX) + 1:].removesuffix(".db")
|
||||
# A collision within one second gets a `-2` suffix, which is not the stamp.
|
||||
stamp = "-".join(rest.split("-")[:2])
|
||||
try:
|
||||
return datetime.strptime(stamp, "%Y%m%d-%H%M%S")
|
||||
except ValueError:
|
||||
return None
|
||||
+1024
-47
File diff suppressed because it is too large
Load Diff
@@ -29,7 +29,9 @@ from ..knowledge import records as knowledge_records
|
||||
from . import encoding, history
|
||||
|
||||
AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
|
||||
CARD_BUDGET_SHARE = 0.4 # max share of non-reserved budget that story cards may take
|
||||
# `CARD_BUDGET_SHARE = 0.4` was here, and is gone with the injection it bounded
|
||||
# (M9). It is named rather than deleted silently because two other places
|
||||
# reasoned about their own share against it.
|
||||
NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene")
|
||||
SEPARATOR = "\n\n"
|
||||
|
||||
@@ -508,30 +510,47 @@ def build_context(
|
||||
adventure, available_after_knowledge, count_tokens, exclude_action_id
|
||||
)
|
||||
|
||||
# ----- Story cards: triggered by recent story text (the window history could fill) -----
|
||||
trigger_window = truncate_to_last_tokens(
|
||||
SEPARATOR.join(a.text for a in actions), available_after_knowledge
|
||||
)
|
||||
triggered = match_cards(adventure.story_cards, trigger_window)
|
||||
|
||||
card_budget = int(available_after_knowledge * CARD_BUDGET_SHARE)
|
||||
card_records = []
|
||||
lore_lines: list[str] = []
|
||||
# ----- Story cards: legacy, and no longer part of the narrator's prompt (M9)
|
||||
#
|
||||
# Until M9 a keyword-triggered story card was injected here as
|
||||
# `World Lore: <entry>`, taking up to 40% of what was left after the
|
||||
# imported knowledge had been placed.
|
||||
#
|
||||
# `IMPORTED-KNOWLEDGE-DESIGN.md` §73 settles that Story Cards are not the
|
||||
# production imported-knowledge store and, in as many words, that they "must
|
||||
# not become an alternate untracked path around the new knowledge
|
||||
# authority/provenance rules". That is exactly what this was. A card entry
|
||||
# arrived in front of the narrator as a world fact with:
|
||||
#
|
||||
# * no class — nothing said whether it was Canon, Reference or Inspiration,
|
||||
# so nothing framed how far the narrator could rely on it;
|
||||
# * no visibility — no narrator-only distinction at all;
|
||||
# * no source, no hash, no lifecycle, nothing to disable it with;
|
||||
# * no browser surface, since M8 removed the editor — so a reader could
|
||||
# neither see it nor switch it off;
|
||||
# * and no row in the context inspector, which renders `knowledge` and
|
||||
# never rendered `cards`.
|
||||
#
|
||||
# It also competed with imported Canon for one budget, which is the
|
||||
# arrangement M7 spent a milestone separating.
|
||||
#
|
||||
# M9's decision, recorded in the milestone report: story cards are
|
||||
# **compatibility-only legacy data**. Nothing is deleted. The rows stay, the
|
||||
# `/api/story-cards` endpoints stay, the bundle carries them out and back so
|
||||
# a round trip destroys nothing, and `memorybank.cast_brief` still reads them
|
||||
# as the summariser's character roster — a roster names who is on stage so a
|
||||
# memory says "Aldric" rather than "he", it never reaches the narrator, and
|
||||
# every memory written from it is authority-classified by the application
|
||||
# afterwards. What stops is the one path that asserted campaign facts to the
|
||||
# narrator without any of the controls §73 requires.
|
||||
#
|
||||
# `cards` stays in the report and is now always empty for a new turn.
|
||||
# Removing the key would break the historical snapshots that have one, which
|
||||
# M9 has just made portable: an old turn's evidence says story cards were
|
||||
# included, and it must go on saying so.
|
||||
card_records: list[dict] = []
|
||||
lore_section = None
|
||||
used = 0
|
||||
for match in triggered:
|
||||
line = f"World Lore: {match['entry'].strip()}"
|
||||
tokens = count_tokens(line)
|
||||
included = used + tokens <= card_budget
|
||||
if included:
|
||||
lore_lines.append(line)
|
||||
used += tokens
|
||||
card_records.append(
|
||||
{"id": match["id"], "name": match["name"], "keyword": match["keyword"],
|
||||
"included": included}
|
||||
)
|
||||
lore_section = (
|
||||
Section("world_lore", "\n".join(lore_lines)) if lore_lines else None
|
||||
)
|
||||
|
||||
# ----- Story history: newest first until the remaining budget is spent -----
|
||||
history_budget = available_after_knowledge - used
|
||||
|
||||
@@ -137,13 +137,53 @@ def index_line(heading_path: str, text_: str) -> str:
|
||||
|
||||
|
||||
def add(db: Session, chunk_id: int, heading_path: str, text_: str) -> None:
|
||||
"""Indexes one passage. The caller supplies the chunk's id as the rowid."""
|
||||
"""Indexes one passage. The caller supplies the chunk's id as the rowid.
|
||||
|
||||
`OR REPLACE`, and the reason is a defect M9 found rather than a defensive
|
||||
habit. The rowid is a chunk's primary key, so a row already sitting at it is
|
||||
by definition stale: the chunk that owned it does not exist, or is being
|
||||
rewritten by the reindex that called this. Either way the new passage is the
|
||||
truth and the old row is not.
|
||||
|
||||
Without it, an orphaned index row makes an ordinary import fail. SQLite
|
||||
reuses primary keys once the highest row is gone, so the next campaign to
|
||||
import a source is handed rowid 1 again, collides with an orphan, and gets a
|
||||
500 from `INSERT` — and `clear_index` cannot clear the orphan, because it
|
||||
finds index rows *through* the chunks, and there are none. That made Reindex,
|
||||
which is the documented repair, unable to repair this. `REPLACE` closes it
|
||||
from both ends: a leaked row is overwritten the moment the id comes round
|
||||
again, so an existing database repairs itself rather than needing a
|
||||
migration, and Reindex is the repair it is described as.
|
||||
|
||||
The leak itself is closed separately, in `importer.clear_campaign_index`.
|
||||
"""
|
||||
db.execute(
|
||||
sql(f"INSERT INTO {TABLE} (rowid, text) VALUES (:id, :text)"),
|
||||
sql(f"INSERT OR REPLACE INTO {TABLE} (rowid, text) VALUES (:id, :text)"),
|
||||
{"id": chunk_id, "text": index_line(heading_path, text_)},
|
||||
)
|
||||
|
||||
|
||||
def remove_adventure(db: Session, adventure_id: int) -> int:
|
||||
"""Drops every index row belonging to one campaign. Returns how many.
|
||||
|
||||
Scoped through the chunks, which is the only place the campaign is
|
||||
recorded — the index deliberately holds no copy of it
|
||||
(see "The table" above). So this has to run **before** the chunk rows go,
|
||||
which is what `importer.clear_campaign_index` is for.
|
||||
"""
|
||||
result = db.execute(
|
||||
sql(
|
||||
f"""
|
||||
DELETE FROM {TABLE} WHERE rowid IN (
|
||||
SELECT id FROM knowledge_chunks WHERE adventure_id = :adventure_id
|
||||
)
|
||||
"""
|
||||
),
|
||||
{"adventure_id": adventure_id},
|
||||
)
|
||||
return result.rowcount or 0
|
||||
|
||||
|
||||
def remove_chunks(db: Session, chunk_ids: list[int]) -> None:
|
||||
"""Drops passages from the index by id.
|
||||
|
||||
|
||||
@@ -362,6 +362,28 @@ def clear_index(db: Session, source: models.KnowledgeSource) -> None:
|
||||
db.expire(source, ["chunks"])
|
||||
|
||||
|
||||
def clear_campaign_index(db: Session, adventure: models.Adventure) -> int:
|
||||
"""Removes a whole campaign's lexical index rows. Returns how many.
|
||||
|
||||
Called before a campaign is deleted, and it has to be: the FTS index is a
|
||||
virtual table, so no foreign key reaches it and no `ON DELETE CASCADE`
|
||||
covers it. Deleting a campaign cascades `knowledge_sources` to
|
||||
`knowledge_chunks` and stops there, leaving one index row per passage
|
||||
belonging to a chunk that no longer exists.
|
||||
|
||||
Found in M9. The leak is not cosmetic. SQLite hands out the lowest free
|
||||
primary key, so once the highest chunk is gone the *next* source imported
|
||||
into *any* campaign is given a chunk id that an orphan already occupies, and
|
||||
the import fails with an integrity error — a 500 on an ordinary upload, in a
|
||||
campaign that has nothing to do with the deleted one. `fts.add` now repairs
|
||||
such a collision when it meets one; this stops it happening.
|
||||
|
||||
Vectors and passages need no equivalent, because both are real tables whose
|
||||
foreign keys cascade.
|
||||
"""
|
||||
return fts.remove_adventure(db, adventure.id)
|
||||
|
||||
|
||||
def delete_source(db: Session, source: models.KnowledgeSource) -> None:
|
||||
"""Removes a source and everything derived from it.
|
||||
|
||||
|
||||
@@ -50,10 +50,17 @@ from .records import Candidate, Result
|
||||
|
||||
#: Share of the non-protected budget that retrieved knowledge may spend.
|
||||
#:
|
||||
#: Story cards already take up to 40% (`CARD_BUDGET_SHARE`), and the history is
|
||||
#: what is left. A third is enough for several passages at the chunker's
|
||||
#: typical size and leaves the majority of the window to the story itself,
|
||||
#: which is the thing the reader came for.
|
||||
#: A third is enough for several passages at the chunker's typical size and
|
||||
#: leaves the majority of the window to the story itself, which is the thing the
|
||||
#: reader came for.
|
||||
#:
|
||||
#: This share was chosen when story cards could take up to 40% of the same
|
||||
#: budget and the history took what was left. M9 removed that injection
|
||||
#: (`IMPORTED-KNOWLEDGE-DESIGN.md` §73), so the history now gets that 40% back.
|
||||
#: The number here is deliberately unchanged: a third of the budget was chosen
|
||||
#: as the right amount of *imported material* to put in front of the narrator,
|
||||
#: not as a leftover, and raising it because room appeared would be changing
|
||||
#: retrieval behaviour under cover of a portability milestone.
|
||||
KNOWLEDGE_SHARE = 0.33
|
||||
|
||||
#: What each class may take of the knowledge budget. Canon may take all of it;
|
||||
|
||||
+5
-1
@@ -10,7 +10,9 @@ from starlette.exceptions import HTTPException as StarletteHTTPException
|
||||
from .database import engine
|
||||
from .limits import BodySizeLimitMiddleware
|
||||
from .migrations import bootstrap
|
||||
from .routers import adventures, chat, debug, scenarios, settings, story_cards
|
||||
from .routers import (
|
||||
adventures, backups, chat, debug, scenarios, settings, story_cards,
|
||||
)
|
||||
from .seed import seed_public_scenarios
|
||||
|
||||
bootstrap(engine)
|
||||
@@ -112,6 +114,8 @@ app.include_router(scenarios.router)
|
||||
app.include_router(adventures.router)
|
||||
app.include_router(story_cards.router)
|
||||
app.include_router(settings.router)
|
||||
# M9: a verified copy of the whole database, taken while the app is running.
|
||||
app.include_router(backups.router)
|
||||
app.include_router(chat.router)
|
||||
app.include_router(debug.router)
|
||||
|
||||
|
||||
@@ -0,0 +1,66 @@
|
||||
"""M10: the seam a future media provider plugs into, and nothing behind it.
|
||||
|
||||
This package is **readiness, not media**. Nothing here generates an image, a
|
||||
video, audio, speech or a transcription; nothing here opens a socket; nothing
|
||||
here is required for the storyteller to run. A campaign plays exactly as it did
|
||||
in M9 with none of this configured, which is M10's central acceptance
|
||||
condition — see `test_m10_no_media.py`.
|
||||
|
||||
## What M10 found already built, and therefore did not build again
|
||||
|
||||
The largest finding of the milestone is how little of it needed inventing.
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §5 asks the story system to persist a structured
|
||||
scene snapshot with a campaign, a lineage, a source position, a location and the
|
||||
characters present. **All of that already exists**, and has since M5:
|
||||
|
||||
state["scene"] = {"summary": …, "location": <entity key>,
|
||||
"present": [<entity keys>],
|
||||
"at": {"branch_id": …, "depth": …}}
|
||||
|
||||
written only by the validated `set_scene` typed event (ADR 010), snapshotted per
|
||||
node in `actions.narrative_state_after` (M5), restored on every head movement by
|
||||
`attempts.restore_state` (M3/M4), and carried per position in the M9 v3 bundle.
|
||||
So it is already authoritative, already lineage-safe, already survives Undo,
|
||||
Redo, Save Point restore, divergence and restart, and already round-trips into a
|
||||
clean data directory.
|
||||
|
||||
Building a `scenes` table beside that would have been a second representation of
|
||||
information the application already stores authoritatively — the one thing the
|
||||
M10 brief forbids — and it would have needed its own lineage rules, its own
|
||||
restore path and its own bundle carriage, each a chance to disagree with the
|
||||
state document. **So M10 stores no scene rows.** It reads the scene that is
|
||||
already there.
|
||||
|
||||
## What was actually missing
|
||||
|
||||
Three things, and this package is each of them:
|
||||
|
||||
* `profiles.py` — **visual profiles.** Stable descriptors for how an entity
|
||||
*looks*, which nothing recorded. Campaign-scoped rather than per-position,
|
||||
because a character does not change appearance when the story forks (K02, K03).
|
||||
* `packet.py` — **the Scene Packet.** A bounded, provider-neutral,
|
||||
hidden-information-safe view of one scene, built on demand from authoritative
|
||||
state. Persisted nowhere, because it is a pure function of things that are.
|
||||
* `providers.py` — **the provider contracts.** Types and protocols for image,
|
||||
video, audio, TTS and STT, with no provider vocabulary anywhere in them, plus
|
||||
the loopback-only endpoint rule the media contract asks for.
|
||||
|
||||
## The authority direction, which never reverses
|
||||
|
||||
accepted story -> narrative state -> scene packet -> future provider
|
||||
|
||||
Every arrow points away from authority. A visual profile is not a story fact; a
|
||||
scene packet is a read; a future asset would be a depiction. Nothing in this
|
||||
package writes `narrative_state`, emits a state event, or moves the head — and
|
||||
`test_m10_authority.py` asserts that by running each operation and comparing the
|
||||
authoritative document byte for byte either side.
|
||||
|
||||
That is the rule `MEDIA-EXTENSION-CONTRACT.md` §35 and §49 state, and the reason
|
||||
it is enforced structurally rather than by convention: the only code that may
|
||||
change authoritative state is the M5 event pipeline, and nothing here imports
|
||||
it.
|
||||
"""
|
||||
|
||||
from . import packet, profiles, providers
|
||||
|
||||
__all__ = ["packet", "profiles", "providers"]
|
||||
@@ -0,0 +1,328 @@
|
||||
"""M10: the Scene Packet — one accepted scene, bounded, for a future provider.
|
||||
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §10-12 asks for a normalised, provider-independent
|
||||
description of a scene, and asks explicitly that a provider **not** normally
|
||||
receive the campaign transcript. This module builds that description.
|
||||
|
||||
## It is constructed, never stored
|
||||
|
||||
A packet is a pure function of things that are already persisted: the
|
||||
authoritative state document at a position, the entity records inside it, and
|
||||
the campaign's visual profiles. Storing one would create a second copy of all of
|
||||
that, which could then disagree with the first — and the packet has no field the
|
||||
source of truth does not already hold.
|
||||
|
||||
So there is no `scene_packets` table, nothing to migrate, nothing to keep in
|
||||
step with the head, and nothing to carry in a bundle. Rebuilding it costs one
|
||||
state read and one profile query. That is the same reasoning M9 applied to the
|
||||
FTS index and the knowledge passages, applied to a smaller thing.
|
||||
|
||||
## Scene identity, without a scenes table
|
||||
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §10 shows a `scene_id`, and the M10 brief asks
|
||||
that a future asset be able to name unambiguously:
|
||||
|
||||
campaign -> lineage/story position -> source turn or turn range -> scene
|
||||
|
||||
That is a **coordinate**, and the application already has one. So the identity
|
||||
is derived rather than allocated:
|
||||
|
||||
c<adventure>:b<branch>:<start>-<end>
|
||||
|
||||
Two properties follow, and both matter more than a surrogate key would have:
|
||||
|
||||
* it is **stable** — the same scene yields the same id on any machine, before
|
||||
and after an export, without a row having to travel;
|
||||
* it is **resolvable** — a future asset holding this string can be turned back
|
||||
into the exact accepted position it depicts, with no lookup table.
|
||||
|
||||
A surrogate `scene_id` would have needed a table, a lineage column, a restore
|
||||
path and bundle carriage, all to name something the coordinate already names.
|
||||
|
||||
## Ranges, because a video is not a turn
|
||||
|
||||
`build` takes a range, not a position. §30-31 of the contract describe a video
|
||||
covering several accepted turns, and the M10 brief is explicit that neither
|
||||
"one turn == one scene" nor "one scene == one asset" may be assumed.
|
||||
|
||||
So `start` and `end` are depths on one branch, the identity carries both, and a
|
||||
single-turn image is the case where they are equal rather than a different kind
|
||||
of request. Several future assets may name the same identity; nothing here
|
||||
allocates or records them, so nothing constrains how many there are.
|
||||
|
||||
## What is deliberately not in a packet
|
||||
|
||||
**The transcript.** Not a summarised version of it either. The packet carries
|
||||
the scene's own summary — the one sentence the story itself accepted through
|
||||
`set_scene` — and the entities present. A provider that needs to depict a room
|
||||
does not need to have read the campaign.
|
||||
|
||||
**Imported knowledge, of any class.** Not canon, not reference, not
|
||||
inspiration, and emphatically not a narrator-only source. This is the hidden
|
||||
information boundary and it is drawn structurally: this module never reads
|
||||
`knowledge_sources`, so there is no filter to get wrong and no marker to
|
||||
overlook. A secret reaches a packet only if the *story* put it into accepted
|
||||
state through a validated event — which is the correct rule, because at that
|
||||
point it is something that happened rather than something the narrator knows.
|
||||
|
||||
**Memories and summaries.** Derived narrative text about the campaign's past,
|
||||
which is not what depicting a present moment needs.
|
||||
|
||||
**Facts, relationships and threads.** These are the campaign's reasoning about
|
||||
itself. A `continuity_constraints` list carries the few that bear on depiction —
|
||||
what a character is holding, where they are — and nothing else.
|
||||
|
||||
The result is that the honest answer to "what could leak through a packet" is
|
||||
"what the accepted scene contains", which is what a picture of that scene would
|
||||
show anyway.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from .. import models
|
||||
from ..context import lineage
|
||||
from ..narrative import model as narrative_model
|
||||
from ..narrative import store as narrative_store
|
||||
from . import profiles as visual_profiles
|
||||
|
||||
#: How many entities one packet will describe. A scene is a moment with people
|
||||
#: in it; a request naming two hundred is a runaway state document rather than a
|
||||
#: picture, and the bound keeps a future provider's prompt finite.
|
||||
MAX_CHARACTERS = 24
|
||||
MAX_OBJECTS = 24
|
||||
MAX_CONSTRAINTS = 24
|
||||
|
||||
|
||||
def scene_id(adventure_id: int, branch_id: int | None, start: int, end: int) -> str:
|
||||
"""The derived, stable identity for one scene. See the module docstring."""
|
||||
branch = branch_id if branch_id is not None else 0
|
||||
return f"c{adventure_id}:b{branch}:{start}-{end}"
|
||||
|
||||
|
||||
def parse_scene_id(value: str) -> dict | None:
|
||||
"""Turns a scene identity back into the coordinate it names, or `None`.
|
||||
|
||||
The half that makes the derived identity worth having: a future asset
|
||||
holding this string can be resolved to an accepted position without a table.
|
||||
"""
|
||||
try:
|
||||
campaign, branch, span = str(value).split(":")
|
||||
start, end = span.split("-")
|
||||
return {
|
||||
"adventure_id": int(campaign.lstrip("c")),
|
||||
"branch_id": int(branch.lstrip("b")),
|
||||
"start": int(start),
|
||||
"end": int(end),
|
||||
}
|
||||
except (ValueError, AttributeError):
|
||||
return None
|
||||
|
||||
|
||||
def build(
|
||||
db: Session,
|
||||
adventure: models.Adventure,
|
||||
*,
|
||||
start: int | None = None,
|
||||
end: int | None = None,
|
||||
) -> dict:
|
||||
"""The Scene Packet for a range of accepted story on the active branch.
|
||||
|
||||
Defaults to the scene at the active head, which is the ordinary case: an
|
||||
image of what is happening now. `start` and `end` are depths on the active
|
||||
branch; passing both describes a stretch, which is what a future video
|
||||
would ask for.
|
||||
|
||||
Reads. Writes nothing, and cannot: this module imports no writer, emits no
|
||||
event and does not touch the head. `test_m10_authority.py` asserts the
|
||||
authoritative document is byte-identical either side of a build.
|
||||
"""
|
||||
state = narrative_store.current(adventure)
|
||||
scene = state.get("scene") if isinstance(state.get("scene"), dict) else {}
|
||||
|
||||
branch_id = adventure.head_branch_id
|
||||
head_depth = adventure.head_depth
|
||||
# The scene's own coordinate is the position `set_scene` last ran at, which
|
||||
# is where the depiction belongs. It can sit behind the head — the story may
|
||||
# have moved on without re-establishing the scene — and that is correct: the
|
||||
# picture is of the moment the scene was set, not of a later turn that did
|
||||
# not change it.
|
||||
at = scene.get("at") if isinstance(scene.get("at"), dict) else {}
|
||||
scene_branch = at.get("branch_id") if at.get("branch_id") is not None else branch_id
|
||||
scene_depth = at.get("depth") if _is_int(at.get("depth")) else head_depth
|
||||
|
||||
first = start if _is_int(start) else scene_depth
|
||||
last = end if _is_int(end) else max(first, scene_depth)
|
||||
if last < first:
|
||||
first, last = last, first
|
||||
|
||||
profiles = visual_profiles.by_key(db, adventure)
|
||||
location_key = scene.get("location") if isinstance(scene.get("location"), str) else None
|
||||
present = [k for k in (scene.get("present") or []) if isinstance(k, str)]
|
||||
|
||||
return {
|
||||
"scene_id": scene_id(adventure.id, scene_branch, first, last),
|
||||
"campaign": {"id": adventure.id, "title": adventure.title},
|
||||
# Where in the story this is, in the vocabulary the application already
|
||||
# uses internally. A future provider does not read these; a future
|
||||
# coordinator resolving an asset back to its source does.
|
||||
"turn_range": {"branch_id": scene_branch, "start": first, "end": last},
|
||||
"lineage": _lineage_of(db, adventure),
|
||||
"location": _entity_view(state, profiles, location_key),
|
||||
"characters": [
|
||||
view for key in present[:MAX_CHARACTERS]
|
||||
if (view := _entity_view(state, profiles, key)) is not None
|
||||
],
|
||||
"objects": _objects(state, profiles, present, location_key),
|
||||
"action_summary": str(scene.get("summary") or ""),
|
||||
"continuity_constraints": _constraints(state, present, location_key),
|
||||
# Present, empty, and deliberately so — see `_ambience`.
|
||||
"ambience": _ambience(scene),
|
||||
"source": {
|
||||
# What produced this, so a future asset's provenance can say which
|
||||
# build's rules bounded the packet it was made from.
|
||||
"packet_version": PACKET_VERSION,
|
||||
"head_depth": head_depth,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
#: The packet's own shape version. A future provider adapter can branch on it if
|
||||
#: the packet gains fields; nothing in the story engine reads it.
|
||||
PACKET_VERSION = 1
|
||||
|
||||
|
||||
def _lineage_of(db: Session, adventure: models.Adventure) -> list[dict]:
|
||||
"""The capped lineage this scene sits on, as provenance.
|
||||
|
||||
Read through `lineage.path_of`, the same helper every story read uses, so a
|
||||
packet cannot describe a position the story could not. M10 builds no media
|
||||
head: there is one head, and this follows it.
|
||||
"""
|
||||
try:
|
||||
path = lineage.path_of(db, adventure)
|
||||
except Exception: # noqa: BLE001 - a packet is a read; it does not raise
|
||||
return []
|
||||
entries = getattr(path, "entries", None)
|
||||
if not entries:
|
||||
return []
|
||||
return [
|
||||
{"branch_id": branch_id, "through_depth": cap}
|
||||
for branch_id, cap in entries
|
||||
]
|
||||
|
||||
|
||||
def _entity_view(state: dict, profiles: dict, key: str | None) -> dict | None:
|
||||
"""One entity as a packet describes it: what it is, plus how it looks."""
|
||||
if not key:
|
||||
return None
|
||||
found = narrative_model.entity(state, key)
|
||||
if found is None:
|
||||
return None
|
||||
return {
|
||||
"key": key,
|
||||
"name": narrative_model.entity_name(state, key),
|
||||
"type": found.get("type") or "other",
|
||||
"status": found.get("status") or "active",
|
||||
"description": found.get("description") or "",
|
||||
# `None` rather than an empty profile, so a provider can tell "nobody
|
||||
# said how this looks" from "somebody said it looks like nothing".
|
||||
"visual_profile": profiles.get(key),
|
||||
}
|
||||
|
||||
|
||||
def _objects(
|
||||
state: dict, profiles: dict, present: list[str], location_key: str | None
|
||||
) -> list[dict]:
|
||||
"""The things visibly in the scene, from what the present entities hold.
|
||||
|
||||
Possession is the only relation in the state document that says an object is
|
||||
*somewhere*, so it is the honest source for "what would be in the picture".
|
||||
An item nobody in the scene is carrying is not depicted, which is the same
|
||||
rule a reader would apply looking at the room.
|
||||
"""
|
||||
possessions = state.get("possessions")
|
||||
if not isinstance(possessions, dict):
|
||||
return []
|
||||
holders = set(present) | ({location_key} if location_key else set())
|
||||
out: list[dict] = []
|
||||
for item_key, holder in possessions.items():
|
||||
if holder not in holders or not isinstance(item_key, str):
|
||||
continue
|
||||
view = _entity_view(state, profiles, item_key)
|
||||
if view is None:
|
||||
continue
|
||||
view["held_by"] = holder
|
||||
out.append(view)
|
||||
if len(out) >= MAX_OBJECTS:
|
||||
break
|
||||
return out
|
||||
|
||||
|
||||
def _constraints(
|
||||
state: dict, present: list[str], location_key: str | None
|
||||
) -> list[str]:
|
||||
"""The few facts that bear on depicting *this* scene, as sentences.
|
||||
|
||||
Deliberately narrow. The state document's `facts` list is the campaign's
|
||||
reasoning about itself and most of it has nothing to do with a picture;
|
||||
forwarding all of it would make the packet a state dump with a different
|
||||
name, and would be the route by which something the scene has not exposed
|
||||
reached a provider.
|
||||
|
||||
So only two kinds are carried: where the present entities are, and what they
|
||||
are holding. Both are already visible in the scene by construction.
|
||||
"""
|
||||
out: list[str] = []
|
||||
for key in present:
|
||||
found = narrative_model.entity(state, key)
|
||||
if found is None:
|
||||
continue
|
||||
name = narrative_model.entity_name(state, key)
|
||||
status = found.get("status")
|
||||
if status and status != "active":
|
||||
out.append(f"{name} is {status}.")
|
||||
if len(out) >= MAX_CONSTRAINTS:
|
||||
return out
|
||||
possessions = state.get("possessions")
|
||||
if isinstance(possessions, dict):
|
||||
for item_key, holder in possessions.items():
|
||||
if holder not in present:
|
||||
continue
|
||||
out.append(
|
||||
f"{narrative_model.entity_name(state, holder)} is carrying "
|
||||
f"{narrative_model.entity_name(state, item_key)}."
|
||||
)
|
||||
if len(out) >= MAX_CONSTRAINTS:
|
||||
break
|
||||
return out
|
||||
|
||||
|
||||
def _ambience(scene: dict) -> dict:
|
||||
"""Time of day, lighting and mood — present in the shape, empty in v1.
|
||||
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §5 lists these among a scene snapshot's
|
||||
conceptual fields, and M10 **does not** add them to the `set_scene` event
|
||||
that would establish them.
|
||||
|
||||
That is a deliberate deferral rather than an oversight. Adding them would
|
||||
mean extending M5's typed-event vocabulary, which means teaching the
|
||||
narrator to emit them, which means changing the prompt — and M10's central
|
||||
acceptance condition is that ordinary story flow is *unchanged*. Buying
|
||||
three optional fields at the price of touching every narration was the wrong
|
||||
trade for a milestone whose deliverable is a seam.
|
||||
|
||||
So the keys are here and are `None`, read from the scene document if a later
|
||||
milestone starts recording them. A provider adapter written today against
|
||||
this shape keeps working when they arrive.
|
||||
"""
|
||||
return {
|
||||
"time_of_day": scene.get("time_of_day") or None,
|
||||
"lighting": scene.get("lighting") or None,
|
||||
"mood": scene.get("mood") or None,
|
||||
}
|
||||
|
||||
|
||||
def _is_int(value) -> bool:
|
||||
return isinstance(value, int) and not isinstance(value, bool)
|
||||
@@ -0,0 +1,220 @@
|
||||
"""M10: reading and writing how an entity looks.
|
||||
|
||||
`models.VisualProfile` carries the design reasoning — why these rows are
|
||||
campaign-scoped rather than per-position, why there is one table for characters,
|
||||
locations and items, and why nothing here is story state. This module is the
|
||||
narrow set of operations on them, and its own job is to make two things true:
|
||||
|
||||
* **a profile can only name an entity the campaign actually has**, so a typo
|
||||
produces an error rather than a row describing nobody;
|
||||
* **writing one changes nothing authoritative**, which is guaranteed by this
|
||||
module not importing anything that could.
|
||||
|
||||
## Why the entity is checked against the current head
|
||||
|
||||
An entity key means something only in a state document, and a campaign has a
|
||||
different document at every position. The check is made against the state at
|
||||
the **active head** — the story the reader is on — for the same reason
|
||||
`narrative/validate.py` resolves its `refs` there: it is the only position the
|
||||
reader is looking at, and a key that means nothing there is a mistake, not a
|
||||
branch subtlety.
|
||||
|
||||
The row that results is campaign-scoped anyway, so a profile written while
|
||||
standing on one branch is visible from every branch. That asymmetry is
|
||||
deliberate and is the continuity the profile exists for: the check is *"does
|
||||
this name someone"*, and the storage answers *"what do they look like"*, which
|
||||
does not vary by path.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from .. import models
|
||||
from ..narrative import model as narrative_model
|
||||
from ..narrative import store as narrative_store
|
||||
|
||||
#: How many descriptors one profile may carry, and how long each may be. A
|
||||
#: profile is a handful of stable traits, not a document: the bound exists so a
|
||||
#: future provider's prompt cannot be grown without limit through this door, and
|
||||
#: so one campaign cannot store an essay per entity.
|
||||
MAX_DESCRIPTORS = 40
|
||||
MAX_FEATURES = 40
|
||||
MAX_VALUE = 400
|
||||
MAX_STYLE_NOTES = 2_000
|
||||
MAX_KEY = 200
|
||||
|
||||
|
||||
class ProfileError(ValueError):
|
||||
"""A visual profile could not be written, and why."""
|
||||
|
||||
|
||||
def entity_exists(state: dict, entity_key: str) -> bool:
|
||||
"""Whether the state document names this entity."""
|
||||
return narrative_model.entity(state, entity_key) is not None
|
||||
|
||||
|
||||
def set_profile(
|
||||
db: Session,
|
||||
adventure: models.Adventure,
|
||||
entity_key: str,
|
||||
*,
|
||||
descriptors: dict | None = None,
|
||||
features: list | None = None,
|
||||
style_notes: str | None = None,
|
||||
) -> models.VisualProfile:
|
||||
"""Records how `entity_key` looks, creating or replacing the profile.
|
||||
|
||||
Replaces rather than merges. A profile is one answer to "what does this look
|
||||
like", and merging would make it impossible to *remove* a descriptor — the
|
||||
caller would be able to add "wearing a red coat" and never take it off,
|
||||
which for continuity metadata is the wrong default. A caller that wants to
|
||||
amend one reads it first.
|
||||
|
||||
Raises `ProfileError` if the campaign's state at the active head does not
|
||||
name the entity, or if the profile is malformed. It writes nothing in either
|
||||
case, and it writes nothing to `narrative_state` in any case.
|
||||
"""
|
||||
key = _checked_key(entity_key)
|
||||
state = narrative_store.current(adventure)
|
||||
if not entity_exists(state, key):
|
||||
raise ProfileError(
|
||||
f"This campaign has no entity called {key!r}, so there is nothing "
|
||||
f"for a visual profile to describe. Profiles attach to the "
|
||||
f"campaign's own entities, not to names."
|
||||
)
|
||||
row = get_profile(db, adventure, key)
|
||||
if row is None:
|
||||
row = models.VisualProfile(adventure_id=adventure.id, entity_key=key)
|
||||
db.add(row)
|
||||
row.descriptors = _checked_descriptors(descriptors)
|
||||
row.features = _checked_features(features)
|
||||
row.style_notes = _checked_notes(style_notes)
|
||||
return row
|
||||
|
||||
|
||||
def get_profile(
|
||||
db: Session, adventure: models.Adventure, entity_key: str
|
||||
) -> models.VisualProfile | None:
|
||||
return db.execute(
|
||||
select(models.VisualProfile).where(
|
||||
models.VisualProfile.adventure_id == adventure.id,
|
||||
models.VisualProfile.entity_key == entity_key,
|
||||
)
|
||||
).scalars().first()
|
||||
|
||||
|
||||
def all_for(db: Session, adventure: models.Adventure) -> list[models.VisualProfile]:
|
||||
return list(db.execute(
|
||||
select(models.VisualProfile)
|
||||
.where(models.VisualProfile.adventure_id == adventure.id)
|
||||
.order_by(models.VisualProfile.entity_key)
|
||||
).scalars().all())
|
||||
|
||||
|
||||
def by_key(db: Session, adventure: models.Adventure) -> dict[str, dict]:
|
||||
"""Every profile in the campaign, keyed by entity, as plain dictionaries.
|
||||
|
||||
One query, because the Scene Packet needs several profiles at once and
|
||||
fetching them per entity would be a query per character in the scene.
|
||||
"""
|
||||
return {row.entity_key: as_dict(row) for row in all_for(db, adventure)}
|
||||
|
||||
|
||||
def as_dict(row: models.VisualProfile) -> dict:
|
||||
"""One profile as it appears in a Scene Packet."""
|
||||
return {
|
||||
"descriptors": dict(row.descriptors or {}),
|
||||
"features": list(row.features or []),
|
||||
"style_notes": row.style_notes or "",
|
||||
}
|
||||
|
||||
|
||||
def delete_profile(
|
||||
db: Session, adventure: models.Adventure, entity_key: str
|
||||
) -> bool:
|
||||
"""Removes a profile. Returns whether there was one.
|
||||
|
||||
Deleting a profile removes a *description*, never the entity: the entity
|
||||
lives in the authoritative state document and nothing here can reach it.
|
||||
"""
|
||||
row = get_profile(db, adventure, entity_key)
|
||||
if row is None:
|
||||
return False
|
||||
db.delete(row)
|
||||
return True
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the checking
|
||||
|
||||
def _checked_key(entity_key) -> str:
|
||||
if not isinstance(entity_key, str) or not entity_key.strip():
|
||||
raise ProfileError("A visual profile has to name an entity.")
|
||||
key = entity_key.strip()
|
||||
if len(key) > MAX_KEY:
|
||||
raise ProfileError(f"Entity keys are at most {MAX_KEY} characters.")
|
||||
return key
|
||||
|
||||
|
||||
def _checked_descriptors(descriptors) -> dict:
|
||||
"""Trait -> value, both short strings.
|
||||
|
||||
Values are text rather than arbitrary JSON on purpose. A descriptor is
|
||||
something a future provider will put in a prompt, and a nested structure
|
||||
would either be flattened by whoever does that — inconsistently — or
|
||||
smuggle a provider-shaped payload through a story-side field, which is the
|
||||
boundary this package exists to keep.
|
||||
"""
|
||||
if descriptors is None:
|
||||
return {}
|
||||
if not isinstance(descriptors, dict):
|
||||
raise ProfileError("`descriptors` must be a map of trait to value.")
|
||||
if len(descriptors) > MAX_DESCRIPTORS:
|
||||
raise ProfileError(
|
||||
f"A profile may carry at most {MAX_DESCRIPTORS} descriptors."
|
||||
)
|
||||
out: dict[str, str] = {}
|
||||
for trait, value in descriptors.items():
|
||||
if not isinstance(trait, str) or not trait.strip():
|
||||
raise ProfileError("Every descriptor needs a name.")
|
||||
if not isinstance(value, str):
|
||||
raise ProfileError(
|
||||
f"The value for {trait!r} must be text — a profile describes "
|
||||
f"how something looks, in words a person could read back."
|
||||
)
|
||||
if len(value) > MAX_VALUE:
|
||||
raise ProfileError(
|
||||
f"The value for {trait!r} is longer than {MAX_VALUE} characters."
|
||||
)
|
||||
out[trait.strip()[:MAX_KEY]] = value
|
||||
return out
|
||||
|
||||
|
||||
def _checked_features(features) -> list:
|
||||
if features is None:
|
||||
return []
|
||||
if not isinstance(features, list):
|
||||
raise ProfileError("`features` must be a list of short phrases.")
|
||||
if len(features) > MAX_FEATURES:
|
||||
raise ProfileError(f"A profile may carry at most {MAX_FEATURES} features.")
|
||||
out = []
|
||||
for feature in features:
|
||||
if not isinstance(feature, str) or not feature.strip():
|
||||
raise ProfileError("Every feature must be a non-empty phrase.")
|
||||
if len(feature) > MAX_VALUE:
|
||||
raise ProfileError(f"A feature is longer than {MAX_VALUE} characters.")
|
||||
out.append(feature.strip())
|
||||
return out
|
||||
|
||||
|
||||
def _checked_notes(style_notes) -> str:
|
||||
if style_notes is None:
|
||||
return ""
|
||||
if not isinstance(style_notes, str):
|
||||
raise ProfileError("`style_notes` must be text.")
|
||||
if len(style_notes) > MAX_STYLE_NOTES:
|
||||
raise ProfileError(
|
||||
f"Style notes are longer than {MAX_STYLE_NOTES} characters."
|
||||
)
|
||||
return style_notes.strip()
|
||||
@@ -0,0 +1,332 @@
|
||||
"""M10: what a future media provider must satisfy, and nothing that satisfies it.
|
||||
|
||||
No provider is implemented here, none is registered by default, and nothing in
|
||||
this module opens a socket. What it defines is the shape of the boundary, so
|
||||
that adding a real image, video, audio, TTS or STT provider later is writing an
|
||||
adapter rather than editing the story engine.
|
||||
|
||||
## The rule these types exist to enforce
|
||||
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §3: the Story Engine must not call ComfyUI, Stable
|
||||
Diffusion, a video pipeline, a TTS engine or a third-party media API. It states
|
||||
that as a recommendation; this module makes it structural. Everything crossing
|
||||
the boundary is expressed in this vocabulary:
|
||||
|
||||
MediaKind image | video | audio | tts | stt
|
||||
MediaRequest a scene packet, a kind, and neutral hints
|
||||
MediaResult bytes-or-path, a type, and provenance
|
||||
DraftTranscription STT's deliberately different answer (see below)
|
||||
|
||||
**No provider vocabulary appears anywhere in this file or in any story module.**
|
||||
There is no workflow JSON, no sampler name, no CFG scale, no LoRA, no
|
||||
`num_inference_steps`, no Whisper option and no voice id. A provider adapter
|
||||
owns that translation, in its own package, and the story engine never learns it.
|
||||
`test_m10_providers.py` greps the story modules for that vocabulary so the rule
|
||||
cannot rot quietly.
|
||||
|
||||
## Why Protocols rather than base classes
|
||||
|
||||
A future adapter should not have to import from here to be usable — it should
|
||||
merely have to *fit*. `typing.Protocol` gives a structural contract that a test
|
||||
double satisfies as readily as a real ComfyUI adapter, which keeps the seam
|
||||
honest: if the only way to satisfy the interface were to inherit from it, the
|
||||
interface would be describing this codebase rather than the boundary.
|
||||
|
||||
## STT is deliberately shaped differently, and that is the point
|
||||
|
||||
Every other provider returns a `MediaResult` — a depiction of something the
|
||||
story already established. STT returns a `DraftTranscription`, which is a
|
||||
different type on purpose, because it flows the other way:
|
||||
|
||||
audio -> local STT -> draft text -> the reader edits it -> normal submission
|
||||
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §24A states the rule as *"STT output is draft user
|
||||
input, not an accepted story event."* A shared return type would have made it
|
||||
possible to hand a transcription to something expecting a finished artefact, and
|
||||
the asymmetry would have survived only as a comment. `DraftTranscription`
|
||||
carries `editable = True` and has no path into the turn pipeline: the reader's
|
||||
edited text enters through the ordinary action endpoint like anything they
|
||||
typed, and is validated, refereed and snapshotted exactly the same way.
|
||||
|
||||
M10 implements no microphone capture and no transcription. The type boundary is
|
||||
the deliverable.
|
||||
|
||||
## Endpoints: loopback only, and stricter than the narrator's on purpose
|
||||
|
||||
`endpoints.py` already decides which *inference* endpoints this product will
|
||||
talk to, and allows an explicitly configured trusted LAN as well as loopback
|
||||
(ADR 011). Media is not given that latitude. `MEDIA-EXTENSION-CONTRACT.md` §27
|
||||
and §28 set the media default at loopback, with any future LAN extension
|
||||
explicit and user-controlled — so `check_endpoint` below reuses the existing,
|
||||
tested address machinery and then applies the stricter rule on top.
|
||||
|
||||
Reusing rather than reimplementing matters: a second endpoint validator would be
|
||||
a second place for the policy to be wrong, and this one inherits the property
|
||||
that makes the first one hard to talk around — it judges the address a host
|
||||
actually resolves to, not the name.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Protocol, runtime_checkable
|
||||
|
||||
from .. import endpoints
|
||||
|
||||
#: The kinds of media this architecture is required to accommodate. A string
|
||||
#: enum rather than free text, so a typo is a failure here rather than a request
|
||||
#: nothing will ever service.
|
||||
IMAGE = "image"
|
||||
VIDEO = "video"
|
||||
AUDIO = "audio"
|
||||
TTS = "tts"
|
||||
STT = "stt"
|
||||
|
||||
MEDIA_KINDS: tuple[str, ...] = (IMAGE, VIDEO, AUDIO, TTS, STT)
|
||||
|
||||
|
||||
def is_media_kind(value) -> bool:
|
||||
return isinstance(value, str) and value in MEDIA_KINDS
|
||||
|
||||
|
||||
class MediaProviderError(RuntimeError):
|
||||
"""A provider could not do what was asked.
|
||||
|
||||
Deliberately its own type, and deliberately not caught anywhere in the story
|
||||
path: nothing in a turn calls a provider, so there is no code path where
|
||||
this could reach an accepted narration. If a future coordinator catches it,
|
||||
it does so on its own side of the boundary — a failed depiction must leave
|
||||
the story exactly as it was (`MEDIA-EXTENSION-CONTRACT.md` §50).
|
||||
"""
|
||||
|
||||
|
||||
class EndpointRejected(endpoints.EndpointRejected):
|
||||
"""A media endpoint outside the loopback-only media policy.
|
||||
|
||||
Subclasses the inference rejection so that a caller which already handles
|
||||
"this endpoint is not allowed" keeps working, while a caller that wants to
|
||||
tell the two policies apart still can.
|
||||
"""
|
||||
|
||||
|
||||
def endpoint_rejection_reason(url: str) -> str | None:
|
||||
"""Why this URL may not be a media endpoint, or `None` if it may.
|
||||
|
||||
Two rules, in order, and the first is somebody else's:
|
||||
|
||||
1. the existing inference policy — an address in an allowed private network,
|
||||
judged by resolution rather than by name (`endpoints.py`);
|
||||
2. **and** loopback specifically, which is the media contract's stricter
|
||||
default (§27, §28).
|
||||
|
||||
So a trusted-LAN address that an Ollama may legitimately use is refused here.
|
||||
That is not an oversight: narrator inference is a deployment the user has
|
||||
already reasoned about and configured, whereas a media endpoint is a new
|
||||
surface with no v1 use, and the safe default for a surface nobody needs yet
|
||||
is the narrowest one. A future milestone may widen it, explicitly and off by
|
||||
default, which is what §27 requires of any such change.
|
||||
"""
|
||||
reason = endpoints.rejection_reason(url)
|
||||
if reason is not None:
|
||||
return reason
|
||||
if not endpoints.is_loopback(url):
|
||||
return (
|
||||
"A media provider endpoint must be on this machine. "
|
||||
f"{url!r} resolves somewhere else — media generation has no "
|
||||
"trusted-LAN mode, and adding one would be an explicit, "
|
||||
"off-by-default change rather than a setting."
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
def check_endpoint(url: str) -> None:
|
||||
"""Raises `EndpointRejected` unless `url` is an allowed media endpoint."""
|
||||
reason = endpoint_rejection_reason(url)
|
||||
if reason is not None:
|
||||
raise EndpointRejected(reason)
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- the types
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ProviderCapabilities:
|
||||
"""What one provider can do, in neutral terms.
|
||||
|
||||
Deliberately small. `MEDIA-EXTENSION-CONTRACT.md` §25 shows a richer example
|
||||
— seeds, reference images, inpainting — and M10 does not model those,
|
||||
because every one of them is a guess until a provider exists to be asked.
|
||||
What is here is what a coordinator would need in order to choose *whether*
|
||||
to route to this provider at all; anything finer belongs to the adapter and
|
||||
its own capability document.
|
||||
"""
|
||||
|
||||
provider_id: str
|
||||
kinds: tuple[str, ...] = ()
|
||||
#: Free-form, provider-owned, and never interpreted by story code. It exists
|
||||
#: so an adapter can advertise what it supports without this module growing
|
||||
#: a field per feature the ecosystem invents.
|
||||
details: dict = field(default_factory=dict)
|
||||
|
||||
def supports(self, kind: str) -> bool:
|
||||
return kind in self.kinds
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MediaRequest:
|
||||
"""What a coordinator would hand a provider: a scene, a kind, and hints.
|
||||
|
||||
`scene` is a Scene Packet (`packet.build`) — a bounded description of one
|
||||
accepted scene, not the transcript. That is the whole point of the packet
|
||||
existing (`MEDIA-EXTENSION-CONTRACT.md` §12): a provider is given what it
|
||||
needs to depict a moment and no more, which bounds prompt size, keeps
|
||||
providers interchangeable, and means swapping one does not hand a new
|
||||
process the campaign's history.
|
||||
|
||||
`hints` is provider-neutral and optional — an aspect ratio, a duration, a
|
||||
count. It is **not** where a workflow graph or a sampler setting goes; those
|
||||
belong to the adapter, which knows what it is talking to.
|
||||
"""
|
||||
|
||||
kind: str
|
||||
scene: dict
|
||||
hints: dict = field(default_factory=dict)
|
||||
|
||||
def __post_init__(self):
|
||||
if not is_media_kind(self.kind):
|
||||
raise ValueError(
|
||||
f"{self.kind!r} is not one of {', '.join(MEDIA_KINDS)}"
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MediaResult:
|
||||
"""What a provider hands back: a depiction, and where it came from.
|
||||
|
||||
Bytes *or* a path, never both, and the caller says which it wanted. Neither
|
||||
is interpreted here; M10 registers no provider, so nothing constructs one of
|
||||
these outside a test.
|
||||
|
||||
`provenance` carries the scene identity the request named, so that a future
|
||||
asset can always be traced to the accepted position it depicts
|
||||
(`MEDIA-EXTENSION-CONTRACT.md` §48). It is a record of what was asked for —
|
||||
it does not make the depiction true.
|
||||
"""
|
||||
|
||||
kind: str
|
||||
media_type: str
|
||||
provenance: dict = field(default_factory=dict)
|
||||
data: bytes | None = None
|
||||
path: str | None = None
|
||||
details: dict = field(default_factory=dict)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DraftTranscription:
|
||||
"""STT's answer, and deliberately not a `MediaResult`.
|
||||
|
||||
See the module docstring. This is **draft user input**: text the reader is
|
||||
expected to read, correct and submit themselves. It is not an accepted turn,
|
||||
not a state event, not canon, and it has no route into the story that the
|
||||
reader's own typing does not also take.
|
||||
|
||||
`editable` is `True` and there is no constructor that sets it otherwise —
|
||||
it is a statement about what this type *is* rather than a setting, and a
|
||||
reader that finds it false has been handed something that is not a draft.
|
||||
"""
|
||||
|
||||
text: str
|
||||
editable: bool = True
|
||||
confidence: float | None = None
|
||||
details: dict = field(default_factory=dict)
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the protocols
|
||||
|
||||
@runtime_checkable
|
||||
class MediaProvider(Protocol):
|
||||
"""Anything that can depict an accepted scene.
|
||||
|
||||
One protocol covers image, video and audio because the boundary is the same
|
||||
for all three: a bounded scene in, a depiction out, nothing written to the
|
||||
story. What differs between them is entirely inside the adapter.
|
||||
"""
|
||||
|
||||
def capabilities(self) -> ProviderCapabilities: ...
|
||||
|
||||
async def generate(self, request: MediaRequest) -> MediaResult: ...
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class SpeechProvider(Protocol):
|
||||
"""Text to speech: still a depiction, of prose the story already accepted."""
|
||||
|
||||
def capabilities(self) -> ProviderCapabilities: ...
|
||||
|
||||
async def speak(self, text: str, hints: dict | None = None) -> MediaResult: ...
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class TranscriptionProvider(Protocol):
|
||||
"""Speech to text, which runs the other way and returns a draft.
|
||||
|
||||
The signature is the asymmetry: it takes audio and returns
|
||||
`DraftTranscription`, so no coordinator can hand its output to something
|
||||
expecting a finished artefact, and nothing can mistake it for an accepted
|
||||
turn.
|
||||
"""
|
||||
|
||||
def capabilities(self) -> ProviderCapabilities: ...
|
||||
|
||||
async def transcribe(
|
||||
self, audio: bytes, hints: dict | None = None
|
||||
) -> DraftTranscription: ...
|
||||
|
||||
|
||||
# -------------------------------------------------------------- the registry
|
||||
|
||||
#: Registered providers, by id. **Empty, and empty on purpose.**
|
||||
#:
|
||||
#: M10 ships no provider, so nothing is registered at import, nothing is
|
||||
#: required at startup, and no configuration is read. `test_m10_no_media.py`
|
||||
#: asserts this is empty after the application has been imported and a campaign
|
||||
#: has been played — media readiness has to be inert until something explicitly
|
||||
#: uses it.
|
||||
_REGISTRY: dict[str, object] = {}
|
||||
|
||||
|
||||
def register(provider_id: str, provider: object) -> None:
|
||||
"""Makes a provider available to a future coordinator.
|
||||
|
||||
Exists to prove the claim in M10's Definition of Done — that a provider can
|
||||
be added *without modifying story authority or history* — by being the only
|
||||
thing an adapter has to call. Nothing in `app/routers`, `app/narrative`,
|
||||
`app/context` or `app/tree` imports this module, so registering one cannot
|
||||
reach them.
|
||||
"""
|
||||
if not isinstance(provider_id, str) or not provider_id.strip():
|
||||
raise ValueError("a provider needs an id")
|
||||
_REGISTRY[provider_id] = provider
|
||||
|
||||
|
||||
def unregister(provider_id: str) -> None:
|
||||
_REGISTRY.pop(provider_id, None)
|
||||
|
||||
|
||||
def registered() -> dict[str, object]:
|
||||
"""The registry, copied — callers must not mutate it in place."""
|
||||
return dict(_REGISTRY)
|
||||
|
||||
|
||||
def for_kind(kind: str) -> list[object]:
|
||||
"""Every registered provider advertising `kind`. Empty in v1."""
|
||||
out = []
|
||||
for provider in _REGISTRY.values():
|
||||
caps = getattr(provider, "capabilities", None)
|
||||
if caps is None:
|
||||
continue
|
||||
try:
|
||||
if caps().supports(kind):
|
||||
out.append(provider)
|
||||
except Exception: # noqa: BLE001 - a broken adapter is not this layer's
|
||||
continue
|
||||
return out
|
||||
@@ -433,6 +433,34 @@ MIGRATIONS: list[tuple[int, str | dict[str, str]]] = [
|
||||
# Such a campaign opens with an empty library and needs no source to play.
|
||||
(92, {"sqlite": fts.DDL,
|
||||
"default": "-- FTS5 is SQLite-only; this build stores campaigns in SQLite"}),
|
||||
|
||||
# M10 adds **no migration**, and that is the whole of its schema story.
|
||||
#
|
||||
# `visual_profiles` is a new table, so `create_all` builds it on every path
|
||||
# — fresh install, existing database, test setup — exactly as it did for
|
||||
# `memories`, `branches`, `checkpoints`, `summaries` and the knowledge
|
||||
# tables. Its one index is declared on the column (`index=True`) rather than
|
||||
# in `__table_args__`, so `create_all` builds that too, which is what
|
||||
# version 92's note above says about the M7 tables: when the index is on the
|
||||
# column there is nothing left for a `CREATE INDEX` here to do.
|
||||
#
|
||||
# A version 93 was written here first, adding
|
||||
# `ix_visual_profiles_adventure`. It was wrong, and the M10 suite's
|
||||
# fresh-versus-upgraded comparison is what found it: an upgraded database
|
||||
# ended up with that index *and* the `ix_visual_profiles_adventure_id` that
|
||||
# `create_all` had already made, while a fresh install had only the latter.
|
||||
# Two schemas that differ by which path the file took is the thing a
|
||||
# migration exists to prevent, and the redundant index was the only
|
||||
# difference between them.
|
||||
#
|
||||
# **No backfill, and there is nothing that could be backfilled.** A profile
|
||||
# says what an entity looks like, and no existing column holds that: the
|
||||
# narrative state records what entities *are* — type, status, description,
|
||||
# location — and inventing an appearance from a description would be
|
||||
# fabricating exactly the kind of visual detail
|
||||
# `MEDIA-EXTENSION-CONTRACT.md` §37 says must never appear without the
|
||||
# reader asking for it. An M9 campaign therefore opens with no profiles,
|
||||
# which is what such a campaign had, and plays unchanged without any.
|
||||
]
|
||||
|
||||
LATEST_VERSION = max((v for v, _ in MIGRATIONS), default=1)
|
||||
|
||||
@@ -246,6 +246,13 @@ class Adventure(Base):
|
||||
cascade="all, delete-orphan",
|
||||
order_by="KnowledgeSource.id",
|
||||
)
|
||||
# M10: how the campaign's entities look. Derived presentation metadata, not
|
||||
# story state — see `VisualProfile`.
|
||||
visual_profiles: Mapped[list["VisualProfile"]] = relationship(
|
||||
back_populates="adventure",
|
||||
cascade="all, delete-orphan",
|
||||
order_by="VisualProfile.id",
|
||||
)
|
||||
|
||||
|
||||
class Branch(Base):
|
||||
@@ -802,6 +809,108 @@ class KnowledgeEmbedding(Base):
|
||||
chunk: Mapped[KnowledgeChunk] = relationship(back_populates="embedding")
|
||||
|
||||
|
||||
class VisualProfile(Base):
|
||||
"""M10: how one entity looks, so a future depiction can be consistent.
|
||||
|
||||
The only thing M10 persists, and the reason is that it was the only thing
|
||||
the media contract asks for that nothing already stored. The scene snapshot
|
||||
§5 asks for already exists as `narrative_state["scene"]` and has since M5;
|
||||
building a second one beside it would have been a duplicate representation
|
||||
with its own lineage rules to get wrong.
|
||||
|
||||
## Not story state, and structurally so
|
||||
|
||||
A visual profile is **presentation metadata**. Nothing here is a fact the
|
||||
story established: `MEDIA-EXTENSION-CONTRACT.md` §35 and §37 are explicit
|
||||
that a depiction — and therefore a description written to guide one — must
|
||||
never become canon on its own, and that promoting a visual detail into canon
|
||||
would have to be a deliberate act by the reader.
|
||||
|
||||
So these rows are deliberately **outside** the M5 pipeline. They are not
|
||||
events, they are not validated by `narrative/validate.py`, they are not in
|
||||
the state document, and they are not snapshotted per position. Writing one
|
||||
cannot change `narrative_state`, because nothing in `media/` imports the
|
||||
code that may. That is the guarantee, and it is a structural one rather than
|
||||
a rule somebody has to remember.
|
||||
|
||||
## Campaign-scoped, not per-position — which is the interesting decision
|
||||
|
||||
Every other derived record in this schema carries a `(branch_id, depth)`
|
||||
coordinate, because it describes a *moment*: a memory summarises a stretch,
|
||||
a summary covers a range, a snapshot records an outcome. A visual profile
|
||||
describes none of those. It says what someone looks like, and a character
|
||||
does not change appearance because the story forked.
|
||||
|
||||
Making it per-position would have been actively wrong twice over. It would
|
||||
have meant a profile written on one branch was invisible on another, so a
|
||||
reader who diverged would lose their cast's appearance — the opposite of the
|
||||
continuity the profile exists for. And it would have put a descriptor
|
||||
document into every per-position state snapshot, which M9 measured as
|
||||
already 74% of a campaign bundle; the profiles would have been duplicated
|
||||
once per turn to say something that never varies.
|
||||
|
||||
So the key is `(adventure_id, entity_key)` and there is exactly one profile
|
||||
per entity per campaign. It is stable across Undo, Redo, Save Point restore
|
||||
and divergence for the same reason it is simple: there is nothing there to
|
||||
move.
|
||||
|
||||
## `entity_key` is the M5 key, and no second identity namespace
|
||||
|
||||
The key is the entity key the narrative state already uses — `"mara"`,
|
||||
`"the_office"`, `"silver_key"` — not a new id, not a name, and not a media
|
||||
identifier. `MEDIA-EXTENSION-CONTRACT.md` §7-9 describe character, location
|
||||
and item profiles separately; this is one table for all three, because M5's
|
||||
entity model is genre-neutral by design (`DATA-MODEL.md` §9) and a
|
||||
character, a location, an item, a vehicle and a spaceship are all entities
|
||||
with a `type`. Splitting them here would have reintroduced the genre shape
|
||||
M5 spent a milestone removing.
|
||||
|
||||
There is no `kind` column for the same reason: the entity already has a
|
||||
`type`, and storing it again would be a second source of truth for one fact.
|
||||
|
||||
## The columns, and why they are shaped this way
|
||||
|
||||
The contract's examples are fantasy-shaped — hair, eyes, build; architecture,
|
||||
hearths, oil lamps — and the brief is explicit that they are examples rather
|
||||
than a schema. A fixed column per fantasy attribute would not hold an
|
||||
orbital station, a corporate office or a car.
|
||||
|
||||
So: `descriptors` is an open map of trait to value, `features` is a list of
|
||||
distinctive visible things, and `style_notes` is free text about how it
|
||||
should be rendered. `{"hair": "dark auburn"}` and
|
||||
`{"hull": "pitted white composite"}` are the same shape, and neither needed
|
||||
a migration to become possible.
|
||||
"""
|
||||
|
||||
__tablename__ = "visual_profiles"
|
||||
__table_args__ = (
|
||||
# One profile per entity per campaign. The uniqueness is the model: a
|
||||
# second profile for the same entity would be a second answer to "what
|
||||
# does this look like", with nothing to decide between them.
|
||||
UniqueConstraint("adventure_id", "entity_key", name="uq_visual_entity"),
|
||||
)
|
||||
|
||||
id: Mapped[int] = mapped_column(primary_key=True)
|
||||
adventure_id: Mapped[int] = mapped_column(
|
||||
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
|
||||
)
|
||||
#: The narrative-state entity key. Not a display name: two characters may
|
||||
#: share a name, and M9's report recorded that the state model permits it.
|
||||
entity_key: Mapped[str] = mapped_column(String(200))
|
||||
#: Trait -> value. Open by construction; see the class docstring.
|
||||
descriptors: Mapped[dict] = mapped_column(JSON, default=dict)
|
||||
#: Distinctive visible things, as short phrases.
|
||||
features: Mapped[list] = mapped_column(JSON, default=list)
|
||||
#: How it should be rendered, rather than what it is.
|
||||
style_notes: Mapped[str] = mapped_column(Text, default="")
|
||||
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
|
||||
updated_at: Mapped[datetime] = mapped_column(
|
||||
DateTime, default=utcnow, onupdate=utcnow
|
||||
)
|
||||
|
||||
adventure: Mapped[Adventure] = relationship(back_populates="visual_profiles")
|
||||
|
||||
|
||||
class StoryCard(Base):
|
||||
"""Owned by either a scenario or an adventure (exactly one set)."""
|
||||
|
||||
|
||||
@@ -42,6 +42,7 @@ from . import ( # noqa: F401
|
||||
memories,
|
||||
actions,
|
||||
knowledge,
|
||||
visuals,
|
||||
)
|
||||
from ... import limits # noqa: F401 `adventures.limits` is patched by tests.
|
||||
from .crud import SNIPPET_MAX, _snippet
|
||||
|
||||
@@ -1,7 +1,31 @@
|
||||
"""Exporting an adventure to a bundle, and importing one back.
|
||||
|
||||
`app/bundle.py` owns the format and the version handling. These two endpoints
|
||||
only check ownership and hand the work over.
|
||||
only check ownership, apply the caps, and hand the work over.
|
||||
|
||||
## Why the import is one transaction and two phases
|
||||
|
||||
`bundle.plan` reads the whole file and returns a checked, normalised tree
|
||||
without opening a session, touching a row or creating an adventure. Everything a
|
||||
hand-edited file can get wrong about its own shape — a node on a branch that is
|
||||
not listed, a fork from a branch listed after it, a head past the story, an
|
||||
audit record naming a turn that is not there — is a 400 from a function with no
|
||||
side effects.
|
||||
|
||||
Only then does `bundle.materialize` write, and it writes inside the single
|
||||
transaction this endpoint commits at the end. So there are exactly two outcomes
|
||||
a caller can see, and M9 requires them to be distinguishable:
|
||||
|
||||
the authoritative import failed 4xx, and no campaign exists
|
||||
the authoritative import succeeded 201, and the campaign is complete
|
||||
|
||||
A third state — the campaign landed and a *rebuildable* index did not — is not a
|
||||
failure of the import and does not roll it back. Passages, the lexical index and
|
||||
vectors are all a deterministic function of content the file carries, so losing
|
||||
them costs a rebuild rather than data. It is reported on the response as a
|
||||
warning, it is visible per source in the Knowledge panel, and Reindex is the
|
||||
repair. Refusing a whole campaign because a search index would not build would
|
||||
trade the valuable thing for the cheap one.
|
||||
"""
|
||||
|
||||
from fastapi import Body, Depends, Request
|
||||
@@ -18,15 +42,15 @@ def export_adventure(
|
||||
db: Session = Depends(get_db),
|
||||
adv: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
"""Returns a full backup: plot components, story cards, scripts, state, and tree.
|
||||
"""Returns a full backup: the story, the tree, the state, and the evidence.
|
||||
|
||||
`app/bundle.py` owns the format, in both of its versions. A backup outlives
|
||||
the schema, so no call site decides anything about its shape.
|
||||
`app/bundle.py` owns the format, in all three of its versions. A backup
|
||||
outlives the schema, so no call site decides anything about its shape.
|
||||
"""
|
||||
return bundle.export(db, adv)
|
||||
|
||||
|
||||
@router.post("/import", response_model=schemas.AdventureOut, status_code=201)
|
||||
@router.post("/import", response_model=schemas.ImportedAdventureOut, status_code=201)
|
||||
def import_adventure(
|
||||
request: Request,
|
||||
payload: dict = Body(...),
|
||||
@@ -58,18 +82,29 @@ def import_adventure(
|
||||
branches=story["branches"],
|
||||
)
|
||||
|
||||
adventure = bundle.materialize(db, payload, story, user.id)
|
||||
|
||||
db.commit()
|
||||
try:
|
||||
adventure, report = bundle.materialize(db, payload, story, user.id)
|
||||
db.commit()
|
||||
except Exception:
|
||||
# Explicit, rather than left to the session closing. The planner has
|
||||
# already refused everything it can see, so anything raising here is a
|
||||
# write that surprised us — the case where leaving a partial campaign
|
||||
# behind would be worst, and the case a test can only assert on if the
|
||||
# rollback is a statement rather than a side effect of teardown.
|
||||
db.rollback()
|
||||
raise
|
||||
db.refresh(adventure)
|
||||
# A campaign exported while undone imports undone (M3), so the history
|
||||
# controls have to be right on the response that opens it — otherwise the
|
||||
# first thing the reader sees about a story with a retained future is a
|
||||
# greyed-out Redo.
|
||||
out = schemas.AdventureOut.model_validate(adventure)
|
||||
out = schemas.ImportedAdventureOut.model_validate(adventure)
|
||||
out.can_undo = head.can_undo(db, adventure)
|
||||
out.can_redo = head.can_redo(db, adventure)
|
||||
# This is not a funnel step. A returning player imports a bundle, so it
|
||||
# says nothing about how far a first-time visitor got. It is counted anyway,
|
||||
# because it is the clearest evidence that anyone uses the export format.
|
||||
out.import_warnings = [
|
||||
f"The search index for “{failure['title']}” could not be rebuilt "
|
||||
f"({failure['detail']}). The file itself imported intact — use Reindex "
|
||||
f"in the Knowledge panel to try again."
|
||||
for failure in report["knowledge_index_failures"]
|
||||
]
|
||||
return out
|
||||
|
||||
@@ -14,6 +14,8 @@ from ... import (
|
||||
worldstate,
|
||||
)
|
||||
from ...database import get_db
|
||||
from ...knowledge import embeddings as knowledge_embeddings
|
||||
from ...knowledge import importer as knowledge_importer
|
||||
|
||||
from .deps import CurrentUser, current_adventure, router
|
||||
from .paging import action_window, annotate_takes
|
||||
@@ -348,8 +350,14 @@ def delete_adventure(
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
# M9. The lexical index first, while the chunks that locate it still exist.
|
||||
# It is a virtual table, so nothing cascades into it, and an orphaned index
|
||||
# row makes the *next* import into *any* campaign fail — see
|
||||
# `knowledge.importer.clear_campaign_index`.
|
||||
knowledge_importer.clear_campaign_index(db, adventure)
|
||||
db.delete(adventure)
|
||||
db.commit()
|
||||
# No later request reads this adventure's vectors, so drop them now. The
|
||||
# cache would otherwise hold them until the process restarted.
|
||||
memorybank.forget_cached_vectors(adventure_id)
|
||||
knowledge_embeddings.forget_cached(adventure_id)
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
"""M10: reading and writing how a campaign's entities look.
|
||||
|
||||
Four endpoints on the campaign, and one on the scene beneath it. They are the
|
||||
only reader-facing surface M10 adds, and they are an API surface rather than a
|
||||
browser one: M10 builds no gallery, no picker and no preview, because there is
|
||||
nothing to generate and a screen for configuring depictions nobody can make
|
||||
would be a feature pretending to be a seam.
|
||||
|
||||
## Why a scene-packet endpoint exists at all
|
||||
|
||||
`GET .../scene-packet` returns exactly what a future media coordinator would be
|
||||
handed (`media/packet.py`). Nothing in v1 calls it, and it generates nothing.
|
||||
|
||||
It is here because it is the one part of M10 whose *contents* are a
|
||||
correctness claim — that a provider is given a bounded view and not the
|
||||
campaign, and that narrator-only material does not travel through it. A claim
|
||||
like that should be inspectable by whoever is reviewing the boundary, not only
|
||||
by a test that imports a private function. It is a read: it writes nothing,
|
||||
emits no event, and cannot move the head.
|
||||
|
||||
## What these endpoints deliberately are not
|
||||
|
||||
They are not a state API. A visual profile is presentation metadata and writing
|
||||
one changes no story fact (`models.VisualProfile`), so there is no event, no
|
||||
proposal, no snapshot and no head movement anywhere below here. The separation
|
||||
is structural — this module reaches `media.profiles`, and that module imports
|
||||
nothing that can write authoritative state.
|
||||
"""
|
||||
|
||||
from fastapi import Body, Depends, HTTPException
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from ... import models
|
||||
from ...database import get_db
|
||||
from ...media import packet as scene_packet
|
||||
from ...media import profiles as visual_profiles
|
||||
|
||||
from .deps import current_adventure, router
|
||||
|
||||
|
||||
@router.get("/{adventure_id}/visual-profiles")
|
||||
def list_visual_profiles(
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
"""Every visual profile in the campaign, by entity key.
|
||||
|
||||
Campaign-scoped rather than scoped to the story being read, because that is
|
||||
what a profile is: a character does not change appearance when the story
|
||||
forks, so there is no position for this list to be relative to.
|
||||
"""
|
||||
return {
|
||||
"profiles": [
|
||||
{"entity_key": row.entity_key, **visual_profiles.as_dict(row)}
|
||||
for row in visual_profiles.all_for(db, adventure)
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
@router.put("/{adventure_id}/visual-profiles/{entity_key}")
|
||||
def set_visual_profile(
|
||||
entity_key: str,
|
||||
payload: dict = Body(...),
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
"""Records how one entity looks. Replaces any existing profile.
|
||||
|
||||
A `PUT` rather than a `PATCH`, and the whole profile rather than a delta,
|
||||
for the reason `profiles.set_profile` gives: merging would make a descriptor
|
||||
impossible to remove.
|
||||
|
||||
The entity must exist in the campaign's state at the active head. A 400 for
|
||||
a name nobody has is better than a row describing nobody, which would then
|
||||
be invisible until a future depiction quietly ignored it.
|
||||
"""
|
||||
try:
|
||||
row = visual_profiles.set_profile(
|
||||
db, adventure, entity_key,
|
||||
descriptors=payload.get("descriptors"),
|
||||
features=payload.get("features"),
|
||||
style_notes=payload.get("style_notes"),
|
||||
)
|
||||
except visual_profiles.ProfileError as exc:
|
||||
raise HTTPException(400, str(exc)) from exc
|
||||
db.commit()
|
||||
db.refresh(row)
|
||||
return {"entity_key": row.entity_key, **visual_profiles.as_dict(row)}
|
||||
|
||||
|
||||
@router.get("/{adventure_id}/visual-profiles/{entity_key}")
|
||||
def read_visual_profile(
|
||||
entity_key: str,
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
row = visual_profiles.get_profile(db, adventure, entity_key)
|
||||
if row is None:
|
||||
raise HTTPException(404, f"No visual profile for {entity_key!r}.")
|
||||
return {"entity_key": row.entity_key, **visual_profiles.as_dict(row)}
|
||||
|
||||
|
||||
@router.delete("/{adventure_id}/visual-profiles/{entity_key}", status_code=204)
|
||||
def delete_visual_profile(
|
||||
entity_key: str,
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
"""Removes a description. Never the entity, which lives in the state."""
|
||||
if not visual_profiles.delete_profile(db, adventure, entity_key):
|
||||
raise HTTPException(404, f"No visual profile for {entity_key!r}.")
|
||||
db.commit()
|
||||
|
||||
|
||||
@router.get("/{adventure_id}/scene-packet")
|
||||
def read_scene_packet(
|
||||
start: int | None = None,
|
||||
end: int | None = None,
|
||||
db: Session = Depends(get_db),
|
||||
adventure: models.Adventure = Depends(current_adventure),
|
||||
):
|
||||
"""What a future media provider would be given for the current scene.
|
||||
|
||||
`start` and `end` are depths on the active branch, and both are optional:
|
||||
omitted, the packet describes the scene at the position the story last set
|
||||
one. Passing a range is what a future video request would do — a scene is
|
||||
not assumed to be one turn (`MEDIA-EXTENSION-CONTRACT.md` §30-31).
|
||||
|
||||
Generates nothing and contacts nothing. There is no provider to send it to.
|
||||
"""
|
||||
return scene_packet.build(db, adventure, start=start, end=end)
|
||||
@@ -0,0 +1,75 @@
|
||||
"""M9: taking a verified copy of the whole database, from the browser.
|
||||
|
||||
Two endpoints and no third. `app/backup.py` owns the procedure and every
|
||||
guarantee it makes; these only decide who may ask.
|
||||
|
||||
## Why there is no restore endpoint, and no download
|
||||
|
||||
**Restore** means replacing the database file the running process has open.
|
||||
Doing that from inside that process is how someone loses both copies at once:
|
||||
the connection pool still holds handles on the old file, the WAL belongs to the
|
||||
old file, and a half-swapped database is not something a running application can
|
||||
notice. The supported procedure is in `DEVELOPMENT.md` — stop the application,
|
||||
move the file into place, start it — and it is a procedure precisely because
|
||||
each step needs the application not to be running. Campaign-level recovery, the
|
||||
common case and the only one that crosses machines, is the export bundle.
|
||||
|
||||
**Download** is not offered either. The file is a copy of every campaign on the
|
||||
machine, and streaming it through the browser would put it in the download
|
||||
directory, in the browser's own cache, and in whatever the reader does with it
|
||||
next — for a local single-user application whose whole premise is that the story
|
||||
does not leave the machine, that is a worse default than a path the reader can
|
||||
copy. So the response names the directory and the reader takes it from there.
|
||||
|
||||
## Where the file goes
|
||||
|
||||
Nowhere a request can name. The destination is derived from the database the
|
||||
application is already using, and the filename is generated from the clock. No
|
||||
part of either comes from the caller, so there is no traversal to attempt (H08),
|
||||
and the endpoints below accept no body at all.
|
||||
"""
|
||||
|
||||
import logging
|
||||
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
|
||||
from .. import auth, backup, models
|
||||
|
||||
router = APIRouter(prefix="/api/backups", tags=["backups"])
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
@router.get("")
|
||||
def list_backups(_user: models.User = Depends(auth.get_current_user)):
|
||||
"""The backups already on disk, newest first, and where they are.
|
||||
|
||||
The directory is reported once here rather than on every row, because it is
|
||||
the same for all of them and it is what the reader needs in order to find
|
||||
the files at all.
|
||||
"""
|
||||
return {
|
||||
"directory": str(backup.directory()),
|
||||
"backups": backup.existing(),
|
||||
}
|
||||
|
||||
|
||||
@router.post("", status_code=201)
|
||||
def create_backup(_user: models.User = Depends(auth.get_current_user)):
|
||||
"""Takes one verified backup, and reports what it wrote.
|
||||
|
||||
Synchronous. A backup of a local single-user database is a page copy that
|
||||
finishes in well under a second, and a reader who pressed the button is
|
||||
entitled to be told whether it worked rather than to be told it started.
|
||||
|
||||
A failure is a 500 carrying the reason. There is nothing for the caller to
|
||||
fix by retrying differently — the request has no parameters — so the useful
|
||||
thing is the message, and `backup.create` guarantees that the source database
|
||||
is untouched and no partial file is left behind.
|
||||
"""
|
||||
try:
|
||||
result = backup.create()
|
||||
except backup.BackupError as exc:
|
||||
log.error("Backup failed: %s", exc)
|
||||
raise HTTPException(500, str(exc)) from exc
|
||||
return {"directory": str(result.path.parent), **result.as_dict()}
|
||||
@@ -497,6 +497,25 @@ class AdventureOut(ORMModel):
|
||||
can_redo: bool = False
|
||||
|
||||
|
||||
class ImportedAdventureOut(AdventureOut):
|
||||
"""A campaign that has just been restored from a bundle (M9).
|
||||
|
||||
Exactly `AdventureOut` plus what could not be rebuilt. The extra field is on
|
||||
a subclass rather than on the base, because "which of your search indexes
|
||||
failed to rebuild" is a fact about one import and not a property of a
|
||||
campaign — putting it on `AdventureOut` would attach it to every read of
|
||||
every campaign forever.
|
||||
|
||||
An empty list is the ordinary answer and means the whole campaign, its
|
||||
evidence and its derived indexes all landed. A non-empty one means the
|
||||
authoritative import succeeded and a rebuildable index did not, which is a
|
||||
distinction M9 requires a caller to be able to draw: the campaign is intact,
|
||||
and Reindex is the repair.
|
||||
"""
|
||||
|
||||
import_warnings: list[str] = []
|
||||
|
||||
|
||||
class ActionPage(BaseModel):
|
||||
"""A slice of the story, counted back from the newest action."""
|
||||
|
||||
|
||||
@@ -63,7 +63,9 @@ def give(db: Session, user: models.User) -> models.Adventure | None:
|
||||
# flush whatever part of the adventure the session still held.
|
||||
with db.begin_nested():
|
||||
story = bundle.plan(payload, bundle.check_format(payload))
|
||||
adventure = bundle.materialize(db, payload, story, user.id)
|
||||
# The starter ships with no imported knowledge, so the derived
|
||||
# report is always empty here and nothing reads it.
|
||||
adventure, _ = bundle.materialize(db, payload, story, user.id)
|
||||
_link_scenario(db, adventure, payload)
|
||||
return adventure
|
||||
except Exception:
|
||||
|
||||
@@ -0,0 +1,136 @@
|
||||
"""M10: a campaign with a scene worth depicting, deliberately not a fantasy one.
|
||||
|
||||
The M10 brief asks for at least one non-fantasy representation, and the reason
|
||||
is a real risk rather than a preference: the media contract's own examples are
|
||||
fantasy-shaped — hair and eyes, timber framing, oil lamps — and a schema written
|
||||
while looking at them can acquire that shape without anyone deciding to give it
|
||||
one. So the fixture is four people in an office, and the same code has to hold
|
||||
it with no change.
|
||||
|
||||
Bill the protagonist
|
||||
Alice a coworker, with a visual profile
|
||||
Roger a coworker, with no profile at all
|
||||
John a coworker who is not in the room
|
||||
|
||||
the office a location, with a visual profile
|
||||
a badge an item Bill is carrying
|
||||
the server room a second location, for divergence
|
||||
|
||||
The cast is the one from the post-M8 playtest finding, and that is deliberate
|
||||
too — but only as *shape*. M10 does not investigate that finding, and nothing
|
||||
here asserts anything about coreference; it is M11's, and §23 of the brief says
|
||||
so. What the shape buys here is a scene with three present characters and one
|
||||
absent, which is what makes "the packet describes who is in the room" a claim
|
||||
with a wrong answer available.
|
||||
|
||||
Roger having no profile is load-bearing: it is how the tests tell "no profile"
|
||||
from "an empty profile", which a future provider has to be able to distinguish.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
#: A narrator-only secret, used by the hidden-information tests. It is imported
|
||||
#: as an M7 hidden knowledge source — the product's real mechanism for
|
||||
#: narrator-only material — rather than as an invented marker, so the test
|
||||
#: exercises the boundary that actually exists.
|
||||
SECRET_SENTINEL = "ZARQUON-CONCEALED-OBSERVER-7731"
|
||||
|
||||
SECRET_MD = f"""# What nobody in the room knows
|
||||
|
||||
There is a concealed observer behind the north wall of the office, watching the
|
||||
meeting through a gap in the panelling. Their code name is {SECRET_SENTINEL}.
|
||||
|
||||
Nobody present is aware of this.
|
||||
"""
|
||||
|
||||
#: A source that is *not* hidden, so a test can show the packet excludes
|
||||
#: imported knowledge as a class rather than only excluding secrets.
|
||||
HANDBOOK_MD = """# Office handbook
|
||||
|
||||
The building was refurbished in the spring. The north wall panelling is new.
|
||||
"""
|
||||
|
||||
|
||||
def play(client, adv_id, text, events, prose="The meeting continues."):
|
||||
ScriptedProvider.replies = [f"{prose}\n" + state_block(events)]
|
||||
response = client.post(
|
||||
f"/api/adventures/{adv_id}/actions", json={"type": "do", "text": text}
|
||||
)
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
return response
|
||||
|
||||
|
||||
def entity(key, kind, name):
|
||||
return {"type": "create_entity", "entity": key, "entity_type": kind,
|
||||
"name": name}
|
||||
|
||||
|
||||
def build(client, adv_id) -> dict:
|
||||
"""Plays the office campaign and returns what a test needs to check it.
|
||||
|
||||
Leaves the campaign with a scene set at the active head, two visual
|
||||
profiles, one character deliberately unprofiled, and one character
|
||||
deliberately not present.
|
||||
"""
|
||||
play(client, adv_id, "arrive at the office", [
|
||||
entity("bill", "character", "Bill"),
|
||||
entity("alice", "character", "Alice"),
|
||||
entity("roger", "character", "Roger"),
|
||||
entity("john", "character", "John"),
|
||||
entity("office", "location", "The office"),
|
||||
entity("server_room", "location", "The server room"),
|
||||
entity("badge", "item", "Security badge"),
|
||||
])
|
||||
play(client, adv_id, "start the meeting", [
|
||||
{"type": "set_possession", "item": "badge", "owner": "bill"},
|
||||
{"type": "set_scene",
|
||||
"summary": "Bill, Alice and Roger meet around the table.",
|
||||
"location": "office",
|
||||
"present": ["bill", "alice", "roger"]},
|
||||
])
|
||||
|
||||
profiles = {
|
||||
"alice": {
|
||||
"descriptors": {"build": "tall", "hair": "short black",
|
||||
"clothing": "grey blazer"},
|
||||
"features": ["tortoiseshell glasses"],
|
||||
"style_notes": "photographic, natural light",
|
||||
},
|
||||
"office": {
|
||||
"descriptors": {"architecture": "open-plan floor",
|
||||
"lighting": "flat fluorescent"},
|
||||
"features": ["whiteboard covered in diagrams"],
|
||||
"style_notes": "",
|
||||
},
|
||||
}
|
||||
for key, profile in profiles.items():
|
||||
response = client.put(
|
||||
f"/api/adventures/{adv_id}/visual-profiles/{key}", json=profile
|
||||
)
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
return {"profiles": profiles}
|
||||
|
||||
|
||||
def upload_secret(client, adv_id) -> int:
|
||||
"""Imports the narrator-only source the hidden-information tests use."""
|
||||
return _upload(client, adv_id, "observer.md", SECRET_MD, "canon",
|
||||
visibility="hidden")
|
||||
|
||||
|
||||
def upload_handbook(client, adv_id) -> int:
|
||||
return _upload(client, adv_id, "handbook.md", HANDBOOK_MD, "reference")
|
||||
|
||||
|
||||
def _upload(client, adv_id, name, body, classification, **fields):
|
||||
data = {"classification": classification}
|
||||
data.update({k: str(v).lower() if isinstance(v, bool) else str(v)
|
||||
for k, v in fields.items()})
|
||||
response = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": (name, body.encode("utf-8"), "text/markdown")},
|
||||
data=data,
|
||||
)
|
||||
assert response.status_code == 201, response.text[:400]
|
||||
return response.json()["id"]
|
||||
@@ -0,0 +1,406 @@
|
||||
"""M9: one campaign that exercises every portable data family at once.
|
||||
|
||||
`TEST-CAMPAIGN-FIXTURE.md` describes the standard Continuity Test campaign, and
|
||||
the acceptance suites use it. This is a different thing and does not replace it:
|
||||
the Continuity Test is shaped to read like a story, and this one is shaped to
|
||||
break a round trip. Every property M9 promises has a source in this campaign that
|
||||
would be silently lost by a plausible mistake in the exporter or the importer.
|
||||
|
||||
Opening
|
||||
|
|
||||
+-- normal turns transcript, state events, snapshots
|
||||
+-- Retry two takes at one coordinate
|
||||
+-- knowledge retrieval imported passages in a stored prompt
|
||||
+-- Save Point S1 a named coordinate on the first line
|
||||
+-- more turns a future the reader will leave
|
||||
|
|
||||
+-- Undo x2 the head steps back
|
||||
|
|
||||
+-- divergent continuation a second branch, and a second future
|
||||
+-- Save Point S2 a named coordinate on the second line
|
||||
+-- manual state correction an event nothing narrated
|
||||
+-- Undo x1 the head ends behind the newest row
|
||||
|
||||
The shape is chosen so that no single fact identifies a position. The active head
|
||||
is not the newest row, not the deepest row, not the last row written, and not on
|
||||
the branch that holds the most story — an importer that guesses any one of those
|
||||
lands somewhere else.
|
||||
|
||||
Two campaigns are built, not one. `build` returns the rich campaign; the fixture
|
||||
also leaves a neighbour beside it, because a bundle that accidentally exported
|
||||
another campaign's rows would otherwise export nothing and pass.
|
||||
|
||||
The builder speaks HTTP throughout. A fixture that wrote rows directly would
|
||||
prove the exporter can read what the fixture wrote, which is not the claim.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
|
||||
from app import memorybank
|
||||
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
# --------------------------------------------------------------- source files
|
||||
# Three imported sources, one per class, plus the two lifecycle states that a
|
||||
# round trip most easily loses: a source someone switched off, and one only the
|
||||
# narrator may see.
|
||||
|
||||
CANON_MD = """# Westhaven
|
||||
|
||||
## The Old Abbey
|
||||
|
||||
The abbey above Westhaven has stood since the founding. Its crypt is sealed,
|
||||
and the seal has never been broken.
|
||||
|
||||
## What cannot happen here
|
||||
|
||||
The dead do not return. No rite, relic or bargain in Westhaven has ever
|
||||
returned anyone from death, and none ever will.
|
||||
"""
|
||||
|
||||
REFERENCE_MD = """# The Crooked Lantern
|
||||
|
||||
The tavern on Fen Street is timber-framed, low-beamed, and older than the
|
||||
street it stands on. The hearth is never allowed to go out.
|
||||
|
||||
## The keeper
|
||||
|
||||
Mara keeps the Crooked Lantern. She was born in Westhaven and has never left
|
||||
it.
|
||||
"""
|
||||
|
||||
INSPIRATION_MD = """# Weather notes
|
||||
|
||||
Rain on shutters. Lantern light through wet glass. The smell of a hearth
|
||||
banked for the night.
|
||||
"""
|
||||
|
||||
SECRET_MD = """# The seal
|
||||
|
||||
The abbey seal was broken once, sixty years ago, and set again by a hand that
|
||||
is still alive. Nobody in Westhaven knows this.
|
||||
"""
|
||||
|
||||
DISABLED_MD = """# Discarded draft
|
||||
|
||||
An earlier draft of the Westhaven material, kept for reference and switched off
|
||||
so it cannot reach the narrator.
|
||||
"""
|
||||
|
||||
#: The campaign's own rule, so the correction and the canon block have something
|
||||
#: real to be measured against.
|
||||
CAMPAIGN_CANON = {"rules": ["The dead do not return."]}
|
||||
|
||||
OPENING = "Aldric sits in the Crooked Lantern with Mara, and the rain starts."
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- helpers
|
||||
|
||||
def _play(client, adv_id, text, prose, events=None, kind="do"):
|
||||
ScriptedProvider.replies = [f"{prose}\n{state_block(events or [])}"]
|
||||
response = client.post(
|
||||
f"/api/adventures/{adv_id}/actions", json={"type": kind, "text": text}
|
||||
)
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
return response
|
||||
|
||||
|
||||
def _fact(predicate, value, fact_id):
|
||||
return {"type": "add_fact", "predicate": predicate, "value": value,
|
||||
"fact_id": fact_id}
|
||||
|
||||
|
||||
def upload(client, adv_id, name, body, classification, **fields):
|
||||
"""Imports a file the way the browser does: multipart, and no pathname."""
|
||||
data = {"classification": classification}
|
||||
data.update({k: str(v).lower() if isinstance(v, bool) else str(v)
|
||||
for k, v in fields.items()})
|
||||
response = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": (name, body.encode("utf-8"), "text/markdown")},
|
||||
data=data,
|
||||
)
|
||||
assert response.status_code == 201, response.text[:400]
|
||||
return response.json()["id"]
|
||||
|
||||
|
||||
def _checkpoint(client, adv_id, name, note=""):
|
||||
response = client.post(
|
||||
f"/api/adventures/{adv_id}/checkpoints", json={"name": name, "note": note}
|
||||
)
|
||||
assert response.status_code == 201, response.text[:400]
|
||||
return response.json()
|
||||
|
||||
|
||||
def _undo(client, adv_id, times=1):
|
||||
for _ in range(times):
|
||||
response = client.post(f"/api/adventures/{adv_id}/undo")
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
|
||||
|
||||
def settle_derived(adv_id):
|
||||
"""Runs the background memory and summary pass to completion.
|
||||
|
||||
The turn endpoint fires this as a fire-and-forget task, which a test client
|
||||
does not wait for. Calling it directly is the same code on the same rows —
|
||||
what is skipped is the scheduling, not the work — and it is what
|
||||
`test_context_realistic.py` does for the same reason.
|
||||
"""
|
||||
asyncio.run(memorybank.run_post_turn(adv_id))
|
||||
|
||||
|
||||
# --------------------------------------------------------------------- build
|
||||
|
||||
def build(client, adv_id) -> dict:
|
||||
"""Plays the fixture campaign onto `adv_id`, and returns what it built.
|
||||
|
||||
The returned dictionary is the assertion source for every round-trip test:
|
||||
it names the properties that must survive, measured from the campaign as it
|
||||
stands here rather than restated as constants, so a test compares the copy
|
||||
against the original instead of against a guess about the original.
|
||||
"""
|
||||
# Story memory and the rolling summary on, because a campaign that
|
||||
# generated neither would let an exporter omit both and still pass. The
|
||||
# abandoned line below gets long enough to earn its own, which is what E03
|
||||
# is about after a round trip.
|
||||
switched_on = client.patch(
|
||||
f"/api/adventures/{adv_id}",
|
||||
json={"auto_summarize": True, "memory_bank_enabled": True},
|
||||
)
|
||||
assert switched_on.status_code == 200, switched_on.text[:400]
|
||||
|
||||
sources = {
|
||||
"canon": upload(client, adv_id, "canon.md", CANON_MD, "canon",
|
||||
always_include=True),
|
||||
"reference": upload(client, adv_id, "reference.md", REFERENCE_MD,
|
||||
"reference"),
|
||||
"inspiration": upload(client, adv_id, "inspiration.md", INSPIRATION_MD,
|
||||
"inspiration"),
|
||||
"secret": upload(client, adv_id, "secret.md", SECRET_MD, "canon",
|
||||
visibility="hidden"),
|
||||
"disabled": upload(client, adv_id, "draft.md", DISABLED_MD, "reference"),
|
||||
}
|
||||
disable = client.patch(
|
||||
f"/api/adventures/{adv_id}/knowledge/{sources['disabled']}",
|
||||
json={"enabled": False},
|
||||
)
|
||||
assert disable.status_code == 200, disable.text[:400]
|
||||
|
||||
# ---- the first line of story -----------------------------------------
|
||||
# Turn 1 asks about the abbey, so the canon source is retrieved and the
|
||||
# stored prompt for this turn holds an imported passage. That turn is the
|
||||
# one the provenance tests read back after the round trip.
|
||||
_play(client, adv_id, "ask Mara about the abbey",
|
||||
"Mara sets down the cloth. The abbey, she says, is sealed.",
|
||||
[_fact("tally", 10, "tally-10")])
|
||||
_play(client, adv_id, "walk up to the abbey",
|
||||
"The path climbs out of the town and the rain follows.",
|
||||
[_fact("tally", 20, "tally-20")])
|
||||
|
||||
# A retry, so one coordinate holds two takes and the earlier one is
|
||||
# retained but not selected.
|
||||
ScriptedProvider.replies = [
|
||||
"The door is oak, and the seal on it is unbroken.\n"
|
||||
+ state_block([_fact("tally", 30, "tally-30")])
|
||||
]
|
||||
_play(client, adv_id, "try the crypt door",
|
||||
"The door will not move.", [_fact("tally", 30, "tally-30")])
|
||||
retry = client.post(f"/api/adventures/{adv_id}/retry")
|
||||
assert retry.status_code == 200, retry.text[:400]
|
||||
|
||||
s1 = _checkpoint(client, adv_id, "At the crypt door",
|
||||
"Before anything is decided.")
|
||||
|
||||
# The future the reader is about to leave behind. It is played out far
|
||||
# enough to earn derived data of its own — `memorybank.MEMORY_INTERVAL` is
|
||||
# six actions — because a summary and a memory belonging to an abandoned
|
||||
# line are what E03 forbids reaching an active prompt, and a round trip is
|
||||
# a new way to leak one.
|
||||
_play(client, adv_id, "force the door",
|
||||
"The seal gives, and the stair below is dark.",
|
||||
[_fact("tally", 40, "tally-40")])
|
||||
_play(client, adv_id, "go down",
|
||||
"The crypt is dry, and the air has not moved in years.",
|
||||
[_fact("tally", 50, "tally-50")])
|
||||
_play(client, adv_id, "read the names on the slabs",
|
||||
"Sixty years of Westhaven dead, and one slab with no name at all.",
|
||||
[_fact("tally", 60, "tally-60")])
|
||||
_play(client, adv_id, "touch the nameless slab",
|
||||
"The stone is warm, which stone in a crypt is not.",
|
||||
[_fact("tally", 70, "tally-70")])
|
||||
|
||||
# Derived data for the line that is about to be abandoned, written while
|
||||
# the head is still on it. This is the summary and the memory that must
|
||||
# come back after a round trip and must still be ineligible there.
|
||||
settle_derived(adv_id)
|
||||
tip_state = client.get(f"/api/adventures/{adv_id}/state").json()
|
||||
|
||||
# ---- step back, and go somewhere else ---------------------------------
|
||||
_undo(client, adv_id, 4)
|
||||
_play(client, adv_id, "turn back and return to the tavern",
|
||||
"The rain has not let up, and the Lantern's windows are lit.",
|
||||
[_fact("tally", 41, "tally-41")])
|
||||
s2 = _checkpoint(client, adv_id, "Back at the Lantern", "The other way.")
|
||||
_play(client, adv_id, "ask Mara what she is not saying",
|
||||
"She looks at the fire for a while before she answers.",
|
||||
[_fact("tally", 51, "tally-51")])
|
||||
_play(client, adv_id, "wait",
|
||||
"The rain fills the silence, and then she starts talking.",
|
||||
[_fact("tally", 61, "tally-61")])
|
||||
|
||||
# A manual correction: an accepted state change with no narration behind
|
||||
# it, which is the one kind of state event a replay could never recreate.
|
||||
correction = client.post(
|
||||
f"/api/adventures/{adv_id}/state/corrections",
|
||||
json={
|
||||
"events": [{
|
||||
"type": "add_fact",
|
||||
"predicate": "keeper_of_the_lantern",
|
||||
"value": "Mara",
|
||||
"fact_id": "keeper",
|
||||
}],
|
||||
"note": "Established in play before the state system saw it.",
|
||||
},
|
||||
)
|
||||
assert correction.status_code == 201, correction.text[:400]
|
||||
|
||||
# Derived data for the line the reader stayed on, so the copy has both an
|
||||
# eligible and an ineligible summary to tell apart. The generated one landed
|
||||
# on the abandoned line, which is the E03 case; this one is typed at the
|
||||
# current head, so it is the eligible case beside it. A round trip has to
|
||||
# keep them on opposite sides of that line.
|
||||
settle_derived(adv_id)
|
||||
|
||||
# One more Undo, so the head finishes behind the retained tip of its own
|
||||
# branch as well as behind the abandoned line's.
|
||||
_undo(client, adv_id, 1)
|
||||
|
||||
# Typed at the final head, so it is the eligible summary and the generated
|
||||
# one on the abandoned line is not. A round trip has to keep them on
|
||||
# opposite sides of that line.
|
||||
typed = client.patch(
|
||||
f"/api/adventures/{adv_id}",
|
||||
json={"story_summary": "Aldric went back to the Lantern instead."},
|
||||
)
|
||||
assert typed.status_code == 200, typed.text[:400]
|
||||
|
||||
return snapshot_of(client, adv_id, sources=sources, s1=s1, s2=s2,
|
||||
tip_state=tip_state)
|
||||
|
||||
|
||||
def snapshot_in(action: dict) -> dict | None:
|
||||
"""The stored prompt in one bundle entry, decoded.
|
||||
|
||||
The export compresses it (`bundle._packed`), so a test that reached for a
|
||||
plain dict would conclude the evidence was missing when it is merely
|
||||
encoded. Both keys are read, plain first, exactly as the importer does.
|
||||
"""
|
||||
from app import bundle
|
||||
|
||||
plain = action.get("contextSnapshot")
|
||||
if isinstance(plain, dict):
|
||||
return plain
|
||||
return bundle._unpacked(action.get("contextSnapshotZ"))
|
||||
|
||||
|
||||
def with_snapshot(action: dict, snapshot: dict | None) -> dict:
|
||||
"""A bundle entry carrying `snapshot`, written in the plain form.
|
||||
|
||||
Tests that break a snapshot on purpose write the readable key, because the
|
||||
importer prefers it and because a test that had to compress its own fixture
|
||||
would be testing the encoding rather than the thing it edited.
|
||||
"""
|
||||
edited = {k: v for k, v in action.items() if k != "contextSnapshotZ"}
|
||||
if snapshot is None:
|
||||
edited.pop("contextSnapshot", None)
|
||||
else:
|
||||
edited["contextSnapshot"] = snapshot
|
||||
return edited
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- reading
|
||||
|
||||
def snapshot_of(client, adv_id, *, sources=None, s1=None, s2=None,
|
||||
tip_state=None) -> dict:
|
||||
"""Everything about a campaign that a round trip has to reproduce.
|
||||
|
||||
Read through the API, so the comparison is between what a reader can see in
|
||||
the source campaign and what a reader can see in the copy. Two campaigns
|
||||
that agree here agree on everything the product promises about a restored
|
||||
campaign; nothing below is a database id, because ids are expected to
|
||||
differ.
|
||||
"""
|
||||
head = client.get(f"/api/adventures/{adv_id}").json()
|
||||
branches = client.get(f"/api/adventures/{adv_id}/branches").json()
|
||||
checkpoints = client.get(f"/api/adventures/{adv_id}/checkpoints").json()
|
||||
knowledge = client.get(f"/api/adventures/{adv_id}/knowledge").json()
|
||||
state = client.get(f"/api/adventures/{adv_id}/state").json()
|
||||
events = client.get(f"/api/adventures/{adv_id}/state/events?limit=500").json()
|
||||
memories = client.get(f"/api/adventures/{adv_id}/memories").json()
|
||||
derived = client.get(f"/api/adventures/{adv_id}/derived").json()
|
||||
return {
|
||||
"id": adv_id,
|
||||
"title": head["title"],
|
||||
"canon_rules": head.get("canon_rules") or [],
|
||||
"can_undo": head.get("can_undo"),
|
||||
"can_redo": head.get("can_redo"),
|
||||
"transcript": [(a["type"], a["text"]) for a in head["actions"]],
|
||||
# Every branch's own story, which is the whole retained tree as text.
|
||||
"branch_count": len(branches),
|
||||
"checkpoints": sorted(
|
||||
(c["name"], c["note"]) for c in checkpoints
|
||||
),
|
||||
"knowledge": sorted(
|
||||
(k["title"], k["classification"], k["enabled"], k["visibility"],
|
||||
k["always_include"], k["content_hash"])
|
||||
for k in knowledge
|
||||
),
|
||||
"state": _comparable_state(state),
|
||||
"state_events": sorted(
|
||||
(e["event_type"], e["source"], _payload_key(e["payload"]))
|
||||
for e in events
|
||||
),
|
||||
"memories": sorted(m["text"] for m in memories),
|
||||
"summaries": sorted(
|
||||
(s["preview"], s["trigger"], s["eligible"])
|
||||
for s in derived.get("summaries", [])
|
||||
),
|
||||
# Carried through from `build`, for the tests that need the original
|
||||
# ids or the state at a position the head has since left.
|
||||
"sources": sources,
|
||||
"s1": s1,
|
||||
"s2": s2,
|
||||
"tip_state": _comparable_state(tip_state) if tip_state else None,
|
||||
}
|
||||
|
||||
|
||||
def _comparable_state(state: dict) -> dict:
|
||||
"""The authoritative state, with only what a reader is shown.
|
||||
|
||||
Groups arrive from the API as display sections, which is the right shape to
|
||||
compare: two campaigns whose State panels read identically hold the same
|
||||
state, whatever ids sit underneath.
|
||||
"""
|
||||
groups = state.get("groups") if isinstance(state, dict) else None
|
||||
if not isinstance(groups, list):
|
||||
return {}
|
||||
return {
|
||||
str(group.get("title")): sorted(
|
||||
", ".join(f"{k}={group_row[k]}" for k in sorted(group_row))
|
||||
for group_row in (group.get("rows") or [])
|
||||
if isinstance(group_row, dict)
|
||||
)
|
||||
for group in groups
|
||||
}
|
||||
|
||||
|
||||
def _payload_key(payload) -> str:
|
||||
"""A stable identity for an event payload, for set comparison."""
|
||||
if not isinstance(payload, dict):
|
||||
return str(payload)
|
||||
for key in ("fact_id", "entity_id", "thread_id", "id", "predicate"):
|
||||
if payload.get(key):
|
||||
return f"{key}={payload[key]}"
|
||||
return ",".join(f"{k}={payload[k]}" for k in sorted(payload))
|
||||
@@ -568,10 +568,29 @@ def test_the_action_cap_counts_the_rows_a_v1_file_expands_into(client, monkeypat
|
||||
assert _adventure_count() == before, "and nothing was written"
|
||||
|
||||
|
||||
def test_an_unknown_format_is_refused(client):
|
||||
r = _import(client, {"format": "ai-dnd-adventure-v3", "title": "From the future"})
|
||||
def test_a_format_from_a_later_build_is_refused(client):
|
||||
"""A version this build has never heard of is refused, not guessed at.
|
||||
|
||||
The placeholder version here has to stay ahead of `bundle.FORMAT`. It was
|
||||
`v3` until M9 made v3 real, at which point this test started importing a
|
||||
bundle it meant to reject — the failure mode a hard-coded "next version"
|
||||
always eventually has, and the reason the message is asserted against
|
||||
`bundle.FORMAT` rather than against a literal.
|
||||
"""
|
||||
r = _import(client, {"format": "ai-dnd-adventure-v99", "title": "From the future"})
|
||||
assert r.status_code == 400, r.text
|
||||
assert bundle.FORMAT in r.json()["detail"]
|
||||
detail = r.json()["detail"]
|
||||
assert bundle.FORMAT in detail
|
||||
assert "ai-dnd-adventure-v99" in detail
|
||||
|
||||
|
||||
def test_something_that_is_not_an_export_at_all_is_refused(client):
|
||||
r = _import(client, {"title": "A file of some other kind"})
|
||||
assert r.status_code == 400, r.text
|
||||
# Every version it can read is named, so the reader can tell whether the
|
||||
# file they have is one of them.
|
||||
for readable in bundle.READABLE:
|
||||
assert readable in r.json()["detail"]
|
||||
|
||||
|
||||
# ------------------------------------------------------- the persona (Phase 18)
|
||||
|
||||
@@ -1660,7 +1660,22 @@ def test_an_edited_content_hash_is_recomputed_and_reported(client):
|
||||
|
||||
|
||||
def test_historical_prompt_evidence_survives_an_export_round_trip(client):
|
||||
"""§33: the round trip does not turn provenance into dangling ids."""
|
||||
"""§33: the round trip does not turn provenance into dangling ids.
|
||||
|
||||
Written in M7 and rewritten in M9, and the rewrite is the point of it.
|
||||
|
||||
In M7 the bundle carried no context snapshots at all, so this test pinned
|
||||
the *absence*: there were no ids to dangle because there was no evidence,
|
||||
and the imported campaign's turns simply had no snapshot. That was recorded
|
||||
at the time as a limit owned by M9 rather than as a property worth keeping —
|
||||
`V1-ACCEPTANCE-TESTS.md` I05 said so in as many words, and the M8 report
|
||||
made it handoff question B.
|
||||
|
||||
M9 answered it: the evidence travels. So the assertion inverts, and what it
|
||||
now pins is the thing M7 was worried about and could not check — that the
|
||||
provenance arriving on the other side names *this* campaign's sources rather
|
||||
than the ids it had on the machine that wrote the file.
|
||||
"""
|
||||
ids = import_fixture(client)
|
||||
play(client, "Aldric asks about the Old Abbey and its broken-circle symbol.")
|
||||
actions = client.get(f"/api/adventures/{client.adv_id}/actions").json()["actions"]
|
||||
@@ -1673,17 +1688,33 @@ def test_historical_prompt_evidence_survives_an_export_round_trip(client):
|
||||
bundle = client.get(f"/api/adventures/{client.adv_id}/export").json()
|
||||
restored = client.post("/api/adventures/import", json=bundle).json()
|
||||
|
||||
# The bundle carries no context snapshots at all — it never has, by the rule
|
||||
# at the top of `bundle.py` — so there are no ids to dangle. The imported
|
||||
# campaign's turns simply have no snapshot, which is what a pre-M7 bundle
|
||||
# already did for every other component of the inspector.
|
||||
new_actions = client.get(
|
||||
f"/api/adventures/{restored['id']}/actions"
|
||||
).json()["actions"]
|
||||
new_ai = next(a for a in reversed(new_actions) if a["type"] == "ai")
|
||||
assert client.get(
|
||||
moved = client.get(
|
||||
f"/api/adventures/{restored['id']}/actions/{new_ai['id']}/context"
|
||||
).status_code == 404
|
||||
)
|
||||
assert moved.status_code == 200, moved.text[:300]
|
||||
moved = moved.json()
|
||||
|
||||
# The evidence itself is identical: the same passages, the same text, the
|
||||
# same prompt the turn was actually assembled from.
|
||||
assert [(r["title"], r["text"]) for r in moved["knowledge"]["used"]] == \
|
||||
[(r["title"], r["text"]) for r in before["knowledge"]["used"]]
|
||||
assert moved["prompt"] == before["prompt"]
|
||||
|
||||
# And the one pointer that is not evidence has been translated, so the
|
||||
# inspector's "open this source" reaches the restored library rather than
|
||||
# whatever holds that id here.
|
||||
theirs = {
|
||||
source["id"] for source in
|
||||
client.get(f"/api/adventures/{restored['id']}/knowledge").json()
|
||||
}
|
||||
named = {r["source_id"] for r in moved["knowledge"]["used"]
|
||||
if r["source_id"] is not None}
|
||||
assert named and named <= theirs
|
||||
|
||||
# And the original campaign's evidence is untouched by having been exported.
|
||||
after = client.get(
|
||||
f"/api/adventures/{client.adv_id}/actions/{ai_action['id']}/context"
|
||||
|
||||
@@ -0,0 +1,342 @@
|
||||
"""M10 §6 and §18: the media layer cannot write the story.
|
||||
|
||||
The architectural claim is one sentence — *media is derived presentation, story
|
||||
state is authoritative, and there is no reverse path* — and this file is the
|
||||
part of it that is checked by running things rather than by reading imports.
|
||||
|
||||
Every test here follows the same shape, which is the shape that makes it
|
||||
evidence rather than assertion:
|
||||
|
||||
record the authoritative document, byte for byte
|
||||
do the media-layer thing
|
||||
record it again
|
||||
require them to be identical
|
||||
|
||||
That catches a write nobody intended as well as one somebody did, and it does
|
||||
not depend on knowing *how* a violation would have happened.
|
||||
|
||||
`test_m10_media_hooks.py` covers what the boundary carries; this covers what it
|
||||
must never push back through.
|
||||
|
||||
python -m pytest tests/test_m10_authority.py -v
|
||||
"""
|
||||
|
||||
import copy
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.media import packet as scene_packet
|
||||
from app.media import profiles as visual_profiles
|
||||
from app.media import providers
|
||||
from app.routers import adventures
|
||||
|
||||
import m10_fixture
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
|
||||
class StubDerived:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25] for _ in texts]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m10auth@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(user_id=user.id, title="Authority")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text="It begins.",
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def office(client):
|
||||
return m10_fixture.build(client, client.adv_id)
|
||||
|
||||
|
||||
def authoritative(adv_id) -> dict:
|
||||
"""Everything the story counts as true, read straight from the database."""
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
return {
|
||||
"state": copy.deepcopy(adventure.narrative_state),
|
||||
"head_branch": adventure.head_branch_id,
|
||||
"head_depth": adventure.head_depth,
|
||||
"events": db.query(models.StateEvent).filter(
|
||||
models.StateEvent.adventure_id == adv_id).count(),
|
||||
"proposals": db.query(models.StateProposal).filter(
|
||||
models.StateProposal.adventure_id == adv_id).count(),
|
||||
"actions": db.query(models.Action).filter(
|
||||
models.Action.adventure_id == adv_id).count(),
|
||||
}
|
||||
|
||||
|
||||
# ------------------------------------------------------ writes that must not
|
||||
|
||||
def test_writing_a_visual_profile_changes_no_story_state(client, office):
|
||||
before = authoritative(client.adv_id)
|
||||
response = client.put(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/bill",
|
||||
json={"descriptors": {"build": "heavyset", "clothing": "navy suit"},
|
||||
"features": ["signet ring"], "style_notes": "photographic"},
|
||||
)
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
assert authoritative(client.adv_id) == before
|
||||
|
||||
|
||||
def test_updating_a_visual_profile_creates_no_state_fact(client, office):
|
||||
"""§6's example, made concrete.
|
||||
|
||||
A profile saying Alice wears a blue coat must not make it true that Alice
|
||||
owns or wears a blue coat. Checked by looking for the words in the
|
||||
authoritative document afterwards, not only by comparing counts.
|
||||
"""
|
||||
before = authoritative(client.adv_id)
|
||||
client.put(f"/api/adventures/{client.adv_id}/visual-profiles/alice",
|
||||
json={"descriptors": {"clothing": "blue coat"}})
|
||||
after = authoritative(client.adv_id)
|
||||
assert after == before
|
||||
assert "blue coat" not in repr(after["state"])
|
||||
|
||||
document = client.get(
|
||||
f"/api/adventures/{client.adv_id}/state").json()["document"]
|
||||
assert not any("blue coat" in repr(f) for f in document["facts"])
|
||||
assert "blue coat" not in repr(document["entities"]["alice"])
|
||||
|
||||
|
||||
def test_deleting_a_visual_profile_changes_no_story_state(client, office):
|
||||
before = authoritative(client.adv_id)
|
||||
assert client.delete(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alice"
|
||||
).status_code == 204
|
||||
assert authoritative(client.adv_id) == before
|
||||
|
||||
|
||||
def test_building_a_scene_packet_changes_nothing(client, office):
|
||||
"""A packet is a read. Built repeatedly, it must still be a read."""
|
||||
before = authoritative(client.adv_id)
|
||||
for _ in range(5):
|
||||
assert client.get(
|
||||
f"/api/adventures/{client.adv_id}/scene-packet"
|
||||
).status_code == 200
|
||||
assert authoritative(client.adv_id) == before
|
||||
|
||||
|
||||
def test_a_scene_packet_does_not_move_the_head(client, office):
|
||||
before = authoritative(client.adv_id)
|
||||
client.get(f"/api/adventures/{client.adv_id}/scene-packet?start=0&end=4")
|
||||
after = authoritative(client.adv_id)
|
||||
assert after["head_branch"] == before["head_branch"]
|
||||
assert after["head_depth"] == before["head_depth"]
|
||||
|
||||
|
||||
def test_a_dummy_media_result_cannot_reach_the_story(client, office):
|
||||
"""§18: adding a depiction, even a wrong one, changes nothing.
|
||||
|
||||
The result claims Alice is wearing a red coat and standing in a corridor.
|
||||
None of that is true in the campaign, and after registering, generating and
|
||||
holding the result, none of it has become true.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
before = authoritative(client.adv_id)
|
||||
packet = client.get(
|
||||
f"/api/adventures/{client.adv_id}/scene-packet").json()
|
||||
|
||||
class WrongProvider:
|
||||
def capabilities(self):
|
||||
return providers.ProviderCapabilities(
|
||||
provider_id="wrong", kinds=(providers.IMAGE,))
|
||||
|
||||
async def generate(self, request):
|
||||
return providers.MediaResult(
|
||||
kind=providers.IMAGE, media_type="image/png",
|
||||
data=b"\x89PNG\r\n\x1a\n",
|
||||
provenance={"scene_id": request.scene["scene_id"]},
|
||||
details={"depicts": "Alice in a red coat in a corridor"},
|
||||
)
|
||||
|
||||
providers.register("wrong", WrongProvider())
|
||||
try:
|
||||
result = asyncio.run(WrongProvider().generate(
|
||||
providers.MediaRequest(kind=providers.IMAGE, scene=packet)))
|
||||
assert "red coat" in result.details["depicts"]
|
||||
finally:
|
||||
providers.unregister("wrong")
|
||||
|
||||
after = authoritative(client.adv_id)
|
||||
assert after == before
|
||||
assert "red coat" not in repr(after["state"])
|
||||
assert "corridor" not in repr(after["state"])
|
||||
|
||||
|
||||
def test_a_provider_failure_cannot_advance_the_head(client, office):
|
||||
"""§18: a media failure is not a story event."""
|
||||
import asyncio
|
||||
|
||||
before = authoritative(client.adv_id)
|
||||
|
||||
class FailingProvider:
|
||||
def capabilities(self):
|
||||
return providers.ProviderCapabilities(
|
||||
provider_id="failing", kinds=(providers.IMAGE,))
|
||||
|
||||
async def generate(self, request):
|
||||
raise providers.MediaProviderError("the local generator is not running")
|
||||
|
||||
providers.register("failing", FailingProvider())
|
||||
try:
|
||||
with pytest.raises(providers.MediaProviderError):
|
||||
asyncio.run(FailingProvider().generate(providers.MediaRequest(
|
||||
kind=providers.IMAGE,
|
||||
scene=client.get(
|
||||
f"/api/adventures/{client.adv_id}/scene-packet").json())))
|
||||
finally:
|
||||
providers.unregister("failing")
|
||||
|
||||
assert authoritative(client.adv_id) == before
|
||||
|
||||
|
||||
def test_a_scene_derivation_failure_does_not_corrupt_an_accepted_turn(client, office):
|
||||
"""§18: if building a packet raised, the story would be untouched.
|
||||
|
||||
The failure is induced in the packet builder itself, which is the only place
|
||||
derivation happens, and the accepted turn either side is compared whole.
|
||||
"""
|
||||
before = authoritative(client.adv_id)
|
||||
original = scene_packet.build
|
||||
|
||||
def explode(*args, **kwargs):
|
||||
raise RuntimeError("scene derivation failed")
|
||||
|
||||
scene_packet.build = explode
|
||||
try:
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/scene-packet")
|
||||
assert response.status_code >= 500
|
||||
except RuntimeError:
|
||||
pass # the TestClient re-raises; either way the story must be intact
|
||||
finally:
|
||||
scene_packet.build = original
|
||||
|
||||
assert authoritative(client.adv_id) == before
|
||||
# And the campaign still plays.
|
||||
m10_fixture.play(client, client.adv_id, "carry on", [])
|
||||
assert authoritative(client.adv_id)["actions"] == before["actions"] + 2
|
||||
|
||||
|
||||
# ------------------------------------------------- rebuilding derived data
|
||||
|
||||
def test_deleting_every_visual_profile_leaves_the_campaign_intact(client, office):
|
||||
"""§18's last clause: derived data can go without taking the story with it.
|
||||
|
||||
Profiles are the only thing M10 persists, and they are recoverable only from
|
||||
a bundle or by being written again — so the promise here is narrower than
|
||||
M9's rebuildable indexes, and the test states the narrow thing: removing
|
||||
them costs the descriptions and nothing else.
|
||||
"""
|
||||
before = authoritative(client.adv_id)
|
||||
with SessionLocal() as db:
|
||||
db.query(models.VisualProfile).filter(
|
||||
models.VisualProfile.adventure_id == client.adv_id
|
||||
).delete(synchronize_session=False)
|
||||
db.commit()
|
||||
|
||||
assert authoritative(client.adv_id) == before
|
||||
assert client.get(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles").json()["profiles"] == []
|
||||
|
||||
# The packet still builds; it simply describes nobody's appearance.
|
||||
p = client.get(f"/api/adventures/{client.adv_id}/scene-packet").json()
|
||||
assert [c["name"] for c in p["characters"]] == ["Bill", "Alice", "Roger"]
|
||||
assert all(c["visual_profile"] is None for c in p["characters"])
|
||||
|
||||
|
||||
def test_the_story_survives_a_profile_naming_a_vanished_entity(client, office):
|
||||
"""A profile whose entity is gone is inert, not a corruption.
|
||||
|
||||
Reachable through an import: a bundle may carry a profile for an entity that
|
||||
only exists on a branch the campaign has left.
|
||||
"""
|
||||
with SessionLocal() as db:
|
||||
db.add(models.VisualProfile(
|
||||
adventure_id=client.adv_id, entity_key="nobody_at_all",
|
||||
descriptors={"hair": "green"}, features=[], style_notes=""))
|
||||
db.commit()
|
||||
before = authoritative(client.adv_id)
|
||||
p = client.get(f"/api/adventures/{client.adv_id}/scene-packet").json()
|
||||
assert "green" not in repr(p)
|
||||
assert authoritative(client.adv_id) == before
|
||||
m10_fixture.play(client, client.adv_id, "carry on", [])
|
||||
|
||||
|
||||
# ---------------------------------------------- the separation, structurally
|
||||
|
||||
def test_the_media_package_imports_nothing_that_writes_state(client):
|
||||
"""The guarantee behind every test above, checked as an import rule.
|
||||
|
||||
`narrative.apply` and `narrative.store` are the only modules that write the
|
||||
authoritative document, and `media/` reaching either of them would make the
|
||||
separation a convention rather than a fact. `narrative.model` and
|
||||
`narrative.store.current` are reads and are used.
|
||||
"""
|
||||
import pathlib
|
||||
|
||||
seam = pathlib.Path(__file__).resolve().parent.parent / "app" / "media"
|
||||
for path in seam.rglob("*.py"):
|
||||
body = path.read_text()
|
||||
assert "narrative.apply" not in body, path.name
|
||||
assert "from ..narrative import apply" not in body, path.name
|
||||
assert "set_current" not in body, path.name
|
||||
assert "head.move_to" not in body, path.name
|
||||
assert "tree.place_action" not in body, path.name
|
||||
|
||||
|
||||
def test_no_state_event_type_was_added_for_media(client):
|
||||
"""M10 adds no way for the media layer to speak in the story's vocabulary."""
|
||||
from app.narrative import events
|
||||
|
||||
assert not any(
|
||||
name.startswith("media") or "visual" in name or "asset" in name
|
||||
for name in events.ALLOWED
|
||||
)
|
||||
@@ -0,0 +1,481 @@
|
||||
"""M10 §14 and §15: the profiles travel, and an M9 database opens.
|
||||
|
||||
Two questions, and they are the ones a reader would ask if they knew what M10
|
||||
had done to their machine:
|
||||
|
||||
* **§14 — does a campaign still move?** A visual profile is part of the campaign
|
||||
the reader built, so it belongs in the bundle. It is also *new*, which is the
|
||||
risk: an exporter that carries it and an importer that drops it both pass a
|
||||
test that only checks the campaign still opens.
|
||||
* **§15 — does the database I already have still work?** M10 adds one table and
|
||||
nothing else. An existing campaign must survive opening under the new build
|
||||
untouched, opening must not care how many times it happens, the schema an M9
|
||||
file reaches must be the schema a fresh install has, and M9's backup must keep
|
||||
working on the result.
|
||||
|
||||
The upgrade needs **no migration**: `create_all` builds a new table and the
|
||||
indexes declared on its columns on every path. A `CREATE INDEX` migration was
|
||||
written here first and `test_a_fresh_database_arrives_at_the_same_place` is what
|
||||
found it wrong — it left an upgraded database holding an index a fresh install
|
||||
did not have. That test is the one to keep pointed at any future schema change.
|
||||
|
||||
The bundle format stays `ai-dnd-adventure-v3`. M9's own test for a version bump
|
||||
is whether omission creates ambiguity about what an older file *could* have
|
||||
recorded, and it does not: a campaign with no visual profiles is the ordinary
|
||||
case, so an absent key means "none" rather than "unknown". The tests below hold
|
||||
that decision to its consequence — an M9-written v3 file must still import, and
|
||||
the M10 exporter must still produce a file an M9 build would recognise.
|
||||
|
||||
python -m pytest tests/test_m10_bundle.py -v
|
||||
"""
|
||||
|
||||
import copy
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
from sqlalchemy import create_engine, text
|
||||
from sqlalchemy.orm import sessionmaker
|
||||
|
||||
from app import auth, backup, limits, migrations, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
import m10_fixture
|
||||
from fakes import ScriptedProvider
|
||||
from test_process_restart import Server, _free_port
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m10bundle@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(user_id=user.id, model="test-model",
|
||||
embedding_model=""))
|
||||
adventure = models.Adventure(user_id=user.id, title="Portable office")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="Bill badges in."))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def export(client, adv_id=None) -> dict:
|
||||
response = client.get(f"/api/adventures/{adv_id or client.adv_id}/export")
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
return response.json()
|
||||
|
||||
|
||||
def bring_back(client, payload) -> int:
|
||||
response = client.post("/api/adventures/import", json=payload)
|
||||
assert response.status_code == 201, response.text[:600]
|
||||
return response.json()["id"]
|
||||
|
||||
|
||||
def profiles_of(client, adv_id) -> dict:
|
||||
body = client.get(f"/api/adventures/{adv_id}/visual-profiles").json()
|
||||
return {p["entity_key"]: p for p in body["profiles"]}
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def moved(client):
|
||||
"""The office campaign, its bundle, and the copy the bundle produced."""
|
||||
m10_fixture.build(client, client.adv_id)
|
||||
payload = export(client)
|
||||
return {"bundle": payload, "copy_id": bring_back(client, payload)}
|
||||
|
||||
|
||||
# --------------------------------------------------------------- §14 the file
|
||||
|
||||
def test_the_format_version_is_unchanged(moved):
|
||||
"""The decision, recorded as a test so a later bump is deliberate."""
|
||||
assert moved["bundle"]["format"] == "ai-dnd-adventure-v3"
|
||||
|
||||
|
||||
def test_the_bundle_carries_the_profiles_that_exist(moved):
|
||||
exported = {p["entityKey"]: p for p in moved["bundle"]["visualProfiles"]}
|
||||
assert set(exported) == {"alice", "office"}
|
||||
assert exported["alice"]["descriptors"]["hair"] == "short black"
|
||||
assert exported["alice"]["features"] == ["tortoiseshell glasses"]
|
||||
assert exported["alice"]["styleNotes"] == "photographic, natural light"
|
||||
|
||||
|
||||
def test_an_unprofiled_character_exports_no_empty_profile(moved):
|
||||
"""Roger has no profile, and the file must say that by omission.
|
||||
|
||||
An exporter that wrote a blank row for every entity would lose the
|
||||
distinction a provider needs: "nobody decided what Roger looks like" is not
|
||||
"Roger looks like nothing".
|
||||
"""
|
||||
keys = [p["entityKey"] for p in moved["bundle"]["visualProfiles"]]
|
||||
assert "roger" not in keys and "bill" not in keys
|
||||
|
||||
|
||||
def test_the_copy_holds_the_same_profiles(client, moved):
|
||||
original = profiles_of(client, client.adv_id)
|
||||
copied = profiles_of(client, moved["copy_id"])
|
||||
assert set(copied) == set(original)
|
||||
for key in original:
|
||||
assert copied[key]["descriptors"] == original[key]["descriptors"]
|
||||
assert copied[key]["features"] == original[key]["features"]
|
||||
assert copied[key]["style_notes"] == original[key]["style_notes"]
|
||||
|
||||
|
||||
def test_the_copys_profiles_are_its_own_rows(client, moved):
|
||||
"""Editing the copy must not reach back into the original."""
|
||||
client.put(f"/api/adventures/{moved['copy_id']}/visual-profiles/alice",
|
||||
json={"descriptors": {"hair": "bleached"}})
|
||||
assert profiles_of(client, client.adv_id)["alice"][
|
||||
"descriptors"]["hair"] == "short black"
|
||||
|
||||
|
||||
def test_the_copys_scene_packet_is_populated_from_the_imported_profiles(
|
||||
client, moved):
|
||||
"""The point of carrying them: the copy can be depicted without redoing work."""
|
||||
packet = client.get(
|
||||
f"/api/adventures/{moved['copy_id']}/scene-packet").json()
|
||||
by_name = {c["name"]: c for c in packet["characters"]}
|
||||
assert by_name["Alice"]["visual_profile"]["descriptors"]["build"] == "tall"
|
||||
assert by_name["Roger"]["visual_profile"] is None
|
||||
assert packet["location"]["visual_profile"]["descriptors"][
|
||||
"lighting"] == "flat fluorescent"
|
||||
|
||||
|
||||
def test_an_m9_era_file_still_imports_and_simply_has_no_profiles(client, moved):
|
||||
"""A v3 file written before M10 existed: the key is absent, not empty."""
|
||||
older = copy.deepcopy(moved["bundle"])
|
||||
del older["visualProfiles"]
|
||||
copy_id = bring_back(client, older)
|
||||
assert profiles_of(client, copy_id) == {}
|
||||
# And the campaign itself arrived intact.
|
||||
assert client.get(f"/api/adventures/{copy_id}/scene-packet").json()[
|
||||
"characters"]
|
||||
|
||||
|
||||
def test_a_malformed_profile_is_dropped_rather_than_refusing_the_campaign(
|
||||
client, moved):
|
||||
"""§14's proportionality rule, in the one place M10 could get it wrong.
|
||||
|
||||
A story that will not import because a description of somebody's coat is
|
||||
malformed would be the wrong trade. The campaign arrives; the bad profile
|
||||
does not; the good one does.
|
||||
"""
|
||||
damaged = copy.deepcopy(moved["bundle"])
|
||||
damaged["visualProfiles"].append(
|
||||
{"entity_key": "", "descriptors": "not an object"})
|
||||
damaged["visualProfiles"].append({"descriptors": {"a": "b"}})
|
||||
copy_id = bring_back(client, damaged)
|
||||
assert set(profiles_of(client, copy_id)) == {"alice", "office"}
|
||||
|
||||
|
||||
def test_a_profile_survives_a_second_round_trip_unchanged(client, moved):
|
||||
"""Export, import, export again: the file is a fixed point."""
|
||||
again = export(client, moved["copy_id"])
|
||||
first = sorted(moved["bundle"]["visualProfiles"], key=lambda p: p["entityKey"])
|
||||
second = sorted(again["visualProfiles"], key=lambda p: p["entityKey"])
|
||||
assert [p["entityKey"] for p in first] == [p["entityKey"] for p in second]
|
||||
for a, b in zip(first, second):
|
||||
assert a["descriptors"] == b["descriptors"]
|
||||
assert a["features"] == b["features"]
|
||||
assert a["styleNotes"] == b["styleNotes"]
|
||||
|
||||
|
||||
def test_a_neighbouring_campaigns_profiles_do_not_travel(client, moved):
|
||||
"""Scoping: the exporter must filter by campaign, not by table."""
|
||||
with SessionLocal() as db:
|
||||
neighbour = models.Adventure(user_id=None, title="Someone else's")
|
||||
db.add(neighbour)
|
||||
db.flush()
|
||||
db.add(models.VisualProfile(
|
||||
adventure_id=neighbour.id, entity_key="intruder",
|
||||
descriptors={"hair": "should not travel"}, features=[],
|
||||
style_notes=""))
|
||||
db.commit()
|
||||
keys = [p["entityKey"] for p in export(client)["visualProfiles"]]
|
||||
assert "intruder" not in keys
|
||||
|
||||
|
||||
def test_the_planner_checks_the_profiles_before_a_row_is_written(moved):
|
||||
"""M9's atomicity rule: everything is checked before anything is written.
|
||||
|
||||
`bundle.plan` is that checkpoint — it has no side effects and is what the
|
||||
importer runs first — so a profile that would fail must fail there rather
|
||||
than halfway through writing a campaign. There is no HTTP preview endpoint;
|
||||
the planner is called directly for the same reason the importer calls it.
|
||||
"""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
planned = bundle_module.plan(moved["bundle"], "ai-dnd-adventure-v3")
|
||||
assert {p["entity_key"] for p in planned["visualProfiles"]} == {
|
||||
"alice", "office"}
|
||||
|
||||
|
||||
# ---------------------------------------------------------- §15 the migration
|
||||
|
||||
@pytest.fixture()
|
||||
def m9_database():
|
||||
"""A database as an M9 build left it, with a campaign already in it.
|
||||
|
||||
M10's only schema change is the `visual_profiles` table, so an M9-era file
|
||||
is exactly this: the current schema without that table, stamped at 92 — the
|
||||
version M9 ended on and, since M10 adds no migration, the version it still
|
||||
ends on. The campaign rows are written before the upgrade, because the claim
|
||||
under test is that they are still there afterwards.
|
||||
"""
|
||||
directory = tempfile.mkdtemp(prefix="m10-migrate-")
|
||||
path = Path(directory) / "campaign.db"
|
||||
older = create_engine(f"sqlite:///{path}")
|
||||
Base.metadata.create_all(bind=older)
|
||||
# Written through the ORM, so the campaign in the file is shaped the way the
|
||||
# application writes one rather than the way a test guessed at.
|
||||
with sessionmaker(bind=older)() as db:
|
||||
adventure = models.Adventure(title="An M9 campaign")
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="The story opened before M10."))
|
||||
db.commit()
|
||||
adv_id = adventure.id
|
||||
with older.begin() as conn:
|
||||
conn.execute(text("DROP TABLE visual_profiles"))
|
||||
conn.execute(text("PRAGMA user_version = 92"))
|
||||
older.dispose()
|
||||
yield path, create_engine(f"sqlite:///{path}"), adv_id
|
||||
|
||||
|
||||
def _indexes(engine_) -> set:
|
||||
with engine_.begin() as conn:
|
||||
return {row[0] for row in conn.execute(text(
|
||||
"SELECT name FROM sqlite_master WHERE type = 'index'"))}
|
||||
|
||||
|
||||
def _version(engine_) -> int:
|
||||
with engine_.begin() as conn:
|
||||
return conn.execute(text("PRAGMA user_version")).scalar()
|
||||
|
||||
|
||||
def test_an_m9_database_gains_the_new_table_when_it_is_opened(m9_database):
|
||||
path, older, adv_id = m9_database
|
||||
assert _version(older) == 92
|
||||
migrations.bootstrap(older)
|
||||
assert _version(older) == migrations.LATEST_VERSION == 92
|
||||
with older.begin() as conn:
|
||||
assert conn.execute(text("SELECT COUNT(*) FROM visual_profiles")).scalar() == 0
|
||||
assert "ix_visual_profiles_adventure_id" in _indexes(older)
|
||||
|
||||
|
||||
def test_the_campaign_that_was_already_there_is_untouched(m9_database):
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
with older.begin() as conn:
|
||||
assert conn.execute(text("SELECT title FROM adventures")).scalar() == (
|
||||
"An M9 campaign")
|
||||
assert conn.execute(text("SELECT text FROM actions")).scalar() == (
|
||||
"The story opened before M10.")
|
||||
assert conn.execute(text("PRAGMA foreign_key_check")).fetchall() == []
|
||||
|
||||
|
||||
def test_opening_the_database_repeatedly_is_a_no_op(m9_database):
|
||||
"""Three starts in a row. Nothing accumulates and nothing errors.
|
||||
|
||||
This is the idempotence §15 asks about. It is stated as "open it again"
|
||||
rather than "run the migration again" because opening is what the
|
||||
application does, and M10 has no migration of its own to rerun.
|
||||
"""
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
after_first = _indexes(older)
|
||||
for _ in range(2):
|
||||
migrations.bootstrap(older)
|
||||
assert _version(older) == 92
|
||||
assert _indexes(older) == after_first
|
||||
with older.begin() as conn:
|
||||
assert conn.execute(text("SELECT COUNT(*) FROM adventures")).scalar() == 1
|
||||
|
||||
|
||||
def test_a_fresh_database_arrives_at_the_same_place(m9_database):
|
||||
"""An upgraded M9 file and a new install must not differ.
|
||||
|
||||
Two schemas that disagree is the failure this catches, and it is the one a
|
||||
version stamp alone would hide.
|
||||
"""
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
fresh_path = path.with_name("fresh.db")
|
||||
fresh = create_engine(f"sqlite:///{fresh_path}")
|
||||
migrations.bootstrap(fresh)
|
||||
assert _version(fresh) == _version(older)
|
||||
|
||||
def shape(e):
|
||||
with e.begin() as conn:
|
||||
return conn.execute(text(
|
||||
"SELECT sql FROM sqlite_master WHERE name = 'visual_profiles'"
|
||||
)).scalar()
|
||||
|
||||
assert shape(fresh) == shape(older)
|
||||
# Including the indexes. This comparison is what caught the redundant
|
||||
# `CREATE INDEX` migration M10 first shipped: the upgraded file had an index
|
||||
# the fresh one did not, which is a difference no test of either database on
|
||||
# its own would have shown.
|
||||
assert _indexes(fresh) == _indexes(older)
|
||||
fresh.dispose()
|
||||
|
||||
|
||||
def test_a_backup_of_the_upgraded_database_still_works(m9_database):
|
||||
"""M9's backup keeps its guarantees on a file M10 added a table to."""
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
older.dispose()
|
||||
|
||||
result = backup.create(path)
|
||||
try:
|
||||
assert result.integrity == "ok"
|
||||
assert result.pages > 0
|
||||
with sqlite3.connect(f"file:{result.path}?mode=ro", uri=True) as copy_db:
|
||||
assert copy_db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
|
||||
assert copy_db.execute("PRAGMA foreign_key_check").fetchall() == []
|
||||
# It opens independently: the new table is in it, and so is the
|
||||
# campaign that predates the migration.
|
||||
assert copy_db.execute(
|
||||
"SELECT COUNT(*) FROM visual_profiles").fetchone()[0] == 0
|
||||
assert copy_db.execute(
|
||||
"SELECT title FROM adventures").fetchone()[0] == "An M9 campaign"
|
||||
assert copy_db.execute("PRAGMA user_version").fetchone()[0] == 92
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_a_backup_carries_the_profiles_written_after_the_upgrade(m9_database):
|
||||
path, older, adv_id = m9_database
|
||||
migrations.bootstrap(older)
|
||||
with older.begin() as conn:
|
||||
conn.execute(text(
|
||||
"INSERT INTO visual_profiles "
|
||||
"(adventure_id, entity_key, descriptors, features, style_notes, "
|
||||
" created_at, updated_at) "
|
||||
"VALUES (:adv, 'bill', '{\"build\": \"heavyset\"}', '[]', '', "
|
||||
" datetime('now'), datetime('now'))"), {"adv": adv_id})
|
||||
older.dispose()
|
||||
|
||||
result = backup.create(path)
|
||||
try:
|
||||
with sqlite3.connect(f"file:{result.path}?mode=ro", uri=True) as copy_db:
|
||||
row = copy_db.execute(
|
||||
"SELECT entity_key, descriptors FROM visual_profiles").fetchone()
|
||||
assert row[0] == "bill" and "heavyset" in row[1]
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
# ------------------------------------- §14 the move to a machine that never saw it
|
||||
|
||||
@pytest.fixture()
|
||||
def machines():
|
||||
"""Two directories, each with its own database, and a server on each.
|
||||
|
||||
The same shape as `test_m9_clean_import.py`, for the same reason: a shared
|
||||
id space, a warm cache or a session still holding the original would let an
|
||||
in-process import pass while a real move failed. M9's version of this test
|
||||
predates visual profiles and carries none, so this is the profile-carrying
|
||||
half of the same claim rather than a duplicate of it.
|
||||
"""
|
||||
root = tempfile.mkdtemp(prefix="m10-clean-")
|
||||
started: list[Server] = []
|
||||
|
||||
def start(name: str) -> Server:
|
||||
directory = os.path.join(root, name)
|
||||
os.makedirs(directory, exist_ok=True)
|
||||
server = Server(os.path.join(directory, "campaign.db"), _free_port())
|
||||
started.append(server)
|
||||
server.wait_until_ready()
|
||||
return server
|
||||
|
||||
try:
|
||||
yield start
|
||||
finally:
|
||||
for server in started:
|
||||
server.stop()
|
||||
shutil.rmtree(root, ignore_errors=True)
|
||||
|
||||
|
||||
def test_profiles_reach_a_clean_data_directory_on_another_machine(machines):
|
||||
"""§14's Definition-of-Done clause, run across two real processes.
|
||||
|
||||
Machine A plays a campaign, profiles two entities and exports. Machine B is
|
||||
a database file that has never existed before, in a different directory, in
|
||||
a different process — migrations run there from nothing. Nothing crosses but
|
||||
the bundle.
|
||||
"""
|
||||
a = machines("machine-a")
|
||||
campaign = a.call("POST", "/adventures",
|
||||
{"title": "Moving day", "opening": "The office is quiet."},
|
||||
expect=201)
|
||||
adv = campaign["id"]
|
||||
a.call("POST", f"/adventures/{adv}/state/corrections", {
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "alice",
|
||||
"entity_type": "character", "name": "Alice"},
|
||||
{"type": "create_entity", "entity": "roger",
|
||||
"entity_type": "character", "name": "Roger"},
|
||||
{"type": "create_entity", "entity": "office",
|
||||
"entity_type": "location", "name": "The office"},
|
||||
{"type": "set_scene", "summary": "Alice and Roger wait in the office.",
|
||||
"location": "office", "present": ["alice", "roger"]},
|
||||
],
|
||||
"note": "setting the scene",
|
||||
}, expect=201)
|
||||
a.call("PUT", f"/adventures/{adv}/visual-profiles/alice",
|
||||
{"descriptors": {"build": "tall", "hair": "short black"},
|
||||
"features": ["tortoiseshell glasses"],
|
||||
"style_notes": "photographic, natural light"}, expect=200)
|
||||
a.call("PUT", f"/adventures/{adv}/visual-profiles/office",
|
||||
{"descriptors": {"lighting": "flat fluorescent"}}, expect=200)
|
||||
payload = a.call("GET", f"/adventures/{adv}/export", expect=200)
|
||||
source_packet = a.call("GET", f"/adventures/{adv}/scene-packet", expect=200)
|
||||
a.stop()
|
||||
assert not a.is_listening()
|
||||
|
||||
b = machines("machine-b")
|
||||
moved = b.call("POST", "/adventures/import", payload, expect=201)["id"]
|
||||
|
||||
profiles = {p["entity_key"]: p for p in b.call(
|
||||
"GET", f"/adventures/{moved}/visual-profiles", expect=200)["profiles"]}
|
||||
assert set(profiles) == {"alice", "office"}
|
||||
assert profiles["alice"]["features"] == ["tortoiseshell glasses"]
|
||||
assert profiles["alice"]["style_notes"] == "photographic, natural light"
|
||||
|
||||
# The packet the copy builds describes the same scene, with the same
|
||||
# profiles attached and Roger still deliberately unprofiled. Only the
|
||||
# campaign id differs, which is what a new machine's id space means.
|
||||
moved_packet = b.call("GET", f"/adventures/{moved}/scene-packet", expect=200)
|
||||
assert moved_packet["action_summary"] == source_packet["action_summary"]
|
||||
by_name = {c["name"]: c for c in moved_packet["characters"]}
|
||||
assert by_name["Alice"]["visual_profile"]["descriptors"]["hair"] == "short black"
|
||||
assert by_name["Roger"]["visual_profile"] is None
|
||||
assert moved_packet["location"]["visual_profile"]["descriptors"][
|
||||
"lighting"] == "flat fluorescent"
|
||||
@@ -0,0 +1,452 @@
|
||||
"""M10 §4 and §17: scene data obeys the history rules, because it *is* story data.
|
||||
|
||||
The claim this file makes is unusual, and worth stating plainly before the
|
||||
tests: **M10 wrote no lineage code.** There is no media head, no `active` flag,
|
||||
no scene branch table and no separate restore path. The scene lives in the
|
||||
authoritative narrative state document, which M3 gave a head, M4 gave Save
|
||||
Points, M5 gave per-position snapshots and M9 gave portability — so it inherits
|
||||
every one of those rules by being the same data rather than by copying them.
|
||||
|
||||
That makes these tests a check on an inheritance rather than on an
|
||||
implementation, and they are written to fail loudly if the inheritance were ever
|
||||
broken by a future scene store appearing beside the state document. The M10
|
||||
brief's §4 sequence is exercised literally, including the restart, and the
|
||||
Mara-in-the-cellar example it names is the first test.
|
||||
|
||||
python -m pytest tests/test_m10_lineage.py -v
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import tempfile
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.media import packet as scene_packet
|
||||
from app.routers import adventures
|
||||
|
||||
import m10_fixture
|
||||
from fakes import ScriptedProvider
|
||||
from test_process_restart import Server, _free_port
|
||||
|
||||
|
||||
class StubDerived:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25] for _ in texts]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m10lin@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(user_id=user.id, title="Lineage")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text="It begins.",
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
def scene_of(client, adv_id=None):
|
||||
return client.get(
|
||||
f"/api/adventures/{adv_id or client.adv_id}/state"
|
||||
).json()["document"].get("scene") or {}
|
||||
|
||||
|
||||
def packet_of(client, adv_id=None):
|
||||
r = client.get(f"/api/adventures/{adv_id or client.adv_id}/scene-packet")
|
||||
assert r.status_code == 200, r.text[:300]
|
||||
return r.json()
|
||||
|
||||
|
||||
def retained_scenes(adv_id) -> list[tuple]:
|
||||
"""Every scene the tree still holds, as (branch, depth, summary).
|
||||
|
||||
Read from the per-position snapshots, which is where a retained scene lives
|
||||
— the point being that a scene the story left is still on disk, attached to
|
||||
the position that established it.
|
||||
"""
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
with SessionLocal() as db:
|
||||
rows = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == adv_id)
|
||||
.options(undefer(models.Action.narrative_state_after))
|
||||
.order_by(models.Action.branch_id, models.Action.depth, models.Action.id)
|
||||
.all()
|
||||
)
|
||||
out = []
|
||||
for row in rows:
|
||||
state = row.narrative_state_after or {}
|
||||
summary = (state.get("scene") or {}).get("summary")
|
||||
if summary:
|
||||
out.append((row.branch_id, row.depth, summary))
|
||||
return out
|
||||
|
||||
|
||||
# --------------------------------------------- the brief's own §4 example
|
||||
|
||||
def test_a_scene_from_an_abandoned_line_does_not_become_current(client):
|
||||
"""§4, literally: Mara in the cellar, then Mara upstairs.
|
||||
|
||||
Path A's scene must remain stored, must not be current on Path B, and
|
||||
Path B's scene must be Path B's.
|
||||
"""
|
||||
m10_fixture.play(client, client.adv_id, "set up", [
|
||||
m10_fixture.entity("mara", "character", "Mara"),
|
||||
m10_fixture.entity("cellar", "location", "The cellar"),
|
||||
m10_fixture.entity("upstairs", "location", "Upstairs"),
|
||||
])
|
||||
m10_fixture.play(client, client.adv_id, "go down", [
|
||||
{"type": "set_scene", "summary": "Mara enters the cellar.",
|
||||
"location": "cellar", "present": ["mara"]},
|
||||
])
|
||||
assert scene_of(client)["summary"] == "Mara enters the cellar."
|
||||
path_a = packet_of(client)["scene_id"]
|
||||
|
||||
assert client.post(f"/api/adventures/{client.adv_id}/undo").status_code == 200
|
||||
m10_fixture.play(client, client.adv_id, "stay put", [
|
||||
{"type": "set_scene", "summary": "Mara remains upstairs.",
|
||||
"location": "upstairs", "present": ["mara"]},
|
||||
])
|
||||
|
||||
current = scene_of(client)
|
||||
assert current["summary"] == "Mara remains upstairs."
|
||||
assert current["location"] == "upstairs"
|
||||
assert packet_of(client)["location"]["name"] == "Upstairs"
|
||||
assert packet_of(client)["scene_id"] != path_a
|
||||
|
||||
# Path A's scene is still on disk, on the branch it belongs to.
|
||||
kept = retained_scenes(client.adv_id)
|
||||
assert ("Mara enters the cellar." in [s for _, _, s in kept]), kept
|
||||
assert ("Mara remains upstairs." in [s for _, _, s in kept]), kept
|
||||
branches = {s: b for b, _, s in kept}
|
||||
assert branches["Mara enters the cellar."] != branches["Mara remains upstairs."]
|
||||
|
||||
|
||||
def test_divergence_deletes_no_scene(client):
|
||||
"""§4: diverging retains the old line rather than replacing it."""
|
||||
m10_fixture.play(client, client.adv_id, "set up", [
|
||||
m10_fixture.entity("mara", "character", "Mara"),
|
||||
m10_fixture.entity("cellar", "location", "The cellar"),
|
||||
])
|
||||
m10_fixture.play(client, client.adv_id, "down", [
|
||||
{"type": "set_scene", "summary": "Scene A.", "location": "cellar",
|
||||
"present": ["mara"]}])
|
||||
before = len(retained_scenes(client.adv_id))
|
||||
client.post(f"/api/adventures/{client.adv_id}/undo")
|
||||
m10_fixture.play(client, client.adv_id, "elsewhere", [
|
||||
{"type": "set_scene", "summary": "Scene C.", "location": "cellar",
|
||||
"present": ["mara"]}])
|
||||
after = retained_scenes(client.adv_id)
|
||||
assert len(after) == before + 1
|
||||
assert "Scene A." in [s for _, _, s in after]
|
||||
|
||||
|
||||
# ------------------------------------------------- the brief's §17 sequence
|
||||
|
||||
def test_the_full_scene_lineage_sequence(client):
|
||||
"""§17, step by step, in one test so the order is the thing under test.
|
||||
|
||||
Scene A, Save Point, Scene B, Undo, Redo, restore, diverge to Scene C — and
|
||||
at every step the active scene must be the one the head is on, while the
|
||||
scenes the story left must still be on disk.
|
||||
"""
|
||||
adv = client.adv_id
|
||||
m10_fixture.play(client, adv, "set up", [
|
||||
m10_fixture.entity("mara", "character", "Mara"),
|
||||
m10_fixture.entity("hall", "location", "The hall"),
|
||||
])
|
||||
client.put(f"/api/adventures/{adv}/visual-profiles/mara",
|
||||
json={"descriptors": {"build": "sturdy"}})
|
||||
|
||||
# 1-2. Scene A, persisted.
|
||||
m10_fixture.play(client, adv, "scene a", [
|
||||
{"type": "set_scene", "summary": "Scene A.", "location": "hall",
|
||||
"present": ["mara"]}])
|
||||
assert scene_of(client)["summary"] == "Scene A."
|
||||
|
||||
# 3. Save Point at Scene A.
|
||||
point = client.post(f"/api/adventures/{adv}/checkpoints",
|
||||
json={"name": "At scene A", "note": ""})
|
||||
assert point.status_code == 201, point.text[:300]
|
||||
point_id = point.json()["id"]
|
||||
|
||||
# 4. Advance to Scene B.
|
||||
m10_fixture.play(client, adv, "scene b", [
|
||||
{"type": "set_scene", "summary": "Scene B.", "location": "hall",
|
||||
"present": ["mara"]}])
|
||||
assert scene_of(client)["summary"] == "Scene B."
|
||||
|
||||
# 5. Undo -> back at Scene A.
|
||||
assert client.post(f"/api/adventures/{adv}/undo").status_code == 200
|
||||
assert scene_of(client)["summary"] == "Scene A."
|
||||
|
||||
# 6. Redo -> Scene B again.
|
||||
assert client.post(f"/api/adventures/{adv}/redo").status_code == 200
|
||||
assert scene_of(client)["summary"] == "Scene B."
|
||||
|
||||
# 7. Restore the Save Point -> Scene A, and Scene B is still retained.
|
||||
restored = client.post(f"/api/adventures/{adv}/checkpoints/{point_id}/restore")
|
||||
assert restored.status_code == 200, restored.text[:300]
|
||||
assert scene_of(client)["summary"] == "Scene A."
|
||||
assert "Scene B." in [s for _, _, s in retained_scenes(adv)]
|
||||
|
||||
# 8. Diverge to Scene C.
|
||||
m10_fixture.play(client, adv, "scene c", [
|
||||
{"type": "set_scene", "summary": "Scene C.", "location": "hall",
|
||||
"present": ["mara"]}])
|
||||
assert scene_of(client)["summary"] == "Scene C."
|
||||
|
||||
# Scene B is retained and is NOT current on Scene C's line.
|
||||
kept = [s for _, _, s in retained_scenes(adv)]
|
||||
assert "Scene B." in kept and "Scene A." in kept and "Scene C." in kept
|
||||
assert scene_of(client)["summary"] == "Scene C."
|
||||
|
||||
# 9-11. Restart, then inspect again. Nothing about eligibility moved.
|
||||
with SessionLocal() as fresh:
|
||||
adventure = fresh.get(models.Adventure, adv)
|
||||
assert adventure.narrative_state["scene"]["summary"] == "Scene C."
|
||||
|
||||
# The profile is stable across every one of those movements.
|
||||
profile = client.get(f"/api/adventures/{adv}/visual-profiles/mara").json()
|
||||
assert profile["descriptors"] == {"build": "sturdy"}
|
||||
|
||||
|
||||
def test_a_visual_profile_is_stable_across_divergence(client):
|
||||
"""§17: a character does not change appearance because the story forked.
|
||||
|
||||
This is the one place M10's storage choice is directly observable: profiles
|
||||
are campaign-scoped, so the same profile is visible from both lines.
|
||||
"""
|
||||
m10_fixture.play(client, client.adv_id, "set up", [
|
||||
m10_fixture.entity("mara", "character", "Mara"),
|
||||
m10_fixture.entity("hall", "location", "The hall"),
|
||||
])
|
||||
client.put(f"/api/adventures/{client.adv_id}/visual-profiles/mara",
|
||||
json={"descriptors": {"hair": "dark auburn"}})
|
||||
m10_fixture.play(client, client.adv_id, "a", [
|
||||
{"type": "set_scene", "summary": "A.", "location": "hall",
|
||||
"present": ["mara"]}])
|
||||
on_a = packet_of(client)["characters"][0]["visual_profile"]
|
||||
|
||||
client.post(f"/api/adventures/{client.adv_id}/undo")
|
||||
m10_fixture.play(client, client.adv_id, "b", [
|
||||
{"type": "set_scene", "summary": "B.", "location": "hall",
|
||||
"present": ["mara"]}])
|
||||
on_b = packet_of(client)["characters"][0]["visual_profile"]
|
||||
|
||||
assert on_a == on_b == {"descriptors": {"hair": "dark auburn"},
|
||||
"features": [], "style_notes": ""}
|
||||
|
||||
|
||||
def test_a_profile_survives_redo_and_a_save_point_restore(client):
|
||||
"""The other two history operations, for the profile rather than the scene.
|
||||
|
||||
Divergence is covered above and is the interesting case; Redo and a Save
|
||||
Point restore are covered here because K02 claims stability across all of
|
||||
them, and a claim in a report should have a test under it rather than an
|
||||
argument. Both move the head, and a profile that moved with it would be the
|
||||
per-position storage M10 deliberately did not build.
|
||||
"""
|
||||
m10_fixture.play(client, client.adv_id, "set up", [
|
||||
m10_fixture.entity("mara", "character", "Mara"),
|
||||
m10_fixture.entity("hall", "location", "The hall"),
|
||||
])
|
||||
profile = {"descriptors": {"hair": "dark auburn"}, "features": ["a scar"],
|
||||
"style_notes": "candlelight"}
|
||||
client.put(f"/api/adventures/{client.adv_id}/visual-profiles/mara",
|
||||
json=profile)
|
||||
point = client.post(f"/api/adventures/{client.adv_id}/checkpoints",
|
||||
json={"name": "Before the hall", "note": ""})
|
||||
assert point.status_code == 201, point.text[:300]
|
||||
|
||||
m10_fixture.play(client, client.adv_id, "into the hall", [
|
||||
{"type": "set_scene", "summary": "Mara stands in the hall.",
|
||||
"location": "hall", "present": ["mara"]}])
|
||||
expected = {"descriptors": {"hair": "dark auburn"}, "features": ["a scar"],
|
||||
"style_notes": "candlelight"}
|
||||
assert packet_of(client)["characters"][0]["visual_profile"] == expected
|
||||
|
||||
assert client.post(f"/api/adventures/{client.adv_id}/undo").status_code == 200
|
||||
assert client.post(f"/api/adventures/{client.adv_id}/redo").status_code == 200
|
||||
assert packet_of(client)["characters"][0]["visual_profile"] == expected
|
||||
|
||||
restored = client.post(
|
||||
f"/api/adventures/{client.adv_id}/checkpoints/{point.json()['id']}/restore")
|
||||
assert restored.status_code == 200, restored.text[:300]
|
||||
# The scene is gone — it was set after the Save Point — and the profile is
|
||||
# not, which is exactly the difference between story state and presentation
|
||||
# metadata.
|
||||
assert packet_of(client)["characters"] == []
|
||||
assert client.get(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/mara"
|
||||
).json()["descriptors"] == {"hair": "dark auburn"}
|
||||
|
||||
|
||||
def test_nothing_relies_on_a_mutable_active_flag(client):
|
||||
"""§4's last clause, checked structurally rather than by behaviour.
|
||||
|
||||
The scene follows the head because it *is* the state at the head. If a
|
||||
future change introduced a scene table with its own `active` column, this
|
||||
would be the test that noticed.
|
||||
"""
|
||||
assert not hasattr(models, "Scene")
|
||||
columns = {c.name for c in models.VisualProfile.__table__.columns}
|
||||
assert "active" not in columns
|
||||
assert "branch_id" not in columns
|
||||
assert "depth" not in columns
|
||||
|
||||
|
||||
# ------------------------------------------------- a genuine process restart
|
||||
|
||||
@pytest.fixture()
|
||||
def spawned():
|
||||
"""A real server process against a real database file, twice.
|
||||
|
||||
`test_process_restart.py` owns the harness; M10 reuses it because "survives
|
||||
a restart" is a claim about bytes on disk, and a same-process fixture cannot
|
||||
tell durable state from a live object.
|
||||
"""
|
||||
directory = tempfile.mkdtemp(prefix="m10-restart-")
|
||||
db_path = os.path.join(directory, "campaign.db")
|
||||
started: list[Server] = []
|
||||
|
||||
def start() -> Server:
|
||||
server = Server(db_path, _free_port())
|
||||
started.append(server)
|
||||
server.wait_until_ready()
|
||||
return server
|
||||
|
||||
try:
|
||||
yield start, db_path
|
||||
finally:
|
||||
for server in started:
|
||||
server.stop()
|
||||
shutil.rmtree(directory, ignore_errors=True)
|
||||
|
||||
|
||||
def test_scene_and_profile_survive_a_genuine_process_restart(spawned):
|
||||
"""K01/K02/K03's durability clause, across a real PID boundary.
|
||||
|
||||
The spawned server narrates with a deterministic provider that emits no
|
||||
state events, so the scene and the entities are established through the
|
||||
ordinary correction endpoint — which is a real, validated write path, not a
|
||||
fixture reaching into the ORM.
|
||||
"""
|
||||
start, db_path = spawned
|
||||
first = start()
|
||||
campaign = first.call("POST", "/adventures", {
|
||||
"title": "Restarted", "opening": "The office is quiet.",
|
||||
}, expect=201)
|
||||
adv = campaign["id"]
|
||||
|
||||
first.call("POST", f"/adventures/{adv}/state/corrections", {
|
||||
"events": [
|
||||
{"type": "create_entity", "entity": "alice",
|
||||
"entity_type": "character", "name": "Alice"},
|
||||
{"type": "create_entity", "entity": "office",
|
||||
"entity_type": "location", "name": "The office"},
|
||||
{"type": "set_scene", "summary": "Alice waits in the office.",
|
||||
"location": "office", "present": ["alice"]},
|
||||
],
|
||||
"note": "setting the scene",
|
||||
}, expect=201)
|
||||
|
||||
first.call("PUT", f"/adventures/{adv}/visual-profiles/alice",
|
||||
{"descriptors": {"hair": "short black"},
|
||||
"features": ["tortoiseshell glasses"]}, expect=200)
|
||||
|
||||
before_scene = first.call("GET", f"/adventures/{adv}/state",
|
||||
expect=200)["document"]["scene"]
|
||||
before_packet = first.call("GET", f"/adventures/{adv}/scene-packet", expect=200)
|
||||
first.stop()
|
||||
assert not first.is_listening()
|
||||
|
||||
second = start()
|
||||
after_scene = second.call("GET", f"/adventures/{adv}/state",
|
||||
expect=200)["document"]["scene"]
|
||||
after_packet = second.call("GET", f"/adventures/{adv}/scene-packet", expect=200)
|
||||
after_profile = second.call(
|
||||
"GET", f"/adventures/{adv}/visual-profiles/alice", expect=200)
|
||||
|
||||
assert after_scene == before_scene
|
||||
assert after_scene["summary"] == "Alice waits in the office."
|
||||
assert after_packet == before_packet
|
||||
assert after_profile["descriptors"] == {"hair": "short black"}
|
||||
assert after_packet["characters"][0]["visual_profile"]["features"] == [
|
||||
"tortoiseshell glasses"
|
||||
]
|
||||
|
||||
|
||||
def test_the_restarted_database_holds_the_profile_row(spawned):
|
||||
"""Read out of the file itself, so "persisted" is not taken on trust."""
|
||||
start, db_path = spawned
|
||||
server = start()
|
||||
campaign = server.call("POST", "/adventures",
|
||||
{"title": "Rows", "opening": "Start."}, expect=201)
|
||||
adv = campaign["id"]
|
||||
server.call("POST", f"/adventures/{adv}/state/corrections", {
|
||||
"events": [{"type": "create_entity", "entity": "ship",
|
||||
"entity_type": "vehicle", "name": "The Persephone"}],
|
||||
"note": "",
|
||||
}, expect=201)
|
||||
server.call("PUT", f"/adventures/{adv}/visual-profiles/ship",
|
||||
{"descriptors": {"hull": "pitted white composite"}}, expect=200)
|
||||
server.stop()
|
||||
|
||||
connection = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
|
||||
try:
|
||||
row = connection.execute(
|
||||
"SELECT entity_key, descriptors FROM visual_profiles "
|
||||
"WHERE adventure_id = ?", (adv,)
|
||||
).fetchone()
|
||||
finally:
|
||||
connection.close()
|
||||
assert row is not None
|
||||
assert row[0] == "ship"
|
||||
assert json.loads(row[1]) == {"hull": "pitted white composite"}
|
||||
@@ -0,0 +1,568 @@
|
||||
"""M10: the media seam — K01-K04, the packet, the profiles, the contracts.
|
||||
|
||||
Lineage behaviour has its own file (`test_m10_lineage.py`), as does the
|
||||
authority separation (`test_m10_authority.py`) and the no-media claim
|
||||
(`test_m10_no_media.py`), because those three are the claims a reviewer will
|
||||
want to find whole rather than scattered.
|
||||
|
||||
python -m pytest tests/test_m10_media_hooks.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.media import packet as scene_packet
|
||||
from app.media import profiles as visual_profiles
|
||||
from app.media import providers
|
||||
from app.routers import adventures
|
||||
|
||||
import m10_fixture
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
|
||||
class StubDerived:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25] for _ in texts]
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m10@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(user_id=user.id, title="The Office")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start",
|
||||
text="Bill badges in on a Tuesday morning.",
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def office(client):
|
||||
return m10_fixture.build(client, client.adv_id)
|
||||
|
||||
|
||||
def packet_of(client, adv_id=None, **params):
|
||||
response = client.get(
|
||||
f"/api/adventures/{adv_id or client.adv_id}/scene-packet", params=params
|
||||
)
|
||||
assert response.status_code == 200, response.text[:400]
|
||||
return response.json()
|
||||
|
||||
|
||||
def state_of(client, adv_id=None):
|
||||
return client.get(
|
||||
f"/api/adventures/{adv_id or client.adv_id}/state"
|
||||
).json()["document"]
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- K01
|
||||
|
||||
def test_k01_a_structured_scene_is_persisted_for_a_multi_character_scene(
|
||||
client, office
|
||||
):
|
||||
"""K01. A scene with several characters and a clear location, **persisted**.
|
||||
|
||||
The acceptance text forbids satisfying this with an ephemeral dictionary
|
||||
built inside a test, so the assertion is made against what a *second*
|
||||
session reads out of the database — not against a value this test computed.
|
||||
"""
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, client.adv_id)
|
||||
stored = adventure.narrative_state["scene"]
|
||||
|
||||
assert stored["summary"] == "Bill, Alice and Roger meet around the table."
|
||||
assert stored["location"] == "office"
|
||||
assert sorted(stored["present"]) == ["alice", "bill", "roger"]
|
||||
# The coordinate is what makes it a scene *snapshot* rather than a note: it
|
||||
# says which accepted position this describes.
|
||||
assert stored["at"]["branch_id"] is not None
|
||||
assert isinstance(stored["at"]["depth"], int)
|
||||
|
||||
|
||||
def test_k01_the_persisted_scene_is_sufficient_to_depict(client, office):
|
||||
"""Sufficiency, checked as "could something draw this?" rather than "is it non-empty?"."""
|
||||
p = packet_of(client)
|
||||
assert p["location"]["name"] == "The office"
|
||||
assert [c["name"] for c in p["characters"]] == ["Bill", "Alice", "Roger"]
|
||||
assert p["action_summary"] == "Bill, Alice and Roger meet around the table."
|
||||
assert p["objects"] and p["objects"][0]["name"] == "Security badge"
|
||||
assert p["scene_id"]
|
||||
|
||||
|
||||
def test_the_scene_snapshot_is_per_position_and_survives_a_restart(client, office):
|
||||
"""Persisted in the ordinary sense: a new session reads the same thing.
|
||||
|
||||
A genuine process restart is exercised in `test_m10_lineage.py`; this is the
|
||||
cheaper claim that the value is on disk rather than in a live object.
|
||||
"""
|
||||
with SessionLocal() as first:
|
||||
before = first.get(models.Adventure, client.adv_id).narrative_state["scene"]
|
||||
with SessionLocal() as second:
|
||||
after = second.get(models.Adventure, client.adv_id).narrative_state["scene"]
|
||||
assert before == after
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- K02/K03
|
||||
|
||||
def test_k02_a_character_keeps_stable_visual_descriptors(client, office):
|
||||
row = client.get(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alice"
|
||||
).json()
|
||||
assert row["descriptors"]["hair"] == "short black"
|
||||
assert row["features"] == ["tortoiseshell glasses"]
|
||||
assert row["style_notes"] == "photographic, natural light"
|
||||
|
||||
|
||||
def test_k03_a_location_keeps_stable_visual_descriptors(client, office):
|
||||
row = client.get(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/office"
|
||||
).json()
|
||||
assert row["descriptors"]["architecture"] == "open-plan floor"
|
||||
assert row["features"] == ["whiteboard covered in diagrams"]
|
||||
|
||||
|
||||
def test_profiles_survive_more_turns(client, office):
|
||||
"""K02/K03 across turns: playing on does not disturb a profile."""
|
||||
for i in range(3):
|
||||
m10_fixture.play(client, client.adv_id, f"talk {i}", [])
|
||||
row = client.get(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alice"
|
||||
).json()
|
||||
assert row["descriptors"]["hair"] == "short black"
|
||||
|
||||
|
||||
def test_an_item_may_have_a_profile_too(client, office):
|
||||
"""§5's optional third kind, and proof the one table holds all three.
|
||||
|
||||
There is no `kind` column: a character, a location and an item are all
|
||||
entities in the M5 model, and the profile attaches to the entity key.
|
||||
"""
|
||||
response = client.put(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/badge",
|
||||
json={"descriptors": {"material": "white plastic"},
|
||||
"features": ["photo in the corner"]},
|
||||
)
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
assert packet_of(client)["objects"][0]["visual_profile"]["descriptors"] == {
|
||||
"material": "white plastic"
|
||||
}
|
||||
|
||||
|
||||
def test_no_profile_is_distinguishable_from_an_empty_one(client, office):
|
||||
"""A future provider must be able to tell "unstated" from "stated as nothing"."""
|
||||
p = packet_of(client)
|
||||
by_name = {c["name"]: c for c in p["characters"]}
|
||||
assert by_name["Roger"]["visual_profile"] is None
|
||||
assert by_name["Alice"]["visual_profile"] is not None
|
||||
|
||||
client.put(f"/api/adventures/{client.adv_id}/visual-profiles/roger", json={})
|
||||
again = {c["name"]: c for c in packet_of(client)["characters"]}
|
||||
assert again["Roger"]["visual_profile"] == {
|
||||
"descriptors": {}, "features": [], "style_notes": ""
|
||||
}
|
||||
|
||||
|
||||
def test_a_profile_must_name_an_entity_the_campaign_has(client, office):
|
||||
"""A typo is an error, not a row describing nobody."""
|
||||
response = client.put(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alicce",
|
||||
json={"descriptors": {"hair": "short black"}},
|
||||
)
|
||||
assert response.status_code == 400
|
||||
assert "no entity called" in response.json()["detail"]
|
||||
|
||||
|
||||
def test_a_profile_replaces_rather_than_merges(client, office):
|
||||
"""So a descriptor can be removed, which a merge would make impossible."""
|
||||
client.put(f"/api/adventures/{client.adv_id}/visual-profiles/alice",
|
||||
json={"descriptors": {"hair": "short black"}})
|
||||
row = client.get(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alice"
|
||||
).json()
|
||||
assert row["descriptors"] == {"hair": "short black"}
|
||||
assert row["features"] == []
|
||||
|
||||
|
||||
def test_deleting_a_profile_leaves_the_entity_alone(client, office):
|
||||
"""A profile is a description. Removing it removes a description."""
|
||||
assert client.delete(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alice"
|
||||
).status_code == 204
|
||||
assert "alice" in state_of(client)["entities"]
|
||||
assert {c["name"] for c in packet_of(client)["characters"]} == {
|
||||
"Bill", "Alice", "Roger"
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("bad", [
|
||||
{"descriptors": {"hair": ["short", "black"]}},
|
||||
{"descriptors": "short black hair"},
|
||||
{"features": "glasses"},
|
||||
{"style_notes": {"note": "photographic"}},
|
||||
{"descriptors": {"hair": "x" * 5_000}},
|
||||
])
|
||||
def test_a_malformed_profile_is_refused(client, office, bad):
|
||||
response = client.put(
|
||||
f"/api/adventures/{client.adv_id}/visual-profiles/alice", json=bad
|
||||
)
|
||||
assert response.status_code == 400, response.text[:200]
|
||||
|
||||
|
||||
# --------------------------------------------------------- scene identity
|
||||
|
||||
def test_scene_identity_resolves_back_to_a_position(client, office):
|
||||
"""§3. A future asset holding this string can find the accepted scene again."""
|
||||
p = packet_of(client)
|
||||
resolved = scene_packet.parse_scene_id(p["scene_id"])
|
||||
assert resolved["adventure_id"] == client.adv_id
|
||||
assert resolved["branch_id"] == p["turn_range"]["branch_id"]
|
||||
assert resolved["start"] == p["turn_range"]["start"]
|
||||
assert resolved["end"] == p["turn_range"]["end"]
|
||||
|
||||
|
||||
def test_a_scene_may_span_several_turns(client, office):
|
||||
"""§3: one turn is not assumed to be one scene, which a video needs."""
|
||||
p = packet_of(client, start=0, end=4)
|
||||
assert p["turn_range"]["start"] == 0
|
||||
assert p["turn_range"]["end"] == 4
|
||||
assert p["scene_id"].endswith(":0-4")
|
||||
assert scene_packet.parse_scene_id(p["scene_id"])["end"] == 4
|
||||
|
||||
|
||||
def test_a_reversed_range_is_read_in_order(client, office):
|
||||
assert packet_of(client, start=4, end=0)["turn_range"] == \
|
||||
packet_of(client, start=0, end=4)["turn_range"]
|
||||
|
||||
|
||||
def test_several_assets_may_name_one_scene(client, office):
|
||||
"""§3: nothing allocates or records a scene, so nothing bounds how many
|
||||
future assets refer to it. Two builds of the same scene agree exactly."""
|
||||
assert packet_of(client)["scene_id"] == packet_of(client)["scene_id"]
|
||||
|
||||
|
||||
# ------------------------------------------------------- the packet's bounds
|
||||
|
||||
def test_the_packet_does_not_carry_the_transcript(client, office):
|
||||
"""§12. A provider gets the scene, not the campaign."""
|
||||
for i in range(4):
|
||||
m10_fixture.play(client, client.adv_id, f"say something memorable {i}", [],
|
||||
prose=f"Roger tells a long story about the printer {i}.")
|
||||
blob = repr(packet_of(client))
|
||||
assert "printer" not in blob
|
||||
assert "Bill badges in on a Tuesday morning" not in blob
|
||||
|
||||
|
||||
def test_the_packet_carries_no_imported_knowledge_at_all(client, office):
|
||||
"""Not just secrets: imported material as a class stays out.
|
||||
|
||||
A positive control comes with it — the source really was imported and really
|
||||
does reach the narrator — so this cannot pass because the upload failed.
|
||||
"""
|
||||
m10_fixture.upload_handbook(client, client.adv_id)
|
||||
m10_fixture.play(client, client.adv_id, "ask about the north wall panelling", [])
|
||||
|
||||
report = client.get(f"/api/adventures/{client.adv_id}/context").json()
|
||||
assert any("handbook" in r["filename"] for r in report["knowledge"]["used"]), (
|
||||
"the control failed: the narrator never saw the handbook, so this "
|
||||
"proves nothing about the packet"
|
||||
)
|
||||
assert "refurbished" not in repr(packet_of(client))
|
||||
|
||||
|
||||
def test_the_packet_is_bounded_when_the_state_is_large(client, office):
|
||||
"""A scene with many entities does not produce an unbounded packet.
|
||||
|
||||
Thirty extras rather than more, because `set_scene`'s `present` is itself
|
||||
capped at `validate.MAX_LABELS` (40) — asking for more gets the *event*
|
||||
refused and leaves the previous scene standing, which would make this test
|
||||
pass by measuring the wrong scene. The precondition is asserted first for
|
||||
exactly that reason.
|
||||
"""
|
||||
extras = [f"extra_{i}" for i in range(30)]
|
||||
m10_fixture.play(client, client.adv_id, "the whole floor arrives",
|
||||
[m10_fixture.entity(k, "character", f"Extra {k[-2:]}")
|
||||
for k in extras])
|
||||
m10_fixture.play(client, client.adv_id, "everyone crowds in", [
|
||||
{"type": "set_scene", "summary": "The whole floor crowds in.",
|
||||
"location": "office",
|
||||
"present": ["bill", "alice", "roger"] + extras},
|
||||
])
|
||||
present = state_of(client)["scene"]["present"]
|
||||
assert len(present) == 33, (
|
||||
f"the scene was not set as this test intends ({len(present)} present), "
|
||||
f"so the bound below would be measuring the wrong scene"
|
||||
)
|
||||
p = packet_of(client)
|
||||
assert len(p["characters"]) == scene_packet.MAX_CHARACTERS
|
||||
assert len(p["continuity_constraints"]) <= scene_packet.MAX_CONSTRAINTS
|
||||
|
||||
|
||||
# ------------------------------------------------------- provider contracts
|
||||
|
||||
def test_no_provider_is_registered(client):
|
||||
"""v1 ships none, and nothing registers one at import."""
|
||||
assert providers.registered() == {}
|
||||
for kind in providers.MEDIA_KINDS:
|
||||
assert providers.for_kind(kind) == []
|
||||
|
||||
|
||||
def test_a_provider_can_be_added_without_touching_story_code(client, office):
|
||||
"""M10's Definition of Done, as an executable claim.
|
||||
|
||||
A provider is registered, asked to depict the current scene, and returns —
|
||||
and nothing in the story engine was modified, imported or subclassed to make
|
||||
that work. The adapter satisfies a `Protocol`, so it did not even have to
|
||||
import the base class.
|
||||
"""
|
||||
seen = {}
|
||||
|
||||
class FakeImageProvider:
|
||||
def capabilities(self):
|
||||
return providers.ProviderCapabilities(
|
||||
provider_id="fake-local", kinds=(providers.IMAGE,),
|
||||
)
|
||||
|
||||
async def generate(self, request):
|
||||
seen["scene_id"] = request.scene["scene_id"]
|
||||
return providers.MediaResult(
|
||||
kind=providers.IMAGE, media_type="image/png",
|
||||
data=b"\x89PNG\r\n\x1a\n",
|
||||
provenance={"scene_id": request.scene["scene_id"]},
|
||||
)
|
||||
|
||||
provider = FakeImageProvider()
|
||||
assert isinstance(provider, providers.MediaProvider)
|
||||
providers.register("fake-local", provider)
|
||||
try:
|
||||
assert providers.for_kind(providers.IMAGE) == [provider]
|
||||
import asyncio
|
||||
|
||||
p = packet_of(client)
|
||||
result = asyncio.run(provider.generate(
|
||||
providers.MediaRequest(kind=providers.IMAGE, scene=p)
|
||||
))
|
||||
assert result.media_type == "image/png"
|
||||
assert result.provenance["scene_id"] == p["scene_id"]
|
||||
assert seen["scene_id"] == p["scene_id"]
|
||||
finally:
|
||||
providers.unregister("fake-local")
|
||||
assert providers.registered() == {}
|
||||
|
||||
|
||||
def test_every_required_media_kind_is_accommodated(client):
|
||||
assert set(providers.MEDIA_KINDS) == {"image", "video", "audio", "tts", "stt"}
|
||||
|
||||
|
||||
def test_a_request_for_an_unknown_kind_is_refused(client, office):
|
||||
with pytest.raises(ValueError, match="hologram"):
|
||||
providers.MediaRequest(kind="hologram", scene=packet_of(client))
|
||||
|
||||
|
||||
def test_stt_returns_a_draft_and_not_a_result(client):
|
||||
"""§10, and the reason the return type differs.
|
||||
|
||||
A transcription cannot be handed to something expecting a finished artefact,
|
||||
because it is not one — it is text the reader is going to edit.
|
||||
"""
|
||||
class FakeStt:
|
||||
def capabilities(self):
|
||||
return providers.ProviderCapabilities(
|
||||
provider_id="fake-stt", kinds=(providers.STT,))
|
||||
|
||||
async def transcribe(self, audio, hints=None):
|
||||
return providers.DraftTranscription(text="i open teh door")
|
||||
|
||||
import asyncio
|
||||
|
||||
stt = FakeStt()
|
||||
assert isinstance(stt, providers.TranscriptionProvider)
|
||||
draft = asyncio.run(stt.transcribe(b"\x00\x01"))
|
||||
assert isinstance(draft, providers.DraftTranscription)
|
||||
assert not isinstance(draft, providers.MediaResult)
|
||||
assert draft.editable is True
|
||||
|
||||
|
||||
def test_an_stt_draft_has_no_route_into_the_story(client, office):
|
||||
"""The corrected text enters the way anything the reader types does.
|
||||
|
||||
Asserted by playing the edited draft through the ordinary action endpoint
|
||||
and observing that it is an ordinary turn — validated, refereed, snapshotted
|
||||
— rather than by asserting that some bypass does not exist.
|
||||
"""
|
||||
draft = providers.DraftTranscription(text="i open teh door")
|
||||
corrected = draft.text.replace("teh", "the")
|
||||
|
||||
before = len(client.get(f"/api/adventures/{client.adv_id}").json()["actions"])
|
||||
m10_fixture.play(client, client.adv_id, corrected, [])
|
||||
after = client.get(f"/api/adventures/{client.adv_id}").json()["actions"]
|
||||
assert len(after) == before + 2
|
||||
assert after[-2]["text"].endswith("i open the door.")
|
||||
|
||||
|
||||
def test_the_story_engine_holds_no_provider_vocabulary(client):
|
||||
"""§9. Provider syntax must not appear in Story Engine code.
|
||||
|
||||
Greps rather than trusting the boundary, so a future adapter's vocabulary
|
||||
cannot leak in unnoticed.
|
||||
|
||||
**`app/media/` is excluded, and the exclusion is the point rather than a
|
||||
hole.** §9's rule is about the *Story Engine*; `media/` is the seam, and its
|
||||
docstrings name ComfyUI, Whisper and `num_inference_steps` precisely in
|
||||
order to say that those belong to a future adapter and not here. A grep that
|
||||
failed on the sentence forbidding a thing would push the explanation out of
|
||||
the code, which is the opposite of what the rule wants.
|
||||
|
||||
What would catch a violation inside `media/` is not this test but the shape
|
||||
of the package: it registers no provider (`test_no_provider_is_registered`),
|
||||
ships no adapter, and imports nothing that could reach one.
|
||||
"""
|
||||
import pathlib
|
||||
|
||||
root = pathlib.Path(__file__).resolve().parent.parent / "app"
|
||||
seam = root / "media"
|
||||
forbidden = ("comfyui", "stable diffusion", "stable-diffusion", "automatic1111",
|
||||
"num_inference_steps", "cfg_scale", "denoising_strength",
|
||||
"safetensors", "whisper", "kokoro", "flux.1")
|
||||
offenders = []
|
||||
for path in root.rglob("*.py"):
|
||||
if seam in path.parents:
|
||||
continue
|
||||
lowered = path.read_text().lower()
|
||||
for word in forbidden:
|
||||
if word in lowered:
|
||||
offenders.append(f"{path.relative_to(root)}: {word}")
|
||||
assert offenders == [], offenders
|
||||
|
||||
|
||||
def test_the_seam_ships_no_adapter(client):
|
||||
"""The other half of the rule above, for `app/media/` itself.
|
||||
|
||||
The seam is allowed to *name* a provider in prose; it is not allowed to
|
||||
*be* one. Checked by what it does rather than by what it says: no provider
|
||||
registered, and no HTTP client imported anywhere in the package.
|
||||
"""
|
||||
import pathlib
|
||||
|
||||
assert providers.registered() == {}
|
||||
seam = pathlib.Path(__file__).resolve().parent.parent / "app" / "media"
|
||||
for path in seam.rglob("*.py"):
|
||||
body = path.read_text()
|
||||
for client_lib in ("import httpx", "import requests", "urllib.request",
|
||||
"import socket", "subprocess"):
|
||||
assert client_lib not in body, f"{path.name} imports {client_lib}"
|
||||
|
||||
|
||||
# ------------------------------------------------------------ endpoint policy
|
||||
|
||||
def test_a_media_endpoint_must_be_loopback(client):
|
||||
"""§11 and contract §27-28: stricter than the narrator's policy, on purpose."""
|
||||
assert providers.endpoint_rejection_reason("http://127.0.0.1:8188") is None
|
||||
assert providers.endpoint_rejection_reason("http://localhost:8188") is None
|
||||
|
||||
|
||||
def test_a_trusted_lan_media_endpoint_is_refused(client):
|
||||
"""Allowed for narrator inference; not for media, which has no v1 use."""
|
||||
reason = providers.endpoint_rejection_reason("http://192.168.1.50:8188")
|
||||
assert reason is not None
|
||||
assert "on this machine" in reason
|
||||
|
||||
|
||||
@pytest.mark.parametrize("url", [
|
||||
"https://api.example.com/v1",
|
||||
"http://8.8.8.8:8188",
|
||||
"",
|
||||
"not a url",
|
||||
])
|
||||
def test_a_non_local_media_endpoint_is_refused(client, url):
|
||||
assert providers.endpoint_rejection_reason(url) is not None
|
||||
|
||||
|
||||
def test_check_endpoint_raises_for_a_refused_endpoint(client):
|
||||
with pytest.raises(providers.EndpointRejected):
|
||||
providers.check_endpoint("https://api.example.com/v1")
|
||||
providers.check_endpoint("http://127.0.0.1:8188")
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- K04
|
||||
|
||||
def test_k04_the_extension_point_a_future_asset_would_attach_through(client, office):
|
||||
"""K04, on the acceptance text's **deferred** branch — see the M10 report §F.
|
||||
|
||||
No media tables exist, so this demonstrates the equivalent extension point
|
||||
rather than a stored asset: a dummy local byte fixture is carried through
|
||||
the provider contract, and the association it needs is proved to resolve.
|
||||
|
||||
What is actually asserted is the part that would matter to a real asset:
|
||||
the provenance it carries names a scene, that name resolves to an accepted
|
||||
position, and the story is untouched either side.
|
||||
"""
|
||||
p = packet_of(client)
|
||||
before_state = state_of(client)
|
||||
before_actions = client.get(f"/api/adventures/{client.adv_id}").json()["actions"]
|
||||
|
||||
dummy = providers.MediaResult(
|
||||
kind=providers.IMAGE,
|
||||
media_type="image/png",
|
||||
data=b"\x89PNG\r\n\x1a\n\x00fixture",
|
||||
provenance={"scene_id": p["scene_id"],
|
||||
"turn_range": p["turn_range"],
|
||||
"campaign_id": p["campaign"]["id"]},
|
||||
)
|
||||
|
||||
resolved = scene_packet.parse_scene_id(dummy.provenance["scene_id"])
|
||||
assert resolved["adventure_id"] == client.adv_id
|
||||
assert resolved["branch_id"] == p["turn_range"]["branch_id"]
|
||||
|
||||
# The position it names is a real accepted turn in this campaign.
|
||||
with SessionLocal() as db:
|
||||
found = db.query(models.Action).filter(
|
||||
models.Action.adventure_id == client.adv_id,
|
||||
models.Action.branch_id == resolved["branch_id"],
|
||||
models.Action.depth == resolved["end"],
|
||||
).count()
|
||||
assert found >= 1
|
||||
|
||||
# And nothing about the story moved.
|
||||
assert state_of(client) == before_state
|
||||
assert client.get(f"/api/adventures/{client.adv_id}").json()["actions"] == \
|
||||
before_actions
|
||||
@@ -0,0 +1,341 @@
|
||||
"""M10 §8, §19 and §20: the storyteller does not know the media layer is there.
|
||||
|
||||
Three claims, and the first is the milestone's central acceptance condition:
|
||||
|
||||
* **§20 — ordinary play is unchanged** with no media configuration of any kind.
|
||||
Not "works with a warning", not "works once you dismiss something": unchanged.
|
||||
* **§19 — nothing is contacted**, nothing is required at startup, and no
|
||||
provider setting exists to be got wrong.
|
||||
* **§8 — the hidden-information boundary.** A future provider must not receive
|
||||
narrator-only material merely because the storyteller knows it.
|
||||
|
||||
The §8 tests use a **hidden M7 knowledge source**, which is this product's real
|
||||
narrator-only mechanism, rather than an invented marker — so what is tested is
|
||||
the boundary that exists. Each carries a **positive control**: the sentinel is
|
||||
shown to reach the narrator's own prompt in the same campaign, so a passing test
|
||||
cannot be one where the secret was never established.
|
||||
|
||||
python -m pytest tests/test_m10_no_media.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.media import providers
|
||||
from app.routers import adventures
|
||||
|
||||
import m10_fixture
|
||||
from fakes import ScriptedProvider
|
||||
|
||||
|
||||
class StubDerived:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory of the meeting."
|
||||
|
||||
async def embed(self, texts):
|
||||
out = []
|
||||
for text in texts:
|
||||
lowered = text.lower()
|
||||
out.append([
|
||||
1.0,
|
||||
1.0 if "observer" in lowered or "panelling" in lowered else 0.0,
|
||||
1.0 if "office" in lowered or "meeting" in lowered else 0.0,
|
||||
])
|
||||
return out
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="m10nm@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="",
|
||||
context_token_budget=4000, max_output_tokens=400, memory_top_k=3,
|
||||
))
|
||||
adventure = models.Adventure(user_id=user.id, title="No media")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start",
|
||||
text="Bill badges in on a Tuesday morning.",
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
# ---------------------------------------------------- §20: unchanged play
|
||||
|
||||
def test_a_whole_campaign_plays_with_no_media_configuration(client):
|
||||
"""§20's list, in one campaign, with no media anything.
|
||||
|
||||
Turns, state extraction, memory and summary activity, knowledge retrieval,
|
||||
Undo, Redo, Retry, a Save Point restore, and a fresh read of what was
|
||||
written — all of it while no provider is registered, no media endpoint is
|
||||
configured, and no media table holds a row. The genuine process restarts
|
||||
live in `test_m10_lineage.py`.
|
||||
"""
|
||||
adv = client.adv_id
|
||||
assert providers.registered() == {}
|
||||
|
||||
m10_fixture.upload_handbook(client, adv)
|
||||
m10_fixture.build(client, adv)
|
||||
|
||||
for i in range(3):
|
||||
m10_fixture.play(client, adv, f"discuss item {i}", [])
|
||||
|
||||
point = client.post(f"/api/adventures/{adv}/checkpoints",
|
||||
json={"name": "Mid-meeting", "note": ""})
|
||||
assert point.status_code == 201, point.text[:300]
|
||||
|
||||
m10_fixture.play(client, adv, "the meeting runs long", [])
|
||||
assert client.post(f"/api/adventures/{adv}/undo").status_code == 200
|
||||
assert client.post(f"/api/adventures/{adv}/redo").status_code == 200
|
||||
|
||||
retried = client.post(f"/api/adventures/{adv}/retry")
|
||||
assert retried.status_code == 200, retried.text[:300]
|
||||
|
||||
restored = client.post(
|
||||
f"/api/adventures/{adv}/checkpoints/{point.json()['id']}/restore")
|
||||
assert restored.status_code == 200, restored.text[:300]
|
||||
|
||||
import asyncio
|
||||
asyncio.run(memorybank.run_post_turn(adv))
|
||||
|
||||
# Retrieval still works, and the state is intact.
|
||||
report = client.get(f"/api/adventures/{adv}/context").json()
|
||||
assert report["prompt"]["system"]
|
||||
assert client.get(f"/api/adventures/{adv}/state").json()["document"]["entities"]
|
||||
|
||||
# Read back through a fresh session — the state is on disk, not in the
|
||||
# request that wrote it. This is *not* a process restart: the genuine
|
||||
# spawned-process restarts are in `test_m10_lineage.py`, which runs them
|
||||
# with profiles written and packets built.
|
||||
with SessionLocal() as db:
|
||||
assert db.get(models.Adventure, adv).narrative_state["scene"]["summary"]
|
||||
|
||||
|
||||
def test_no_media_row_exists_after_ordinary_play(client):
|
||||
"""Media readiness is inert until something uses it."""
|
||||
m10_fixture.build(client, client.adv_id)
|
||||
for i in range(3):
|
||||
m10_fixture.play(client, client.adv_id, f"turn {i}", [])
|
||||
with SessionLocal() as db:
|
||||
# The fixture writes two profiles deliberately; ordinary *play* writes
|
||||
# none, which is the claim. Counting after a campaign built without the
|
||||
# fixture's profile step would be the same assertion said less clearly.
|
||||
played_only = models.Adventure(user_id=None, title="untouched")
|
||||
db.add(played_only)
|
||||
db.flush()
|
||||
assert db.query(models.VisualProfile).filter(
|
||||
models.VisualProfile.adventure_id == played_only.id).count() == 0
|
||||
|
||||
|
||||
def test_the_prompt_is_unchanged_by_media_readiness(client):
|
||||
"""M10 touches no prompt path, and the assembled prompt shows it.
|
||||
|
||||
The context builder is the one place a new subsystem would leak into every
|
||||
turn. No section M10 could have added appears, and the packet's own
|
||||
vocabulary is absent.
|
||||
"""
|
||||
m10_fixture.build(client, client.adv_id)
|
||||
report = client.get(f"/api/adventures/{client.adv_id}/context").json()
|
||||
labels = {section["label"] for section in report["sections"]}
|
||||
for absent in ("scene_packet", "visual_profile", "visual_profiles", "media"):
|
||||
assert absent not in labels
|
||||
blob = report["prompt"]["system"] + report["prompt"]["story"]
|
||||
assert "visual_profile" not in blob
|
||||
assert "scene_id" not in blob
|
||||
|
||||
|
||||
def test_the_turn_path_does_not_import_the_media_package(client):
|
||||
"""Structural: a turn cannot reach the media layer even by accident.
|
||||
|
||||
Checked on the modules' import statements rather than on their text, so the
|
||||
test says "does not import the media package" and not "does not contain the
|
||||
letters m-e-d-i-a" — which `immediately` would fail.
|
||||
"""
|
||||
import ast
|
||||
import pathlib
|
||||
|
||||
root = pathlib.Path(__file__).resolve().parent.parent / "app"
|
||||
for name in ("routers/adventures/turns.py", "context/builder.py",
|
||||
"narrative/apply.py", "narrative/store.py", "tree.py",
|
||||
"head.py", "memorybank.py"):
|
||||
for node in ast.walk(ast.parse((root / name).read_text())):
|
||||
if isinstance(node, ast.Import):
|
||||
names = [a.name for a in node.names]
|
||||
elif isinstance(node, ast.ImportFrom):
|
||||
names = [node.module or ""] + [a.name for a in node.names]
|
||||
else:
|
||||
continue
|
||||
assert not any(
|
||||
n == "media" or n.endswith(".media") or n.startswith("media.")
|
||||
for n in names
|
||||
), f"{name} imports the media package"
|
||||
|
||||
|
||||
# ----------------------------------------------------- §19: nothing outbound
|
||||
|
||||
def test_no_media_provider_is_required_at_startup(client):
|
||||
"""The application imports, serves and plays with an empty registry."""
|
||||
assert providers.registered() == {}
|
||||
assert client.get("/api/health").json() == {"ok": True}
|
||||
m10_fixture.play(client, client.adv_id, "play a turn", [])
|
||||
|
||||
|
||||
def test_no_media_setting_exists_to_be_misconfigured(client):
|
||||
"""§11's last clause: if no provider configuration is needed, none exists.
|
||||
|
||||
M10 invents no media endpoint setting, so there is nothing to point at a
|
||||
cloud by mistake. The endpoint *policy* exists and is tested; a stored
|
||||
endpoint does not.
|
||||
"""
|
||||
settings = client.get("/api/settings").json()
|
||||
assert not any(
|
||||
"media" in key or "image" in key or "video" in key or "tts" in key
|
||||
or "stt" in key
|
||||
for key in settings
|
||||
), settings.keys()
|
||||
assert not any(
|
||||
"media" in column.name
|
||||
for column in models.Settings.__table__.columns
|
||||
)
|
||||
|
||||
|
||||
def test_the_media_package_opens_no_socket(client):
|
||||
"""§19: no new required outbound connection, checked by import.
|
||||
|
||||
`test_egress.py` owns the general no-outbound guarantee; this is the narrow
|
||||
M10 claim that the new package could not participate in one.
|
||||
"""
|
||||
import pathlib
|
||||
|
||||
seam = pathlib.Path(__file__).resolve().parent.parent / "app" / "media"
|
||||
for path in seam.rglob("*.py"):
|
||||
body = path.read_text()
|
||||
for forbidden in ("httpx", "requests.", "urlopen", "socket.socket",
|
||||
"aiohttp", "subprocess"):
|
||||
assert forbidden not in body, f"{path.name} references {forbidden}"
|
||||
|
||||
|
||||
def test_a_media_endpoint_cannot_be_pointed_at_a_cloud(client):
|
||||
"""The policy, applied where a future coordinator would apply it."""
|
||||
for url in ("https://api.openai.com/v1", "http://8.8.8.8:8188",
|
||||
"https://replicate.com", "http://example.com"):
|
||||
assert providers.endpoint_rejection_reason(url) is not None
|
||||
|
||||
|
||||
# ------------------------------------------- §8: the hidden-information line
|
||||
|
||||
def test_a_narrator_only_secret_does_not_reach_the_scene_packet(client):
|
||||
"""§8, with a positive control.
|
||||
|
||||
The sentinel lives in a **hidden** imported source, which is the product's
|
||||
narrator-only mechanism. The control proves it genuinely reaches the
|
||||
narrator's prompt in this very campaign — so the packet's silence is a
|
||||
boundary rather than an accident of the source never being retrieved.
|
||||
"""
|
||||
adv = client.adv_id
|
||||
m10_fixture.upload_secret(client, adv)
|
||||
m10_fixture.build(client, adv)
|
||||
m10_fixture.play(client, adv, "look at the north wall panelling of the office", [])
|
||||
|
||||
report = client.get(f"/api/adventures/{adv}/context").json()
|
||||
narrator_prompt = report["prompt"]["system"] + report["prompt"]["story"]
|
||||
assert m10_fixture.SECRET_SENTINEL in narrator_prompt, (
|
||||
"the control failed: the narrator was never told the secret, so the "
|
||||
"packet's not containing it proves nothing"
|
||||
)
|
||||
|
||||
packet = client.get(f"/api/adventures/{adv}/scene-packet").json()
|
||||
assert m10_fixture.SECRET_SENTINEL not in repr(packet)
|
||||
assert "concealed observer" not in repr(packet).lower()
|
||||
|
||||
|
||||
def test_the_packet_carries_no_imported_source_even_when_visible(client):
|
||||
"""The boundary is drawn by class, not by filtering secrets one at a time.
|
||||
|
||||
A *visible* reference source is excluded too, which is what makes the rule
|
||||
hold for a secret nobody thought to mark: the packet never reads imported
|
||||
knowledge at all, so there is no filter to forget to apply.
|
||||
"""
|
||||
adv = client.adv_id
|
||||
m10_fixture.upload_handbook(client, adv)
|
||||
m10_fixture.build(client, adv)
|
||||
m10_fixture.play(client, adv, "ask about the north wall panelling", [])
|
||||
|
||||
report = client.get(f"/api/adventures/{adv}/context").json()
|
||||
assert any("handbook" in r["filename"] for r in report["knowledge"]["used"]), (
|
||||
"the control failed: the handbook never reached the narrator"
|
||||
)
|
||||
assert "refurbished" not in repr(
|
||||
client.get(f"/api/adventures/{adv}/scene-packet").json())
|
||||
|
||||
|
||||
def test_a_secret_the_story_accepted_does_reach_the_packet(client):
|
||||
"""The other side of the line, and the reason the rule is the right one.
|
||||
|
||||
Once the *story* establishes something through a validated event, it is no
|
||||
longer narrator-only knowledge — it is something that happened, at a
|
||||
position, in the accepted state. A picture of that scene should show it, and
|
||||
a packet that hid it would be hiding the story from itself.
|
||||
"""
|
||||
adv = client.adv_id
|
||||
m10_fixture.upload_secret(client, adv)
|
||||
m10_fixture.build(client, adv)
|
||||
m10_fixture.play(client, adv, "the panel swings open", [
|
||||
m10_fixture.entity("observer", "character", "The observer"),
|
||||
{"type": "set_scene",
|
||||
"summary": "The panel swings open and the observer steps out.",
|
||||
"location": "office",
|
||||
"present": ["bill", "alice", "roger", "observer"]},
|
||||
])
|
||||
packet = client.get(f"/api/adventures/{adv}/scene-packet").json()
|
||||
assert "The observer" in [c["name"] for c in packet["characters"]]
|
||||
# And still not the sentinel, which the story never said aloud.
|
||||
assert m10_fixture.SECRET_SENTINEL not in repr(packet)
|
||||
|
||||
|
||||
def test_memories_and_summaries_stay_out_of_the_packet(client):
|
||||
"""§7's bound: derived narrative text about the past is not depiction input."""
|
||||
import asyncio
|
||||
|
||||
adv = client.adv_id
|
||||
m10_fixture.build(client, adv)
|
||||
for i in range(8):
|
||||
m10_fixture.play(client, adv, f"talk {i}", [],
|
||||
prose=f"Roger recounts the printer incident again {i}.")
|
||||
asyncio.run(memorybank.run_post_turn(adv))
|
||||
|
||||
packet = client.get(f"/api/adventures/{adv}/scene-packet").json()
|
||||
assert "printer" not in repr(packet)
|
||||
assert "memor" not in repr(packet).lower()
|
||||
@@ -0,0 +1,451 @@
|
||||
"""M9: a consistent copy of the whole database, taken while it is being written.
|
||||
|
||||
`app/backup.py` explains why a plain file copy is not a backup. This file is the
|
||||
evidence for the claim, and the shape of it matters: **every test below opens the
|
||||
backup as its own database and reads what is in it.** A test that only checked a
|
||||
file appeared, or that the endpoint returned 201, would pass against a `cp` — and
|
||||
a `cp` is exactly what this replaces.
|
||||
|
||||
The load test is the one that separates the two. It writes to the source
|
||||
database *while* the backup is being taken, from a second thread, and then asks
|
||||
the copy for a story it can check turn by turn. A page-torn copy would show a
|
||||
transcript with a hole in it, a campaign whose head points past its own story, or
|
||||
a `quick_check` failure — and would show none of those on a quiet database, which
|
||||
is why the quiet case is not the interesting one.
|
||||
|
||||
python -m pytest tests/test_m9_backup.py -v
|
||||
"""
|
||||
|
||||
import os
|
||||
import sqlite3
|
||||
import tempfile
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, backup, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
from fakes import ScriptedProvider, tally_of, tally_reply
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
"""The app, and a campaign with enough in it to recognise afterwards."""
|
||||
Base.metadata.create_all(bind=engine)
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="backup@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(user_id=user.id, model="test-model"))
|
||||
adventure = models.Adventure(user_id=user.id, title="Backed up")
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text="The story opens.",
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def elsewhere(tmp_path, monkeypatch):
|
||||
"""Backups land under a temporary directory, not beside the real database."""
|
||||
fake_db = tmp_path / "campaign.db"
|
||||
fake_db.write_bytes(Path(str(engine.url.database)).read_bytes())
|
||||
return fake_db
|
||||
|
||||
|
||||
def _play(client, text, total):
|
||||
ScriptedProvider.replies = [tally_reply(f"Beat {total // 10}.", total)]
|
||||
response = client.post(f"/api/adventures/{client.adv_id}/actions",
|
||||
json={"type": "do", "text": text})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
|
||||
|
||||
def _open(path) -> sqlite3.Connection:
|
||||
"""The backup, as its own database, read-only."""
|
||||
connection = sqlite3.connect(f"file:{path}?mode=ro", uri=True)
|
||||
connection.row_factory = sqlite3.Row
|
||||
return connection
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the copy
|
||||
|
||||
def test_the_backup_is_a_database_that_passes_its_own_integrity_check(client):
|
||||
for turn in range(1, 4):
|
||||
_play(client, f"turn {turn}", turn * 10)
|
||||
result = backup.create()
|
||||
try:
|
||||
assert result.integrity == "ok"
|
||||
assert result.pages > 0
|
||||
assert result.bytes > 0
|
||||
with _open(result.path) as db:
|
||||
assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
|
||||
assert db.execute("PRAGMA foreign_key_check").fetchall() == []
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_the_backup_holds_the_schema_and_every_family_of_row(client):
|
||||
"""Not "the file exists": the copy is opened and asked what is in it."""
|
||||
for turn in range(1, 4):
|
||||
_play(client, f"turn {turn}", turn * 10)
|
||||
checkpoint = client.post(f"/api/adventures/{client.adv_id}/checkpoints",
|
||||
json={"name": "Here", "note": "A position."})
|
||||
assert checkpoint.status_code == 201
|
||||
upload = client.post(
|
||||
f"/api/adventures/{client.adv_id}/knowledge",
|
||||
files={"file": ("canon.md", b"# Rule\n\nThe dead do not return.\n",
|
||||
"text/markdown")},
|
||||
data={"classification": "canon"},
|
||||
)
|
||||
assert upload.status_code == 201, upload.text[:300]
|
||||
|
||||
result = backup.create()
|
||||
try:
|
||||
with _open(result.path) as db:
|
||||
tables = {
|
||||
row["name"] for row in
|
||||
db.execute("SELECT name FROM sqlite_master WHERE type='table'")
|
||||
}
|
||||
for expected in ("adventures", "actions", "branches", "checkpoints",
|
||||
"knowledge_sources", "knowledge_chunks",
|
||||
"state_events", "summaries", "settings"):
|
||||
assert expected in tables, f"{expected} is missing from the backup"
|
||||
|
||||
campaign = db.execute(
|
||||
"SELECT * FROM adventures WHERE id = ?", (client.adv_id,)
|
||||
).fetchone()
|
||||
assert campaign["title"] == "Backed up"
|
||||
# The head, which is the thing a restore has to reproduce.
|
||||
assert campaign["head_depth"] >= 0
|
||||
assert campaign["head_branch_id"] is not None
|
||||
|
||||
texts = [row["text"] for row in db.execute(
|
||||
"SELECT text FROM actions WHERE adventure_id = ? ORDER BY id",
|
||||
(client.adv_id,),
|
||||
)]
|
||||
assert "The story opens." in texts
|
||||
assert any("Beat 3." in text for text in texts)
|
||||
|
||||
assert db.execute(
|
||||
"SELECT name FROM checkpoints WHERE adventure_id = ?",
|
||||
(client.adv_id,),
|
||||
).fetchone()["name"] == "Here"
|
||||
assert db.execute(
|
||||
"SELECT COUNT(*) c FROM knowledge_sources WHERE adventure_id = ?",
|
||||
(client.adv_id,),
|
||||
).fetchone()["c"] == 1
|
||||
assert db.execute(
|
||||
"SELECT COUNT(*) c FROM state_events WHERE adventure_id = ?",
|
||||
(client.adv_id,),
|
||||
).fetchone()["c"] > 0
|
||||
# And the head names a turn that is actually in the copy.
|
||||
assert db.execute(
|
||||
"SELECT COUNT(*) c FROM actions WHERE adventure_id = ? "
|
||||
"AND branch_id = ? AND depth = ?",
|
||||
(client.adv_id, campaign["head_branch_id"], campaign["head_depth"]),
|
||||
).fetchone()["c"] > 0
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_the_state_in_the_backup_is_the_state_the_campaign_had(client):
|
||||
"""The authoritative document, read out of the copy and compared."""
|
||||
for turn in range(1, 5):
|
||||
_play(client, f"turn {turn}", turn * 10)
|
||||
live = client.get(f"/api/adventures/{client.adv_id}/state").json()["document"]
|
||||
result = backup.create()
|
||||
try:
|
||||
with _open(result.path) as db:
|
||||
from app import compression
|
||||
|
||||
blob = db.execute(
|
||||
"SELECT narrative_state FROM adventures WHERE id = ?",
|
||||
(client.adv_id,),
|
||||
).fetchone()["narrative_state"]
|
||||
assert tally_of(compression.unpack(blob)) == tally_of(live) == 40
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
# ------------------------------------------------------------ while it is live
|
||||
|
||||
def test_a_backup_taken_during_writes_is_consistent(client):
|
||||
"""The claim a plain file copy cannot make.
|
||||
|
||||
Turns are played from a second thread throughout the copy. The backup that
|
||||
comes out is a snapshot of *some* committed point — which point is not
|
||||
determined, and asserting on a particular one would be asserting on a race —
|
||||
so what is checked is that it is a coherent one: `quick_check` passes, no
|
||||
foreign key dangles, the transcript has no gap in it, and the head names a
|
||||
turn that exists.
|
||||
"""
|
||||
stop = threading.Event()
|
||||
written: list[int] = []
|
||||
failures: list[Exception] = []
|
||||
|
||||
def keep_writing():
|
||||
turn = 0
|
||||
while not stop.is_set() and turn < 40:
|
||||
turn += 1
|
||||
try:
|
||||
_play(client, f"concurrent {turn}", turn * 10)
|
||||
written.append(turn)
|
||||
except Exception as exc: # noqa: BLE001 - reported to the test
|
||||
failures.append(exc)
|
||||
return
|
||||
time.sleep(0.005)
|
||||
|
||||
writer = threading.Thread(target=keep_writing, daemon=True)
|
||||
writer.start()
|
||||
# Let a few turns land, so the copy is taken over a database that is moving
|
||||
# rather than one that has not started.
|
||||
while len(written) < 3 and writer.is_alive():
|
||||
time.sleep(0.01)
|
||||
|
||||
result = backup.create()
|
||||
stop.set()
|
||||
writer.join(timeout=30)
|
||||
assert not failures, f"the writer failed: {failures[0]}"
|
||||
assert written, "no turn was written during the backup"
|
||||
|
||||
try:
|
||||
with _open(result.path) as db:
|
||||
assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
|
||||
assert db.execute("PRAGMA foreign_key_check").fetchall() == []
|
||||
|
||||
rows = db.execute(
|
||||
"SELECT depth, type FROM actions WHERE adventure_id = ? "
|
||||
"AND live = 1 ORDER BY depth",
|
||||
(client.adv_id,),
|
||||
).fetchall()
|
||||
depths = [row["depth"] for row in rows]
|
||||
assert depths == list(range(len(depths))), (
|
||||
f"the transcript in the backup has a gap: {depths}"
|
||||
)
|
||||
campaign = db.execute(
|
||||
"SELECT head_branch_id, head_depth FROM adventures WHERE id = ?",
|
||||
(client.adv_id,),
|
||||
).fetchone()
|
||||
assert db.execute(
|
||||
"SELECT COUNT(*) c FROM actions WHERE adventure_id = ? "
|
||||
"AND branch_id = ? AND depth = ?",
|
||||
(client.adv_id, campaign["head_branch_id"], campaign["head_depth"]),
|
||||
).fetchone()["c"] > 0, "the head points past the story in the backup"
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_the_source_database_is_untouched_by_a_backup(client):
|
||||
"""Opened read-only, so this is a guarantee rather than an observation."""
|
||||
_play(client, "one", 10)
|
||||
source = Path(str(engine.url.database))
|
||||
before = source.read_bytes()
|
||||
result = backup.create()
|
||||
try:
|
||||
assert source.read_bytes() == before
|
||||
assert client.get(f"/api/adventures/{client.adv_id}").status_code == 200
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- the rules
|
||||
|
||||
def test_an_existing_backup_is_never_overwritten(client):
|
||||
"""Yesterday's backup surviving today's mistake is most of the point."""
|
||||
first = backup.create()
|
||||
second = backup.create()
|
||||
try:
|
||||
assert first.path != second.path
|
||||
assert first.path.exists() and second.path.exists()
|
||||
finally:
|
||||
first.path.unlink(missing_ok=True)
|
||||
second.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_two_backups_in_the_same_second_do_not_collide(client, monkeypatch):
|
||||
from datetime import datetime
|
||||
|
||||
fixed = datetime(2026, 9, 7, 4, 30, 0)
|
||||
first = backup.create(now=fixed)
|
||||
second = backup.create(now=fixed)
|
||||
try:
|
||||
assert first.path != second.path
|
||||
assert first.path.exists() and second.path.exists()
|
||||
finally:
|
||||
first.path.unlink(missing_ok=True)
|
||||
second.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_a_failed_verification_leaves_nothing_behind(client, monkeypatch):
|
||||
"""A backup nobody verified is a belief, and one that fails is not kept."""
|
||||
monkeypatch.setattr(
|
||||
backup, "_verify",
|
||||
lambda path: (_ for _ in ()).throw(backup.BackupError("bad pages")),
|
||||
)
|
||||
root = backup.directory()
|
||||
before = set(root.iterdir())
|
||||
with pytest.raises(backup.BackupError, match="bad pages"):
|
||||
backup.create()
|
||||
assert set(root.iterdir()) == before, "a failed backup left a file behind"
|
||||
|
||||
|
||||
def test_a_failed_copy_leaves_nothing_behind_and_reports_the_reason(
|
||||
client, monkeypatch
|
||||
):
|
||||
monkeypatch.setattr(
|
||||
backup, "_copy",
|
||||
lambda source, working: (_ for _ in ()).throw(OSError("disk full")),
|
||||
)
|
||||
root = backup.directory()
|
||||
before = set(root.iterdir())
|
||||
with pytest.raises(backup.BackupError, match="disk full"):
|
||||
backup.create()
|
||||
assert set(root.iterdir()) == before
|
||||
|
||||
|
||||
def test_a_missing_source_database_is_reported_rather_than_guessed_at(tmp_path):
|
||||
with pytest.raises(backup.BackupError, match="no database"):
|
||||
backup.create(tmp_path / "not-here.db")
|
||||
|
||||
|
||||
def test_the_partial_file_is_never_left_wearing_a_backups_name(client, monkeypatch):
|
||||
"""The rename is the last step, so an interrupted run is invisible."""
|
||||
seen: list[Path] = []
|
||||
real_copy = backup._copy
|
||||
|
||||
def watch(source, working):
|
||||
seen.append(Path(working))
|
||||
return real_copy(source, working)
|
||||
|
||||
monkeypatch.setattr(backup, "_copy", watch)
|
||||
result = backup.create()
|
||||
try:
|
||||
assert seen and seen[0].name.endswith(".partial")
|
||||
assert not seen[0].exists(), "the temporary file survived"
|
||||
assert result.path.exists()
|
||||
assert not result.path.name.endswith(".partial")
|
||||
finally:
|
||||
result.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
# --------------------------------------------------------------- the endpoint
|
||||
|
||||
def test_the_endpoint_takes_a_backup_and_says_where_it_went(client):
|
||||
response = client.post("/api/backups")
|
||||
assert response.status_code == 201, response.text[:300]
|
||||
body = response.json()
|
||||
path = Path(body["directory"]) / body["filename"]
|
||||
try:
|
||||
assert body["integrity"] == "ok"
|
||||
assert body["bytes"] > 0
|
||||
assert path.exists()
|
||||
with _open(path) as db:
|
||||
assert db.execute("PRAGMA quick_check").fetchone()[0] == "ok"
|
||||
finally:
|
||||
path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_the_endpoint_lists_what_is_there_newest_first(client):
|
||||
"""Ordered by when the backup was taken, which is what its name records.
|
||||
|
||||
Both files here are written in the same instant, so their modification times
|
||||
are indistinguishable and only the stamp in the name says which is which.
|
||||
That is not a contrived case: copying a backup to another disk or restoring
|
||||
one from an archive rewrites its mtime, and a list that reordered itself
|
||||
afterwards would report when the file was last handled rather than when the
|
||||
backup was taken.
|
||||
"""
|
||||
from datetime import datetime
|
||||
|
||||
older = backup.create(now=datetime(2026, 9, 1, 10, 0, 0))
|
||||
newer = backup.create(now=datetime(2026, 9, 6, 10, 0, 0))
|
||||
try:
|
||||
listed = client.get("/api/backups")
|
||||
assert listed.status_code == 200
|
||||
rows = listed.json()["backups"]
|
||||
names = [row["filename"] for row in rows]
|
||||
assert names.index(newer.path.name) < names.index(older.path.name)
|
||||
by_name = {row["filename"]: row["taken_at"] for row in rows}
|
||||
assert by_name[newer.path.name].startswith("2026-09-06T10:00")
|
||||
assert by_name[older.path.name].startswith("2026-09-01T10:00")
|
||||
finally:
|
||||
older.path.unlink(missing_ok=True)
|
||||
newer.path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_a_backup_this_build_did_not_name_still_lists(client):
|
||||
"""A file in the directory whose name carries no stamp is still shown.
|
||||
|
||||
The modification time answers instead. The fallback exists to keep a
|
||||
hand-renamed or third-party file visible rather than silently absent from
|
||||
the list a reader uses to find their backups.
|
||||
"""
|
||||
stray = backup.directory() / f"{backup.PREFIX}-handwritten.db"
|
||||
stray.write_bytes(b"SQLite format 3\x00")
|
||||
try:
|
||||
rows = client.get("/api/backups").json()["backups"]
|
||||
listed = {row["filename"]: row for row in rows}
|
||||
assert stray.name in listed
|
||||
assert listed[stray.name]["taken_at"]
|
||||
finally:
|
||||
stray.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_the_endpoint_accepts_no_path_from_the_caller(client):
|
||||
"""H08. There is no field to attempt a traversal in.
|
||||
|
||||
The destination is derived from the database the application already has
|
||||
open and the name from the clock, so a body is not merely ignored — there is
|
||||
nothing for one to name.
|
||||
"""
|
||||
from app.main import app as application
|
||||
|
||||
schema = application.openapi()["paths"]["/api/backups"]["post"]
|
||||
assert "requestBody" not in schema
|
||||
assert not schema.get("parameters")
|
||||
# And sending one anyway changes nothing about where the file lands.
|
||||
response = client.post("/api/backups", json={"path": "../../../tmp/escape.db"})
|
||||
assert response.status_code == 201, response.text[:300]
|
||||
body = response.json()
|
||||
path = Path(body["directory"]) / body["filename"]
|
||||
try:
|
||||
assert path.parent == backup.directory()
|
||||
assert ".." not in body["filename"]
|
||||
finally:
|
||||
path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def test_a_failure_is_a_clear_error_rather_than_a_silent_success(
|
||||
client, monkeypatch
|
||||
):
|
||||
monkeypatch.setattr(
|
||||
backup, "create",
|
||||
lambda *a, **k: (_ for _ in ()).throw(backup.BackupError("no space left")),
|
||||
)
|
||||
response = client.post("/api/backups")
|
||||
assert response.status_code == 500
|
||||
assert "no space left" in response.json()["detail"]
|
||||
@@ -0,0 +1,460 @@
|
||||
"""M9: the campaign moves to a machine that has never seen it.
|
||||
|
||||
This is the milestone's Definition of Done, and it is the one claim the rest of
|
||||
the M9 suite cannot make. `test_m9_portability.py` imports beside the original,
|
||||
in one process, against one database — which is the right place to check the
|
||||
*contract* and the wrong place to check *portability*. A shared id space, a
|
||||
warm cache, a row the exporter forgot to scope, a session still holding the
|
||||
original: every one of those would pass there and fail here.
|
||||
|
||||
So each test below:
|
||||
|
||||
1. starts a real server process against database A, and plays a campaign;
|
||||
2. exports it over HTTP and stops that process;
|
||||
3. starts a **second** server process against database B, **a file that has
|
||||
never existed before**, in a different directory;
|
||||
4. imports the file over HTTP, and asks the second process what it has.
|
||||
|
||||
Nothing crosses between them but the bundle. Migrations run on B from nothing,
|
||||
because it is a new file — so this is also the fresh-install path, and the
|
||||
"clean data directory" in the Definition of Done is a directory, not a metaphor.
|
||||
|
||||
The final test restarts the *importing* server, which is L03 after a move: a
|
||||
Save Point restored in the third process must reach the same position and the
|
||||
same state as it did in the second.
|
||||
|
||||
python -m pytest tests/test_m9_clean_import.py -v
|
||||
"""
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import sqlite3
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from fakes import TALLY_PER_TURN, tally_of
|
||||
from test_process_restart import Server, _free_port
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def machines():
|
||||
"""Two directories, each with its own database, and the servers on them.
|
||||
|
||||
Two directories rather than two filenames, because the backup directory and
|
||||
anything else the application derives from the database's location must land
|
||||
in the importing machine's own space rather than beside the exporter's.
|
||||
"""
|
||||
root = tempfile.mkdtemp(prefix="m9-clean-")
|
||||
started: list[Server] = []
|
||||
|
||||
def start(name: str) -> Server:
|
||||
directory = os.path.join(root, name)
|
||||
os.makedirs(directory, exist_ok=True)
|
||||
server = Server(os.path.join(directory, "campaign.db"), _free_port())
|
||||
started.append(server)
|
||||
server.wait_until_ready()
|
||||
return server
|
||||
|
||||
def path_of(name: str) -> str:
|
||||
return os.path.join(root, name, "campaign.db")
|
||||
|
||||
try:
|
||||
yield start, path_of
|
||||
finally:
|
||||
for server in started:
|
||||
server.stop()
|
||||
shutil.rmtree(root, ignore_errors=True)
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ building
|
||||
|
||||
def _campaign(server: Server) -> int:
|
||||
"""A campaign with everything a move has to carry, played over HTTP.
|
||||
|
||||
Deliberately not `m9_fixture`: that builds through a `TestClient` and this
|
||||
file exists to avoid one. What it reproduces is the same shape — a retry, a
|
||||
Save Point, an imported source that a turn actually used, an undone head and
|
||||
a retained future.
|
||||
"""
|
||||
adventure = server.call("POST", "/adventures", {
|
||||
"title": "Moved between machines",
|
||||
"canon_rules": ["The dead do not return."],
|
||||
"opening": "Aldric sits in the Crooked Lantern with Mara.",
|
||||
}, expect=201)
|
||||
adv_id = adventure["id"]
|
||||
|
||||
_upload(server, adv_id, "canon.md", "canon", (
|
||||
"# Westhaven\n\n## The Old Abbey\n\nThe abbey above Westhaven has stood "
|
||||
"since the founding. Its crypt is sealed, its door is oak, and the seal "
|
||||
"on it has never been broken.\n"
|
||||
))
|
||||
_upload(server, adv_id, "secret.md", "canon", (
|
||||
"# The seal\n\nIt was broken once, sixty years ago.\n"
|
||||
), visibility="hidden")
|
||||
disabled = _upload(server, adv_id, "draft.md", "reference", (
|
||||
"# Discarded draft\n\nAn earlier version, switched off.\n"
|
||||
))
|
||||
server.call("PATCH", f"/adventures/{adv_id}/knowledge/{disabled}",
|
||||
{"enabled": False}, expect=200)
|
||||
|
||||
# The spawned narrator writes "Beat N." and nothing else, so every term the
|
||||
# retrieval has to work with comes from the player's own words. They are
|
||||
# written to name things the Canon file names.
|
||||
server.play(adv_id, "ask Mara about the abbey crypt in Westhaven")
|
||||
server.play(adv_id, "walk up the hill to the abbey")
|
||||
server.play(adv_id, "try the sealed crypt door of the abbey")
|
||||
_retry(server, adv_id)
|
||||
server.call("POST", f"/adventures/{adv_id}/checkpoints",
|
||||
{"name": "At the door", "note": "Before deciding."}, expect=201)
|
||||
server.play(adv_id, "force the door")
|
||||
server.play(adv_id, "go down the stair")
|
||||
server.call("POST", f"/adventures/{adv_id}/state/corrections", {
|
||||
"events": [{"type": "add_fact", "predicate": "keeper", "value": "Mara",
|
||||
"fact_id": "keeper"}],
|
||||
"note": "Established in play before the state system saw it.",
|
||||
}, expect=201)
|
||||
# Two Undos, so the export is taken behind the retained tip.
|
||||
server.call("POST", f"/adventures/{adv_id}/undo", expect=200)
|
||||
server.call("POST", f"/adventures/{adv_id}/undo", expect=200)
|
||||
return adv_id
|
||||
|
||||
|
||||
def _retry(server: Server, adv_id: int) -> None:
|
||||
"""Retries the newest turn, over the streaming endpoint it actually uses.
|
||||
|
||||
`Server.call` parses JSON, and `/retry` answers with an SSE stream as
|
||||
`/actions` does — so calling it as JSON reads `data: {...}` as a document and
|
||||
fails on the first character. Draining the stream is what the browser does.
|
||||
"""
|
||||
request = urllib.request.Request(
|
||||
f"http://127.0.0.1:{server.port}/api/adventures/{adv_id}/retry",
|
||||
data=b"{}", method="POST",
|
||||
headers={"Content-Type": "application/json"},
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=120) as response:
|
||||
body = response.read()
|
||||
assert b'"type": "error"' not in body, body[:300]
|
||||
|
||||
|
||||
def _upload(server: Server, adv_id: int, name: str, classification: str,
|
||||
body: str, **fields) -> int:
|
||||
"""A multipart knowledge upload over real HTTP, without a client library."""
|
||||
boundary = "----m9cleanimport"
|
||||
parts = []
|
||||
for key, value in {"classification": classification, **fields}.items():
|
||||
parts.append(
|
||||
f"--{boundary}\r\nContent-Disposition: form-data; name=\"{key}\"\r\n"
|
||||
f"\r\n{value}\r\n"
|
||||
)
|
||||
parts.append(
|
||||
f"--{boundary}\r\nContent-Disposition: form-data; name=\"file\"; "
|
||||
f"filename=\"{name}\"\r\nContent-Type: text/markdown\r\n\r\n{body}\r\n"
|
||||
)
|
||||
payload = ("".join(parts) + f"--{boundary}--\r\n").encode()
|
||||
request = urllib.request.Request(
|
||||
f"http://127.0.0.1:{server.port}/api/adventures/{adv_id}/knowledge",
|
||||
data=payload, method="POST",
|
||||
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"},
|
||||
)
|
||||
with urllib.request.urlopen(request, timeout=60) as response:
|
||||
return json.loads(response.read())["id"]
|
||||
|
||||
|
||||
def _snapshot(server: Server, adv_id: int) -> dict:
|
||||
"""What a reader can see, read over HTTP through the API they read."""
|
||||
page = server.call("GET", f"/adventures/{adv_id}", expect=200)
|
||||
return {
|
||||
"title": page["title"],
|
||||
"canon_rules": page["canon_rules"],
|
||||
"transcript": [(a["type"], a["text"]) for a in page["actions"]],
|
||||
"can_undo": page["can_undo"],
|
||||
"can_redo": page["can_redo"],
|
||||
"state": server.call("GET", f"/adventures/{adv_id}/state", expect=200)["document"],
|
||||
"checkpoints": sorted(
|
||||
(c["name"], c["note"], c["depth"])
|
||||
for c in server.call("GET", f"/adventures/{adv_id}/checkpoints", expect=200)
|
||||
),
|
||||
"knowledge": sorted(
|
||||
(k["title"], k["classification"], k["enabled"], k["visibility"],
|
||||
k["content_hash"], k["index_state"], k["chunk_count"] > 0)
|
||||
for k in server.call("GET", f"/adventures/{adv_id}/knowledge", expect=200)
|
||||
),
|
||||
"events": sorted(
|
||||
(e["event_type"], e["source"], json.dumps(e["payload"], sort_keys=True))
|
||||
for e in server.call("GET", f"/adventures/{adv_id}/state/events?limit=500",
|
||||
expect=200)
|
||||
),
|
||||
"rows": server.total_rows(adv_id),
|
||||
}
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- the move
|
||||
|
||||
@pytest.fixture()
|
||||
def moved(machines):
|
||||
"""The campaign, exported from machine A and imported into a clean B."""
|
||||
start, path_of = machines
|
||||
source = start("a")
|
||||
adv_id = _campaign(source)
|
||||
before = _snapshot(source, adv_id)
|
||||
# What the source machine retrieves at this position, recorded while it is
|
||||
# still running. It is the only thing the copy can honestly be compared to.
|
||||
retrieved = {
|
||||
record["filename"] for record in
|
||||
source.call("GET", f"/adventures/{adv_id}/context", expect=200)
|
||||
["knowledge"]["used"]
|
||||
}
|
||||
bundle = source.call("GET", f"/adventures/{adv_id}/export", expect=200)
|
||||
source.stop()
|
||||
|
||||
assert not os.path.exists(path_of("b")), "machine B must not exist yet"
|
||||
target = start("b")
|
||||
assert target.call("GET", "/adventures", expect=200) == [], \
|
||||
"machine B is not empty"
|
||||
|
||||
imported = target.call("POST", "/adventures/import", bundle, expect=201)
|
||||
return {
|
||||
"bundle": bundle, "before": before, "target": target,
|
||||
"retrieved": retrieved,
|
||||
"copy_id": imported["id"], "imported": imported,
|
||||
"path": path_of, "start": start,
|
||||
}
|
||||
|
||||
|
||||
def test_the_campaign_arrives_whole_on_a_machine_that_never_had_it(moved):
|
||||
"""The Definition of Done, in one assertion per family."""
|
||||
after = _snapshot(moved["target"], moved["copy_id"])
|
||||
before = moved["before"]
|
||||
assert after["transcript"] == before["transcript"]
|
||||
assert after["state"] == before["state"]
|
||||
assert after["canon_rules"] == before["canon_rules"]
|
||||
assert after["checkpoints"] == before["checkpoints"]
|
||||
assert after["knowledge"] == before["knowledge"]
|
||||
assert after["events"] == before["events"]
|
||||
assert after["rows"] == before["rows"], "the retained tree is a different size"
|
||||
|
||||
|
||||
def test_it_opens_at_the_exact_head_it_was_exported_at(moved):
|
||||
"""I07, across the boundary the acceptance test names.
|
||||
|
||||
The export was taken two Undos behind the tip, so a machine that opened the
|
||||
campaign at its newest retained turn would show a story two turns longer
|
||||
than the one that was saved.
|
||||
"""
|
||||
after = _snapshot(moved["target"], moved["copy_id"])
|
||||
assert after["transcript"] == moved["before"]["transcript"]
|
||||
assert after["can_redo"] is True, "the retained future is not reachable"
|
||||
assert moved["imported"]["can_redo"] is True, (
|
||||
"the response that opens the campaign says Redo is unavailable"
|
||||
)
|
||||
assert after["rows"] > len(after["transcript"]), (
|
||||
"the retained future is not in the database"
|
||||
)
|
||||
|
||||
|
||||
def test_the_state_audit_arrives_and_still_names_its_author(moved):
|
||||
"""The manual correction is still a manual correction on the new machine."""
|
||||
events = moved["target"].call(
|
||||
f"GET", f"/adventures/{moved['copy_id']}/state/events?limit=500", expect=200
|
||||
)
|
||||
manual = [e for e in events if e["source"] == "manual_correction"]
|
||||
assert len(manual) == 1
|
||||
assert manual[0]["payload"]["predicate"] == "keeper"
|
||||
assert any(e["source"] == "accepted_story" for e in events), (
|
||||
"and the story's own events are there beside it"
|
||||
)
|
||||
|
||||
|
||||
def test_the_knowledge_works_with_no_access_to_the_original_machine(moved):
|
||||
"""§11. The exporting machine is stopped; nothing may reach back to it.
|
||||
|
||||
Its process is dead and its directory holds a database this server has never
|
||||
opened. If retrieval works here, it works from the content the file carried.
|
||||
|
||||
The comparison is against what the *source* retrieved, recorded before that
|
||||
process was killed, and the source's own result is asserted first. A test
|
||||
that only checked the copy retrieved something would pass by accident on a
|
||||
day the fixture happened to match, and — worse — would report a portability
|
||||
failure when what had actually happened is that neither side retrieved
|
||||
anything. That is M8's finding 10: assert your own precondition.
|
||||
"""
|
||||
assert moved["retrieved"], (
|
||||
"the source campaign retrieved nothing, so this proves nothing about "
|
||||
"the copy"
|
||||
)
|
||||
report = moved["target"].call(
|
||||
"GET", f"/adventures/{moved['copy_id']}/context", expect=200
|
||||
)
|
||||
used = {record["filename"] for record in report["knowledge"]["used"]}
|
||||
assert used == moved["retrieved"], (
|
||||
f"the copy retrieved {used} where the source retrieved {moved['retrieved']}"
|
||||
)
|
||||
assert "draft.md" not in used, "the disabled source was re-enabled by the move"
|
||||
assert "canon.md" in used
|
||||
|
||||
|
||||
def test_a_historical_turn_still_shows_what_it_was_given(moved):
|
||||
"""The M8 handoff, across the boundary that made it a handoff.
|
||||
|
||||
Inspect Context on an old narrator turn works on a machine that never
|
||||
assembled that prompt and could not reassemble it — the sources are here but
|
||||
the state, the head and the canon have all moved on since.
|
||||
"""
|
||||
target, copy_id = moved["target"], moved["copy_id"]
|
||||
page = target.call("GET", f"/adventures/{copy_id}/actions?limit=200", expect=200)
|
||||
narrator = [a for a in page["actions"] if a["type"] == "ai"]
|
||||
assert narrator, "the imported campaign has no narrator turn"
|
||||
inspected = 0
|
||||
for action in narrator:
|
||||
response = target.call(
|
||||
"GET", f"/adventures/{copy_id}/actions/{action['id']}/context"
|
||||
)
|
||||
if response is None:
|
||||
continue
|
||||
assert response["prompt"]["system"], "a restored prompt is empty"
|
||||
assert response["sections"], "a restored prompt has no sections"
|
||||
inspected += 1
|
||||
assert inspected, "no turn on the new machine can say what it was told"
|
||||
|
||||
|
||||
def test_no_secret_and_no_path_from_the_old_machine_travelled(moved):
|
||||
"""I06, and the private-detail half of it.
|
||||
|
||||
The bundle is checked as text, because that is what actually left the
|
||||
machine — a field added to a model the exporter walks would reach the file
|
||||
without any test of a column noticing.
|
||||
"""
|
||||
text = json.dumps(moved["bundle"])
|
||||
assert "api_key" not in text
|
||||
assert "11434" not in text, "an inference endpoint travelled with the campaign"
|
||||
assert "/tmp/" not in text and "campaign.db" not in text, (
|
||||
"a filesystem path from the exporting machine travelled"
|
||||
)
|
||||
|
||||
|
||||
def test_the_importing_machine_keeps_its_own_settings(moved):
|
||||
"""§15. A campaign is not a way to reconfigure the destination.
|
||||
|
||||
The bundle carries per-turn model provenance, which is a record of what
|
||||
happened. It does not carry the endpoint, the model or the context budget,
|
||||
because those describe the machine rather than the campaign — and importing
|
||||
a campaign must not silently repoint the destination's inference at the
|
||||
source's.
|
||||
"""
|
||||
settings = moved["target"].call("GET", "/settings", expect=200)
|
||||
assert settings["endpoint_url"] == "http://localhost:11434/v1", (
|
||||
"the import changed the destination's inference endpoint"
|
||||
)
|
||||
assert settings["context_token_budget"] == 16384
|
||||
|
||||
|
||||
def test_a_missing_model_does_not_stop_the_campaign_arriving(moved):
|
||||
"""§15. The campaign and its data are portable independently of a model.
|
||||
|
||||
The importing server has no model configured at all — nothing has ever
|
||||
written a `model` into its settings — and the import still succeeds, opens,
|
||||
and shows its state. Play would fail; recovery does not.
|
||||
"""
|
||||
settings = moved["target"].call("GET", "/settings", expect=200)
|
||||
assert settings["model"] == "", "this test needs an unconfigured destination"
|
||||
after = _snapshot(moved["target"], moved["copy_id"])
|
||||
assert after["transcript"] == moved["before"]["transcript"]
|
||||
|
||||
|
||||
# ------------------------------------------------- L03, after the campaign moved
|
||||
|
||||
def test_l03_a_save_point_restored_on_the_new_machine_survives_its_restart(moved):
|
||||
"""L03, with the move in front of it.
|
||||
|
||||
Restore a Save Point in the second process, record the position and the
|
||||
state, kill the process, start a **third** against the same file, and ask
|
||||
again. What crosses is bytes on disk.
|
||||
"""
|
||||
target, copy_id = moved["target"], moved["copy_id"]
|
||||
points = target.call("GET", f"/adventures/{copy_id}/checkpoints", expect=200)
|
||||
assert points, "the Save Point did not survive the move"
|
||||
point = points[0]
|
||||
assert point["resolved"] is True
|
||||
|
||||
target.call("POST", f"/adventures/{copy_id}/checkpoints/{point['id']}/restore",
|
||||
expect=200)
|
||||
restored = _snapshot(target, copy_id)
|
||||
rows_before = restored["rows"]
|
||||
target.stop()
|
||||
assert not target.is_listening()
|
||||
|
||||
third = moved["start"]("b")
|
||||
again = _snapshot(third, copy_id)
|
||||
assert again["transcript"] == restored["transcript"]
|
||||
assert again["state"] == restored["state"]
|
||||
assert again["rows"] == rows_before, "restoring deleted later history"
|
||||
|
||||
|
||||
# ------------------------------------------------------ the database it wrote
|
||||
|
||||
def test_the_importing_machines_database_passes_its_own_integrity_check(moved):
|
||||
"""A campaign written by an import is a database SQLite is happy with."""
|
||||
moved["target"].stop()
|
||||
connection = sqlite3.connect(moved["path"]("b"))
|
||||
try:
|
||||
assert connection.execute("PRAGMA quick_check").fetchone()[0] == "ok"
|
||||
assert connection.execute("PRAGMA foreign_key_check").fetchall() == []
|
||||
finally:
|
||||
connection.close()
|
||||
|
||||
|
||||
def test_the_import_left_no_orphan_behind(moved):
|
||||
"""§17's list, checked against the database rather than against the API.
|
||||
|
||||
Every one of these would be invisible from the outside until the moment it
|
||||
mattered: a Save Point pointing at a turn that is not there, knowledge owned
|
||||
by a campaign that does not exist, an action on a branch belonging to
|
||||
something else.
|
||||
"""
|
||||
moved["target"].stop()
|
||||
connection = sqlite3.connect(moved["path"]("b"))
|
||||
try:
|
||||
def one(sql):
|
||||
return connection.execute(sql).fetchone()[0]
|
||||
|
||||
assert one("""
|
||||
SELECT COUNT(*) FROM checkpoints c
|
||||
LEFT JOIN actions a
|
||||
ON a.branch_id = c.branch_id AND a.depth = c.depth
|
||||
AND a.adventure_id = c.adventure_id
|
||||
WHERE a.id IS NULL
|
||||
""") == 0, "a Save Point names a position with no turn at it"
|
||||
assert one("""
|
||||
SELECT COUNT(*) FROM actions a
|
||||
LEFT JOIN branches b ON b.id = a.branch_id
|
||||
WHERE a.branch_id IS NOT NULL
|
||||
AND (b.id IS NULL OR b.adventure_id <> a.adventure_id)
|
||||
""") == 0, "an action sits on another campaign's branch"
|
||||
assert one("""
|
||||
SELECT COUNT(*) FROM knowledge_sources k
|
||||
LEFT JOIN adventures adv ON adv.id = k.adventure_id
|
||||
WHERE adv.id IS NULL
|
||||
""") == 0, "knowledge owned by no campaign"
|
||||
assert one("""
|
||||
SELECT COUNT(*) FROM state_events e
|
||||
LEFT JOIN actions a ON a.id = e.action_id
|
||||
WHERE e.action_id IS NOT NULL
|
||||
AND (a.id IS NULL OR a.adventure_id <> e.adventure_id)
|
||||
""") == 0, "a state event names a turn in another campaign"
|
||||
assert one("""
|
||||
SELECT COUNT(*) FROM adventures adv
|
||||
LEFT JOIN actions a
|
||||
ON a.branch_id = adv.head_branch_id AND a.depth = adv.head_depth
|
||||
AND a.adventure_id = adv.id
|
||||
WHERE adv.head_depth >= 0 AND a.id IS NULL
|
||||
""") == 0, "the head points outside the retained story"
|
||||
finally:
|
||||
connection.close()
|
||||
@@ -0,0 +1,747 @@
|
||||
"""M9: what a broken bundle does, and what it must never do.
|
||||
|
||||
A campaign bundle is a file on a disk. It can be truncated by a full volume,
|
||||
mangled by a text editor, hand-written by somebody curious, or produced by a
|
||||
build that does not exist yet. Every case below starts from a real export of the
|
||||
M9 fixture and breaks exactly one thing about it, so what each test measures is
|
||||
that one break rather than a fixture nobody would recognise.
|
||||
|
||||
## The two rules
|
||||
|
||||
**Nothing lands.** A refused import leaves no campaign, no branch, no orphan
|
||||
action, no Save Point pointing at nothing, and no knowledge owned by a campaign
|
||||
that does not exist. `bundle.plan` has no side effects and runs before a row is
|
||||
written, and the endpoint commits once, so a refusal is a refusal — checked here
|
||||
by counting rows before and after rather than by trusting the status code.
|
||||
|
||||
**Nothing is fetched, read or run.** A bundle is data. A URL in it is text, a
|
||||
filename in it is text, and a path in it is text. No test here needs a network
|
||||
guard to pass, which is the point: there is no code path that would use one.
|
||||
|
||||
## Refuse or repair, and why each is which
|
||||
|
||||
The two are not interchangeable and the choice is made per field, on one
|
||||
question — *does a wrong value here make the rest of the campaign wrong?*
|
||||
|
||||
refuse the head, the tree, the audit trail
|
||||
a head past the story misplaces every read of it; a node on a
|
||||
branch that is not listed is a story with a hole; an audit record
|
||||
naming a turn that is not there leaves state nobody can explain
|
||||
repair a knowledge classification that is unreadable, a filename with a
|
||||
path in it, a live flag nobody set
|
||||
the value is not load-bearing for anything but itself
|
||||
drop a Save Point that names no turn, a summary with no coordinate
|
||||
a bookmark costs a bookmark; refusing the campaign to save it
|
||||
would lose the story
|
||||
|
||||
What none of them ever is: **retarget**. A Save Point whose position is not in
|
||||
the file does not get moved to a nearby one, because the reader named a position
|
||||
and no other position is the one they named.
|
||||
|
||||
python -m pytest tests/test_m9_corrupt_bundles.py -v
|
||||
"""
|
||||
|
||||
import copy
|
||||
import json
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
import m9_fixture
|
||||
from fakes import ScriptedProvider
|
||||
from test_m9_portability import StubDerived
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="corrupt@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="stub-embed",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Source campaign",
|
||||
campaign_canon=m9_fixture.CAMPAIGN_CANON,
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def _cache():
|
||||
"""One place to keep the exported fixture between tests in this module."""
|
||||
return {}
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def good(client):
|
||||
"""A real, valid export of the M9 fixture, ready to be broken."""
|
||||
m9_fixture.build(client, client.adv_id)
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/export")
|
||||
assert response.status_code == 200
|
||||
return response.json()
|
||||
|
||||
|
||||
# ------------------------------------------------------------------ the rules
|
||||
|
||||
def _counts() -> dict:
|
||||
"""Every row that an import can create, per table."""
|
||||
with SessionLocal() as db:
|
||||
return {
|
||||
model.__name__: db.query(model).count()
|
||||
for model in (
|
||||
models.Adventure, models.Branch, models.Action, models.Memory,
|
||||
models.Summary, models.Checkpoint, models.StateEvent,
|
||||
models.StateProposal, models.KnowledgeSource,
|
||||
models.KnowledgeChunk, models.StoryCard,
|
||||
)
|
||||
}
|
||||
|
||||
|
||||
def refused(client, payload, *, status=(400, 409, 413, 422)) -> str:
|
||||
"""Imports expecting a refusal, and asserts that nothing at all landed."""
|
||||
before = _counts()
|
||||
response = client.post("/api/adventures/import", json=payload)
|
||||
assert response.status_code in status, (
|
||||
f"expected a refusal, got {response.status_code}: {response.text[:400]}"
|
||||
)
|
||||
assert _counts() == before, (
|
||||
"a refused import wrote rows: "
|
||||
f"{ {k: (before[k], v) for k, v in _counts().items() if before[k] != v} }"
|
||||
)
|
||||
body = response.json()
|
||||
return str(body.get("detail", body))
|
||||
|
||||
|
||||
def accepted(client, payload) -> int:
|
||||
response = client.post("/api/adventures/import", json=payload)
|
||||
assert response.status_code == 201, response.text[:500]
|
||||
return response.json()["id"]
|
||||
|
||||
|
||||
def broken(good: dict, **changes) -> dict:
|
||||
return dict(copy.deepcopy(good), **changes)
|
||||
|
||||
|
||||
# -------------------------------------------------------- format and version
|
||||
|
||||
def test_a_payload_that_is_not_an_object_is_refused(client):
|
||||
for payload in ([], "a string", 7):
|
||||
response = client.post("/api/adventures/import", json=payload)
|
||||
assert response.status_code in (400, 422), response.text[:200]
|
||||
|
||||
|
||||
def test_an_empty_object_is_refused(client):
|
||||
assert "format" in refused(client, {}).lower() or "export" in refused(client, {})
|
||||
|
||||
|
||||
def test_a_missing_format_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
del payload["format"]
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_a_format_of_the_wrong_type_is_refused(client, good):
|
||||
for wrong in (3, None, ["ai-dnd-adventure-v3"], {"v": 3}):
|
||||
refused(client, broken(good, format=wrong))
|
||||
|
||||
|
||||
def test_an_unsupported_future_version_is_refused_with_its_name(client, good):
|
||||
detail = refused(client, broken(good, format="ai-dnd-adventure-v42"))
|
||||
assert "ai-dnd-adventure-v42" in detail
|
||||
|
||||
|
||||
# ------------------------------------------------------------- the tree graph
|
||||
|
||||
def test_an_action_on_a_branch_the_file_does_not_list_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["actions"][0]["branch"] = 99
|
||||
assert "99" in refused(client, payload)
|
||||
|
||||
|
||||
def test_a_branch_forking_from_one_listed_after_it_is_refused(client, good):
|
||||
"""Which is also how a cycle is made impossible rather than detected.
|
||||
|
||||
A branch may only fork from a branch listed before it, so the graph is
|
||||
acyclic by construction. Without it a lineage walk on a hand-edited file
|
||||
would not terminate.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["branches"][0] = {"parent": 1, "forkDepth": 0}
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_a_branch_that_forks_from_itself_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["branches"][1] = {"parent": 1, "forkDepth": 3}
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_a_fork_with_no_depth_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["branches"][1] = {"parent": 0}
|
||||
assert "depth" in refused(client, payload)
|
||||
|
||||
|
||||
def test_an_action_with_no_depth_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["actions"][1]["depth"] = None
|
||||
assert "depth" in refused(client, payload)
|
||||
|
||||
|
||||
def test_an_action_with_a_negative_depth_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["actions"][1]["depth"] = -4
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_a_head_past_the_story_is_refused(client, good):
|
||||
assert "ends at" in refused(client, broken(good, headDepth=10_000))
|
||||
|
||||
|
||||
def test_a_head_depth_of_the_wrong_type_is_refused(client, good):
|
||||
for wrong in ("3", 3.5, True, [3]):
|
||||
refused(client, broken(good, headDepth=wrong))
|
||||
|
||||
|
||||
def test_a_head_branch_that_is_not_listed_falls_back_to_the_root(client, good):
|
||||
"""Repaired rather than refused, and the repair is the safe direction.
|
||||
|
||||
The head *depth* is checked against the story and refused when it disagrees,
|
||||
because a wrong depth silently moves the reader. A head *branch* that names
|
||||
nothing cannot be read at all, so there is no wrong position to land at —
|
||||
the root is where a campaign with no chosen branch is read.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["headBranch"] = 77
|
||||
payload.pop("headDepth") # the depth belongs to the branch it names
|
||||
copy_id = accepted(client, payload)
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, copy_id)
|
||||
root = (
|
||||
db.query(models.Branch)
|
||||
.filter(models.Branch.adventure_id == copy_id,
|
||||
models.Branch.parent_branch_id.is_(None))
|
||||
.first()
|
||||
)
|
||||
assert adventure.head_branch_id == root.id
|
||||
|
||||
|
||||
def test_two_actions_claiming_one_identity_are_refused(client, good):
|
||||
"""Take parentage and the whole audit trail hang off these ids."""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["actions"][1]["id"] = payload["actions"][0]["id"]
|
||||
assert "both call themselves" in refused(client, payload)
|
||||
|
||||
|
||||
def test_a_turn_whose_takes_are_all_dead_still_tells_one(client, good):
|
||||
"""Repaired, because a turn with no live attempt disappears from the story."""
|
||||
payload = copy.deepcopy(good)
|
||||
for action in payload["actions"]:
|
||||
action["live"] = False
|
||||
copy_id = accepted(client, payload)
|
||||
with SessionLocal() as db:
|
||||
rows = (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == copy_id)
|
||||
.all()
|
||||
)
|
||||
per_turn = {}
|
||||
for row in rows:
|
||||
per_turn.setdefault((row.branch_id, row.depth), []).append(row)
|
||||
for group in per_turn.values():
|
||||
assert sum(1 for row in group if row.live) == 1
|
||||
|
||||
|
||||
def test_a_parent_naming_a_node_the_file_does_not_hold_is_ignored(client, good):
|
||||
"""Dropped, not refused: a wrong parent costs a pager, not a campaign."""
|
||||
payload = copy.deepcopy(good)
|
||||
for action in payload["actions"]:
|
||||
if action.get("parentId") is not None:
|
||||
action["parentId"] = 999_999
|
||||
copy_id = accepted(client, payload)
|
||||
story = client.get(f"/api/adventures/{copy_id}").json()
|
||||
assert story["actions"], "the campaign did not import"
|
||||
|
||||
|
||||
def test_a_node_that_is_its_own_parent_does_not_loop(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
for action in payload["actions"]:
|
||||
if action.get("id") is not None:
|
||||
action["parentId"] = action["id"]
|
||||
copy_id = accepted(client, payload)
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.Action).filter(
|
||||
models.Action.adventure_id == copy_id,
|
||||
models.Action.parent_id == models.Action.id,
|
||||
).count() == 0
|
||||
# And the pager still resolves rather than recursing.
|
||||
assert client.get(f"/api/adventures/{copy_id}").status_code == 200
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- save points
|
||||
|
||||
def test_a_save_point_beyond_the_retained_story_is_dropped_not_retargeted(
|
||||
client, good
|
||||
):
|
||||
payload = copy.deepcopy(good)
|
||||
original = payload["checkpoints"][0]["name"]
|
||||
payload["checkpoints"][0]["depth"] = 5_000
|
||||
copy_id = accepted(client, payload)
|
||||
landed = client.get(f"/api/adventures/{copy_id}/checkpoints").json()
|
||||
assert original not in {point["name"] for point in landed}
|
||||
assert all(point["depth"] < 5_000 for point in landed)
|
||||
assert landed, "the good Save Point was lost with the bad one"
|
||||
|
||||
|
||||
def test_a_save_point_on_a_branch_that_is_not_listed_is_dropped(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["checkpoints"][0]["branch"] = 44
|
||||
copy_id = accepted(client, payload)
|
||||
landed = client.get(f"/api/adventures/{copy_id}/checkpoints").json()
|
||||
assert len(landed) == len(good["checkpoints"]) - 1
|
||||
|
||||
|
||||
def test_a_save_point_with_a_blank_name_is_dropped(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["checkpoints"][0]["name"] = " "
|
||||
copy_id = accepted(client, payload)
|
||||
assert len(client.get(f"/api/adventures/{copy_id}/checkpoints").json()) == \
|
||||
len(good["checkpoints"]) - 1
|
||||
|
||||
|
||||
def test_a_checkpoints_section_that_is_not_a_list_costs_the_bookmarks_only(
|
||||
client, good
|
||||
):
|
||||
copy_id = accepted(client, broken(good, checkpoints={"nope": 1}))
|
||||
assert client.get(f"/api/adventures/{copy_id}/checkpoints").json() == []
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["actions"]
|
||||
|
||||
|
||||
# ------------------------------------------------------------ state and audit
|
||||
|
||||
def test_a_state_section_that_is_not_a_list_is_refused(client, good):
|
||||
assert "list" in refused(client, broken(good, stateEvents={"a": 1}))
|
||||
assert "list" in refused(client, broken(good, stateProposals="events"))
|
||||
|
||||
|
||||
def test_a_state_event_that_is_not_an_object_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["stateEvents"][0] = "an event"
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_a_state_event_with_no_type_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["stateEvents"][0]["eventType"] = ""
|
||||
assert "type" in refused(client, payload)
|
||||
|
||||
|
||||
def test_an_event_naming_a_turn_the_file_does_not_hold_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["stateEvents"][0]["action"] = 424_242
|
||||
assert "424242" in refused(client, payload).replace(",", "")
|
||||
|
||||
|
||||
def test_a_proposal_naming_a_turn_the_file_does_not_hold_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["stateProposals"][0]["action"] = 424_242
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_an_event_naming_a_proposal_that_is_gone_keeps_its_coordinate(client, good):
|
||||
"""`ON DELETE SET NULL`, as a file. The event is the accepted change.
|
||||
|
||||
A proposal can be deleted while the event it produced stands — the schema
|
||||
says so — so an event whose proposal is not in the file is not a broken
|
||||
file. It loses the pointer and keeps everything that makes it an audit
|
||||
record: what changed, where, and who asserted it.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["stateProposals"] = []
|
||||
copy_id = accepted(client, payload)
|
||||
events = client.get(
|
||||
f"/api/adventures/{copy_id}/state/events?limit=500"
|
||||
).json()
|
||||
assert len(events) == len(good["stateEvents"])
|
||||
assert any(e["source"] == "manual_correction" for e in events)
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.StateEvent).filter(
|
||||
models.StateEvent.adventure_id == copy_id,
|
||||
models.StateEvent.proposal_id.isnot(None),
|
||||
).count() == 0
|
||||
|
||||
|
||||
def test_a_malformed_narrative_state_costs_the_state_and_not_the_campaign(
|
||||
client, good
|
||||
):
|
||||
"""M5's rule, unchanged: a malformed document is normalised, not fatal.
|
||||
|
||||
The story is the valuable thing. A state section that arrives as nonsense
|
||||
becomes an empty document — which is honest, because nothing in it can be
|
||||
trusted — and every turn still imports.
|
||||
"""
|
||||
copy_id = accepted(client, broken(good, narrativeState={"entities": "wrong"}))
|
||||
story = client.get(f"/api/adventures/{copy_id}").json()
|
||||
assert len(story["actions"]) == len(
|
||||
client.get(f"/api/adventures/{client.adv_id}").json()["actions"]
|
||||
)
|
||||
state = client.get(f"/api/adventures/{copy_id}/state").json()
|
||||
assert state["document"]["entities"] == {}
|
||||
|
||||
|
||||
def test_a_per_position_snapshot_that_is_not_an_object_is_dropped(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
for action in payload["actions"]:
|
||||
if "narrativeStateAfter" in action:
|
||||
action["narrativeStateAfter"] = "not a document"
|
||||
copy_id = accepted(client, payload)
|
||||
assert client.get(f"/api/adventures/{copy_id}").status_code == 200
|
||||
# Arriving at such a position gives the empty document rather than a
|
||||
# later position's state, which is M5's finding 3.
|
||||
client.post(f"/api/adventures/{copy_id}/undo")
|
||||
assert client.get(f"/api/adventures/{copy_id}/state").json()["document"]["facts"] == []
|
||||
|
||||
|
||||
# -------------------------------------------------------------- knowledge
|
||||
|
||||
def test_a_knowledge_section_that_is_not_a_list_is_refused(client, good):
|
||||
assert "list" in refused(client, broken(good, knowledge={"a": 1}))
|
||||
|
||||
|
||||
def test_a_source_with_no_content_is_refused(client, good):
|
||||
"""Refused rather than dropped, and M7 chose that deliberately.
|
||||
|
||||
A campaign whose imported Canon quietly did not arrive is a campaign whose
|
||||
narrator has stopped being told the rules, and the reader has no way to
|
||||
notice.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["knowledge"][0]["content"] = ""
|
||||
assert "content" in refused(client, payload)
|
||||
|
||||
|
||||
def test_a_source_with_an_unknown_classification_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["knowledge"][0]["classification"] = "gospel"
|
||||
assert "classification" in refused(client, payload)
|
||||
|
||||
|
||||
def test_a_source_that_is_not_an_object_is_refused(client, good):
|
||||
payload = copy.deepcopy(good)
|
||||
payload["knowledge"][0] = "canon.md"
|
||||
refused(client, payload)
|
||||
|
||||
|
||||
def test_an_unreadable_visibility_becomes_normal_rather_than_hidden(client, good):
|
||||
"""Repaired, and in the direction that reveals rather than conceals.
|
||||
|
||||
Visibility is not a permission system — the person who imported the file can
|
||||
always read it — so a source that should have been narrator-only and lands
|
||||
as normal costs a spoiler in the prompt framing. The other direction would
|
||||
silently withhold material the reader expects the narrator to use, with
|
||||
nothing saying so.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
for source in payload["knowledge"]:
|
||||
source["visibility"] = "invisible"
|
||||
copy_id = accepted(client, payload)
|
||||
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
|
||||
assert all(source["visibility"] == "normal" for source in library)
|
||||
|
||||
|
||||
def test_a_content_hash_that_disagrees_is_recomputed_and_reported(client, good):
|
||||
"""The one derived value in the file, and the only reason it is there.
|
||||
|
||||
The stored hash is recomputed from what actually arrived, so it always
|
||||
describes the content. The file's own claim is not silently discarded
|
||||
either: a mismatch means the file was edited after it was written, and the
|
||||
reader is told on the source itself.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["knowledge"][0]["contentHash"] = "0" * 64
|
||||
copy_id = accepted(client, payload)
|
||||
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
|
||||
edited = [s for s in library if s["content_hash"] != "0" * 64]
|
||||
assert len(edited) == len(library)
|
||||
detail = client.get(
|
||||
f"/api/adventures/{copy_id}/knowledge/{library[0]['id']}"
|
||||
).json()
|
||||
assert "did not match" in detail["notes"]
|
||||
|
||||
|
||||
def test_more_sources_than_the_cap_is_refused(client, good, monkeypatch):
|
||||
from app.knowledge import importer
|
||||
|
||||
monkeypatch.setattr(importer, "MAX_SOURCES_PER_ADVENTURE", 2)
|
||||
assert "limit" in refused(client, good)
|
||||
|
||||
|
||||
def test_an_oversized_source_is_refused(client, good, monkeypatch):
|
||||
from app.knowledge import importer
|
||||
|
||||
monkeypatch.setattr(importer, "MAX_SOURCE_BYTES", 32)
|
||||
assert "larger than" in refused(client, good)
|
||||
|
||||
|
||||
# ------------------------------------------------------------- provenance
|
||||
|
||||
def test_a_context_snapshot_that_is_not_an_object_is_dropped(client, good):
|
||||
"""Evidence is restored verbatim or not at all. It is never guessed at."""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["actions"] = [
|
||||
{k: v for k, v in action.items() if k != "contextSnapshotZ"}
|
||||
| ({"contextSnapshot": "the prompt was long"}
|
||||
if m9_fixture.snapshot_in(action) else {})
|
||||
for action in payload["actions"]
|
||||
]
|
||||
copy_id = accepted(client, payload)
|
||||
page = client.get(f"/api/adventures/{copy_id}").json()
|
||||
narrator = [a for a in page["actions"] if a["type"] == "ai"]
|
||||
assert narrator
|
||||
for action in narrator:
|
||||
response = client.get(
|
||||
f"/api/adventures/{copy_id}/actions/{action['id']}/context"
|
||||
)
|
||||
assert response.status_code == 404, "a mangled snapshot was restored"
|
||||
|
||||
|
||||
def test_a_snapshot_whose_knowledge_block_is_nonsense_does_not_break_the_import(
|
||||
client, good
|
||||
):
|
||||
payload = copy.deepcopy(good)
|
||||
rewritten = []
|
||||
for action in payload["actions"]:
|
||||
snapshot = m9_fixture.snapshot_in(action)
|
||||
if isinstance(snapshot, dict) and "knowledge" in snapshot:
|
||||
snapshot["knowledge"] = ["not", "a", "report"]
|
||||
rewritten.append(m9_fixture.with_snapshot(action, snapshot))
|
||||
else:
|
||||
rewritten.append(action)
|
||||
payload["actions"] = rewritten
|
||||
copy_id = accepted(client, payload)
|
||||
assert client.get(f"/api/adventures/{copy_id}").status_code == 200
|
||||
|
||||
|
||||
def test_a_snapshot_naming_an_impossible_source_is_relinked_to_nothing(
|
||||
client, good
|
||||
):
|
||||
payload = copy.deepcopy(good)
|
||||
rewritten = []
|
||||
for action in payload["actions"]:
|
||||
snapshot = m9_fixture.snapshot_in(action)
|
||||
if not isinstance(snapshot, dict):
|
||||
rewritten.append(action)
|
||||
continue
|
||||
for record in (snapshot.get("knowledge") or {}).get("used") or []:
|
||||
record["source_id"] = -1
|
||||
rewritten.append(m9_fixture.with_snapshot(action, snapshot))
|
||||
payload["actions"] = rewritten
|
||||
copy_id = accepted(client, payload)
|
||||
with SessionLocal() as db:
|
||||
from sqlalchemy.orm import undefer
|
||||
|
||||
for row in (
|
||||
db.query(models.Action)
|
||||
.filter(models.Action.adventure_id == copy_id)
|
||||
.options(undefer(models.Action.context_snapshot))
|
||||
):
|
||||
snapshot = row.context_snapshot
|
||||
if not isinstance(snapshot, dict):
|
||||
continue
|
||||
for record in (snapshot.get("knowledge") or {}).get("used") or []:
|
||||
assert record["source_id"] is None
|
||||
|
||||
|
||||
# ------------------------------------------------------------ summaries
|
||||
|
||||
def test_a_summary_with_no_coordinate_is_dropped_not_placed(client, good):
|
||||
"""Placing it at a guess is how E03's leak would arrive by a new route."""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["summaries"][0]["depth"] = None
|
||||
copy_id = accepted(client, payload)
|
||||
with SessionLocal() as db:
|
||||
landed = db.query(models.Summary).filter(
|
||||
models.Summary.adventure_id == copy_id
|
||||
).count()
|
||||
assert landed == len(good["summaries"]) - 1
|
||||
|
||||
|
||||
def test_a_summaries_section_that_is_not_a_list_costs_the_summaries_only(
|
||||
client, good
|
||||
):
|
||||
copy_id = accepted(client, broken(good, summaries="a paragraph"))
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["actions"]
|
||||
with SessionLocal() as db:
|
||||
assert db.query(models.Summary).filter(
|
||||
models.Summary.adventure_id == copy_id
|
||||
).count() == 0
|
||||
|
||||
|
||||
# ------------------------------------------------------------ caps and size
|
||||
|
||||
def test_more_actions_than_the_cap_is_refused(client, good, monkeypatch):
|
||||
monkeypatch.setattr(limits, "MAX_ACTIONS_PER_ADVENTURE", 3)
|
||||
monkeypatch.setitem(limits._BUNDLE_LIST_CAPS, "actions", 3)
|
||||
assert "limit" in refused(client, good)
|
||||
|
||||
|
||||
def test_more_branches_than_the_cap_is_refused(client, good, monkeypatch):
|
||||
monkeypatch.setattr(limits, "MAX_BRANCHES_PER_ADVENTURE", 1)
|
||||
monkeypatch.setitem(limits._BUNDLE_LIST_CAPS, "branches", 1)
|
||||
assert "limit" in refused(client, good)
|
||||
|
||||
|
||||
def test_a_body_past_the_import_ceiling_is_refused_before_it_is_parsed(client):
|
||||
"""413 from the middleware, on the declared length, before any read."""
|
||||
padding = "x" * (limits.MAX_IMPORT_BODY_BYTES + 1024)
|
||||
response = client.post(
|
||||
"/api/adventures/import",
|
||||
content=json.dumps({"format": "ai-dnd-adventure-v3", "title": padding}),
|
||||
headers={"Content-Type": "application/json"},
|
||||
)
|
||||
assert response.status_code == 413
|
||||
assert "too large" in response.json()["detail"].lower()
|
||||
|
||||
|
||||
# ------------------------------------------------- the transaction, not the plan
|
||||
|
||||
def test_a_failure_deep_inside_the_write_leaves_nothing_behind(
|
||||
client, good, monkeypatch
|
||||
):
|
||||
"""The other half of atomicity, and the half the planner cannot provide.
|
||||
|
||||
Every test above is refused by `bundle.plan`, which has no side effects — so
|
||||
they prove the *planner*, and a passing planner would look identical if the
|
||||
write phase left debris. This one breaks something the planner has already
|
||||
approved, half way through writing: the branches, the nodes, their
|
||||
parentage, the memories, the head and the Save Points are all in the session
|
||||
by then.
|
||||
|
||||
What must survive that is the whole transaction rolling back — every table,
|
||||
not merely the adventure row. A half-written campaign is the outcome L01
|
||||
forbids for a turn, and an import is the other place it could happen.
|
||||
"""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
def explode(*args, **kwargs):
|
||||
raise RuntimeError("simulated failure deep inside the write")
|
||||
|
||||
monkeypatch.setattr(bundle_module, "_write_summaries", explode)
|
||||
before = _counts()
|
||||
with pytest.raises(RuntimeError, match="simulated failure"):
|
||||
client.post("/api/adventures/import", json=good)
|
||||
assert _counts() == before, (
|
||||
"a failed write left rows behind: "
|
||||
f"{ {k: (before[k], v) for k, v in _counts().items() if before[k] != v} }"
|
||||
)
|
||||
|
||||
|
||||
def test_the_session_is_usable_after_a_failed_import(client, good, monkeypatch):
|
||||
"""The rollback is explicit, so the next request is not poisoned by it.
|
||||
|
||||
Left to the session closing, a failure would leave the request's session in
|
||||
a state the next caller inherits only by luck of pooling. `bundle_io` rolls
|
||||
back and re-raises, so the very next import succeeds.
|
||||
"""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
calls = {"n": 0}
|
||||
original = bundle_module._write_summaries
|
||||
|
||||
def once(*args, **kwargs):
|
||||
calls["n"] += 1
|
||||
if calls["n"] == 1:
|
||||
raise RuntimeError("simulated, once")
|
||||
return original(*args, **kwargs)
|
||||
|
||||
monkeypatch.setattr(bundle_module, "_write_summaries", once)
|
||||
with pytest.raises(RuntimeError):
|
||||
client.post("/api/adventures/import", json=good)
|
||||
copy_id = accepted(client, good)
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["actions"]
|
||||
|
||||
|
||||
# --------------------------------------------------------------- inert data
|
||||
|
||||
def test_a_url_in_a_bundle_stays_text(client, good):
|
||||
"""H01/G08 for the import path: nothing in a file is ever fetched.
|
||||
|
||||
There is no allowlist to test and no request to intercept, which is the
|
||||
result rather than a gap — the import has no code that could make one. What
|
||||
is asserted is that the text arrives as text.
|
||||
"""
|
||||
payload = copy.deepcopy(good)
|
||||
payload["knowledge"][0]["content"] = (
|
||||
"# Sources\n\nSee https://example.invalid/secret.txt and "
|
||||
"file:///etc/passwd and \n"
|
||||
)
|
||||
copy_id = accepted(client, payload)
|
||||
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
|
||||
detail = client.get(
|
||||
f"/api/adventures/{copy_id}/knowledge/{library[0]['id']}"
|
||||
).json()
|
||||
assert "https://example.invalid/secret.txt" in detail["content"]
|
||||
|
||||
|
||||
def test_a_path_in_a_bundle_never_becomes_a_path(client, good):
|
||||
"""H08. `originalFilename` is metadata; the import stores no file."""
|
||||
payload = copy.deepcopy(good)
|
||||
for hostile in ("../../../etc/passwd", "/etc/shadow", "C:\\Windows\\hosts",
|
||||
"....//....//etc/passwd"):
|
||||
payload["knowledge"][0]["originalFilename"] = hostile
|
||||
copy_id = accepted(client, payload)
|
||||
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
|
||||
for source in library:
|
||||
assert "/" not in source["original_filename"]
|
||||
assert "\\" not in source["original_filename"]
|
||||
assert ".." not in source["original_filename"]
|
||||
|
||||
|
||||
def test_a_title_that_looks_like_a_command_is_stored_as_a_title(client, good):
|
||||
payload = broken(good, title="; rm -rf / #")
|
||||
copy_id = accepted(client, payload)
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["title"] == "; rm -rf / #"
|
||||
|
||||
|
||||
def test_an_over_long_title_is_truncated_rather_than_refused(client, good):
|
||||
copy_id = accepted(client, broken(good, title="A" * 5_000))
|
||||
title = client.get(f"/api/adventures/{copy_id}").json()["title"]
|
||||
assert 0 < len(title) <= 200
|
||||
@@ -0,0 +1,365 @@
|
||||
"""M9: every older bundle still imports, and none is reinterpreted.
|
||||
|
||||
A backup that stops importing is not a backup, so the importer keeps every
|
||||
version it has ever written. That is the easy half. The hard half is the rule
|
||||
`V1-ACCEPTANCE-TESTS.md` I07 states about the head and this file generalises:
|
||||
|
||||
> Do not reinterpret missing legacy data using modern assumptions that did not
|
||||
> exist when the file was written.
|
||||
|
||||
An older file is missing things because its **format** could not carry them, not
|
||||
because the campaign lacked them, and the two demand opposite treatment. A file
|
||||
written before the head was carried opens at its tip, because tip was the only
|
||||
position that format could represent — reproducing what it recorded. A file
|
||||
written before state events existed opens with no state events, because
|
||||
manufacturing an audit trail from the snapshots it does carry would be this
|
||||
build's reading of a history it never saw, handed to a reader as the record of
|
||||
what happened.
|
||||
|
||||
Each seam below is built by taking a real v3 export and removing exactly what
|
||||
the older format could not hold. That is deliberate: a checked-in fixture file
|
||||
drifts, and a hand-written one tests a shape nothing ever wrote.
|
||||
|
||||
python -m pytest tests/test_m9_legacy_bundles.py -v
|
||||
"""
|
||||
|
||||
import copy
|
||||
|
||||
import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, bundle, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.knowledge import embeddings
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
|
||||
import m9_fixture
|
||||
from fakes import ScriptedProvider
|
||||
from test_m9_portability import StubDerived
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def client(monkeypatch):
|
||||
Base.metadata.create_all(bind=engine)
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
setup = SessionLocal()
|
||||
user = models.User(is_guest=False, email="legacy@example.com")
|
||||
setup.add(user)
|
||||
setup.flush()
|
||||
setup.add(models.Settings(
|
||||
user_id=user.id, model="test-model", embedding_model="stub-embed",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Source",
|
||||
campaign_canon=m9_fixture.CAMPAIGN_CANON,
|
||||
)
|
||||
setup.add(adventure)
|
||||
setup.flush()
|
||||
setup.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
|
||||
))
|
||||
setup.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
setup.close()
|
||||
|
||||
monkeypatch.setattr(limits, "check_row_cap", lambda *a, **k: None)
|
||||
monkeypatch.setattr(adventures.turns, "OpenAICompatibleProvider", ScriptedProvider)
|
||||
monkeypatch.setattr(memorybank, "embedding_provider", lambda s: StubDerived())
|
||||
monkeypatch.setattr(memorybank, "summary_provider", lambda s: StubDerived())
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
test_client = TestClient(app)
|
||||
test_client.adv_id = adv_id
|
||||
try:
|
||||
yield test_client
|
||||
finally:
|
||||
app.dependency_overrides.clear()
|
||||
adventures.turns._active_turns.clear()
|
||||
memorybank._vector_cache.clear()
|
||||
embeddings._cache.clear()
|
||||
Base.metadata.drop_all(bind=engine)
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def current(client):
|
||||
"""A real v3 export of the M9 fixture, to age backwards from."""
|
||||
m9_fixture.build(client, client.adv_id)
|
||||
response = client.get(f"/api/adventures/{client.adv_id}/export")
|
||||
assert response.status_code == 200
|
||||
return response.json()
|
||||
|
||||
|
||||
# ---------------------------------------------------- ageing a bundle backwards
|
||||
|
||||
def as_of(payload: dict, era: str) -> dict:
|
||||
"""The same campaign as an export from an earlier era.
|
||||
|
||||
Each step removes only what that era's format genuinely could not carry, so
|
||||
the result is the file a build of that vintage would have produced from this
|
||||
campaign — not a mutilated modern one.
|
||||
"""
|
||||
older = copy.deepcopy(payload)
|
||||
eras = ("pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head")
|
||||
assert era in eras, era
|
||||
reached = eras.index(era)
|
||||
|
||||
# M9 (v3): the evidence sections and the node identities.
|
||||
older["format"] = bundle.TREE_FORMAT
|
||||
for key in ("stateEvents", "stateProposals", "summaries"):
|
||||
older.pop(key, None)
|
||||
for action in older["actions"]:
|
||||
for key in ("contextSnapshot", "contextSnapshotZ", "id", "parentId"):
|
||||
action.pop(key, None)
|
||||
for memory in older.get("memories") or []:
|
||||
memory.pop("authority", None)
|
||||
for source in older.get("knowledge") or []:
|
||||
for key in ("sourceId", "parserVersion", "chunkingVersion"):
|
||||
source.pop(key, None)
|
||||
if reached == 0:
|
||||
return older
|
||||
|
||||
# M7: the imported knowledge library.
|
||||
older.pop("knowledge", None)
|
||||
if reached == 1:
|
||||
return older
|
||||
|
||||
# M5: the authoritative narrative state, its per-position snapshots, and
|
||||
# the campaign's own canon.
|
||||
for key in ("narrativeState", "campaignCanon"):
|
||||
older.pop(key, None)
|
||||
for action in older["actions"]:
|
||||
for key in ("narrativeStateAfter", "stateChanges"):
|
||||
action.pop(key, None)
|
||||
if reached == 2:
|
||||
return older
|
||||
|
||||
# M4: named Save Points.
|
||||
older.pop("checkpoints", None)
|
||||
if reached == 3:
|
||||
return older
|
||||
|
||||
# M3: the chosen head. Such a file could only ever be read at its tip.
|
||||
older.pop("headDepth", None)
|
||||
return older
|
||||
|
||||
|
||||
def bring_back(client, payload) -> int:
|
||||
response = client.post("/api/adventures/import", json=payload)
|
||||
assert response.status_code == 201, response.text[:500]
|
||||
return response.json()["id"]
|
||||
|
||||
|
||||
def _rows(adv_id, model) -> int:
|
||||
with SessionLocal() as db:
|
||||
return db.query(model).filter(model.adventure_id == adv_id).count()
|
||||
|
||||
|
||||
def _tree_size(client, adv_id) -> int:
|
||||
"""Every retained row, which is what "no accepted story was lost" means."""
|
||||
return len(client.get(f"/api/adventures/{adv_id}/export").json()["actions"])
|
||||
|
||||
|
||||
# --------------------------------------------------------------- every era
|
||||
|
||||
@pytest.mark.parametrize("era", [
|
||||
"pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head",
|
||||
])
|
||||
def test_no_accepted_story_is_lost_at_any_seam(client, current, era):
|
||||
"""The floor under every case below: the turns all arrive.
|
||||
|
||||
Counted over the whole retained tree rather than the active path, because
|
||||
the head moves between eras and a count of what is on screen would move
|
||||
with it.
|
||||
"""
|
||||
copy_id = bring_back(client, as_of(current, era))
|
||||
assert _tree_size(client, copy_id) == len(current["actions"])
|
||||
|
||||
|
||||
@pytest.mark.parametrize("era", [
|
||||
"pre-m9", "pre-m7", "pre-m5", "pre-save-points", "pre-active-head",
|
||||
])
|
||||
def test_nothing_is_invented_to_fill_a_gap_the_format_left(client, current, era):
|
||||
"""Absent means the format could not say. It never means "make one up".
|
||||
|
||||
Each era is checked against what that era's files could hold: a pre-M9 file
|
||||
gets no audit trail and no summaries, a pre-M7 file no knowledge, a pre-M5
|
||||
file no state, a pre-Save-Point file no Save Points.
|
||||
"""
|
||||
copy_id = bring_back(client, as_of(current, era))
|
||||
reached = ("pre-m9", "pre-m7", "pre-m5", "pre-save-points",
|
||||
"pre-active-head").index(era)
|
||||
|
||||
assert _rows(copy_id, models.StateEvent) == 0
|
||||
assert _rows(copy_id, models.StateProposal) == 0
|
||||
assert _rows(copy_id, models.Summary) == 0
|
||||
if reached >= 1:
|
||||
assert _rows(copy_id, models.KnowledgeSource) == 0
|
||||
assert _rows(copy_id, models.KnowledgeChunk) == 0
|
||||
if reached >= 2:
|
||||
state = client.get(f"/api/adventures/{copy_id}/state").json()
|
||||
assert state["document"]["facts"] == []
|
||||
assert state["document"]["entities"] == {}
|
||||
assert client.get(f"/api/adventures/{copy_id}").json()["canon_rules"] == []
|
||||
if reached >= 3:
|
||||
assert _rows(copy_id, models.Checkpoint) == 0
|
||||
|
||||
|
||||
# ------------------------------------------------------ the head, era by era
|
||||
|
||||
def test_a_pre_m9_file_still_opens_at_the_head_it_recorded(client, current):
|
||||
"""v2 carried the head, so it is honoured exactly as before."""
|
||||
copy_id = bring_back(client, as_of(current, "pre-m9"))
|
||||
with SessionLocal() as db:
|
||||
assert db.get(models.Adventure, copy_id).head_depth == current["headDepth"]
|
||||
|
||||
|
||||
def test_a_pre_active_head_file_opens_at_its_tip(client, current):
|
||||
"""I07's compatibility clause. Not a degraded path.
|
||||
|
||||
Such a file was written when the head could not be anywhere but the tip, so
|
||||
opening it there reproduces the position it recorded. An import that refused
|
||||
it, or that guessed some other position, would be the failure.
|
||||
"""
|
||||
copy_id = bring_back(client, as_of(current, "pre-active-head"))
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, copy_id)
|
||||
tip = max(
|
||||
row.depth for row in
|
||||
db.query(models.Action).filter(
|
||||
models.Action.adventure_id == copy_id,
|
||||
models.Action.branch_id == adventure.head_branch_id,
|
||||
)
|
||||
)
|
||||
assert adventure.head_depth == tip
|
||||
assert adventure.head_depth > current["headDepth"], (
|
||||
"the fixture's head must really be behind its tip, or this proves nothing"
|
||||
)
|
||||
|
||||
|
||||
def test_a_pre_active_head_file_offers_no_redo_because_it_is_at_the_tip(
|
||||
client, current
|
||||
):
|
||||
copy_id = bring_back(client, as_of(current, "pre-active-head"))
|
||||
page = client.get(f"/api/adventures/{copy_id}").json()
|
||||
assert page["can_redo"] is False
|
||||
assert page["can_undo"] is True
|
||||
|
||||
|
||||
# ----------------------------------------------------- what each era can do
|
||||
|
||||
def test_a_pre_m5_campaign_can_be_played_on_and_gains_state_from_there(
|
||||
client, current
|
||||
):
|
||||
"""The M5 rule, applied to an import: no backfill, and no obstacle either.
|
||||
|
||||
An old campaign starts with an empty state because its narration was never
|
||||
read by a state extractor. The next turn fills it in, which is what makes
|
||||
"no backfill" a decision rather than a loss.
|
||||
"""
|
||||
copy_id = bring_back(client, as_of(current, "pre-m5"))
|
||||
assert client.get(f"/api/adventures/{copy_id}/state").json()["empty"] is True
|
||||
|
||||
ScriptedProvider.replies = [
|
||||
"The door gives at last.\n" + __import__("fakes").state_block([
|
||||
{"type": "add_fact", "predicate": "tally", "value": 500,
|
||||
"fact_id": "tally-500"}
|
||||
])
|
||||
]
|
||||
played = client.post(f"/api/adventures/{copy_id}/actions",
|
||||
json={"type": "do", "text": "push harder"})
|
||||
assert played.status_code == 200, played.text[:300]
|
||||
after = client.get(f"/api/adventures/{copy_id}/state").json()
|
||||
assert after["empty"] is False
|
||||
assert any(f["predicate"] == "tally" for f in after["document"]["facts"])
|
||||
|
||||
|
||||
def test_a_pre_m7_campaign_needs_no_source_and_can_import_one(client, current):
|
||||
copy_id = bring_back(client, as_of(current, "pre-m7"))
|
||||
assert client.get(f"/api/adventures/{copy_id}/knowledge").json() == []
|
||||
# It plays without one.
|
||||
assert client.get(f"/api/adventures/{copy_id}/context").status_code == 200
|
||||
# And gains one.
|
||||
landed = m9_fixture.upload(
|
||||
client, copy_id, "canon.md", m9_fixture.CANON_MD, "canon",
|
||||
)
|
||||
library = client.get(f"/api/adventures/{copy_id}/knowledge").json()
|
||||
assert [s["id"] for s in library] == [landed]
|
||||
assert library[0]["index_state"] == "ready"
|
||||
|
||||
|
||||
def test_a_pre_save_point_campaign_can_be_given_one(client, current):
|
||||
copy_id = bring_back(client, as_of(current, "pre-save-points"))
|
||||
assert client.get(f"/api/adventures/{copy_id}/checkpoints").json() == []
|
||||
made = client.post(f"/api/adventures/{copy_id}/checkpoints",
|
||||
json={"name": "From here", "note": ""})
|
||||
assert made.status_code == 201, made.text[:300]
|
||||
assert made.json()["resolved"] is True
|
||||
|
||||
|
||||
def test_a_pre_m9_campaign_re_exports_as_v3_without_gaining_evidence(
|
||||
client, current
|
||||
):
|
||||
"""Re-exporting an old campaign does not turn absence into presence.
|
||||
|
||||
The file it writes is a v3 file, because that is what this build writes. Its
|
||||
evidence sections are empty, because the campaign genuinely has none — and a
|
||||
later reader can therefore trust a v3 file's empty `stateEvents` to mean
|
||||
"this campaign has no audit trail" rather than "the file could not say".
|
||||
"""
|
||||
copy_id = bring_back(client, as_of(current, "pre-m9"))
|
||||
again = client.get(f"/api/adventures/{copy_id}/export").json()
|
||||
assert again["format"] == bundle.FORMAT
|
||||
assert again["stateEvents"] == []
|
||||
assert again["stateProposals"] == []
|
||||
assert again["summaries"] == []
|
||||
assert not any(a.get("contextSnapshotZ") for a in again["actions"])
|
||||
# And the story it does have survives a second round trip unchanged.
|
||||
twice = bring_back(client, again)
|
||||
assert _tree_size(client, twice) == _tree_size(client, copy_id)
|
||||
|
||||
|
||||
def test_a_v1_file_still_imports_and_reads_in_order(client):
|
||||
"""The flat format, with its retries as a repeating group."""
|
||||
copy_id = bring_back(client, {
|
||||
"format": bundle.LEGACY_FORMAT,
|
||||
"title": "An old flat file",
|
||||
"memory": "Kept from before the tree.",
|
||||
"actions": [
|
||||
{"index": 0, "type": "start", "text": "It begins."},
|
||||
{"index": 1, "type": "do", "text": "look around"},
|
||||
{"index": 2, "type": "ai", "text": "Take two.",
|
||||
"variants": [{"text": "Take one."}, {"text": "Take two."}],
|
||||
"variantIndex": 1},
|
||||
],
|
||||
})
|
||||
page = client.get(f"/api/adventures/{copy_id}").json()
|
||||
assert [a["text"] for a in page["actions"]] == [
|
||||
"It begins.", "look around", "Take two.",
|
||||
]
|
||||
assert page["memory"] == "Kept from before the tree."
|
||||
# Both attempts arrived; only one is the story.
|
||||
with SessionLocal() as db:
|
||||
rows = db.query(models.Action).filter(
|
||||
models.Action.adventure_id == copy_id, models.Action.type == "ai",
|
||||
).all()
|
||||
assert sorted(r.text for r in rows) == ["Take one.", "Take two."]
|
||||
assert sum(1 for r in rows if r.live) == 1
|
||||
|
||||
|
||||
def test_a_pre_m2_file_with_scripting_still_imports(client, current):
|
||||
"""M2 removed campaign scripting. Its keys are ignored, not rejected.
|
||||
|
||||
The story, the tree and everything else in such a file are still worth
|
||||
importing, and refusing the campaign over a subsystem that no longer exists
|
||||
would lose all of it to reject one key.
|
||||
"""
|
||||
payload = as_of(current, "pre-m5")
|
||||
payload["scripts"] = [{"name": "onTurn", "code": "state.gold += 10"}]
|
||||
payload["scriptState"] = {"gold": 70}
|
||||
copy_id = bring_back(client, payload)
|
||||
assert _tree_size(client, copy_id) == len(current["actions"])
|
||||
File diff suppressed because it is too large
Load Diff
@@ -350,13 +350,17 @@ def test_export_and_import_round_trips_variants(client):
|
||||
assert [a["type"] for a in actions] == ["start", "do", "ai"]
|
||||
assert actions[-1]["text"] == "Two."
|
||||
|
||||
# The pager reads 1/1 on the copy, because the import writes no
|
||||
# `parent_id` and `annotate_takes` groups on it. The attempts are both
|
||||
# there, at one coordinate, and `GET .../variants` still lists them. This
|
||||
# is a gap in the import rather than in the drop: `take_count` has been the
|
||||
# only number the client reads since SP9, and the import has never set the
|
||||
# column it is derived from.
|
||||
assert actions[-1]["take_count"] == 1
|
||||
# The pager reads 2/2 on the copy, as it does on the original.
|
||||
#
|
||||
# It read 1/1 until M9, and this test recorded that as a gap in the import
|
||||
# rather than in the export: the attempts were both there at one coordinate
|
||||
# and `GET .../variants` listed them, but the import wrote no `parent_id`,
|
||||
# so `annotate_takes` grouped on the coordinate instead. That is right for a
|
||||
# plain retry and wrong the moment two takes of one turn each have takes of
|
||||
# their own beneath them, which is why M9 carried the parentage rather than
|
||||
# leaving the pager to a fallback. See `bundle._link_take_parents`.
|
||||
assert actions[-1]["take_count"] == 2
|
||||
assert actions[-1]["take_index"] == 1
|
||||
variants = client.get(
|
||||
f"/api/adventures/{imported}/actions/{actions[-1]['id']}/variants").json()
|
||||
assert [v["text"] for v in variants] == ["One.", "Two."]
|
||||
|
||||
@@ -50,10 +50,20 @@ def payload() -> dict:
|
||||
|
||||
|
||||
def test_the_shipped_file_is_a_bundle_this_build_can_import():
|
||||
"""The file is written by an export, so a format change can strand it."""
|
||||
"""The file is written by an export, so a format change can strand it.
|
||||
|
||||
It is checked against every version the importer reads rather than against
|
||||
the newest one it writes, which is the property that actually matters and
|
||||
the one the shipped file has to keep. M9 bumped the format to v3 and did not
|
||||
regenerate this asset: the starter is a linear story with no state events,
|
||||
no summaries and no stored prompts, so a v3 rewrite of it would differ from
|
||||
the v2 file in the version string alone — and rewriting a shipped asset to
|
||||
keep a test's equality holding would be changing the evidence to fit the
|
||||
test. What it does need is to go on importing, which is asserted below.
|
||||
"""
|
||||
data = payload()
|
||||
version = bundle.check_format(data)
|
||||
assert version == bundle.FORMAT
|
||||
assert version in bundle.READABLE
|
||||
story = bundle.plan(data, version)
|
||||
assert story["nodes"]
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ import pytest
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, models
|
||||
from app import auth, bundle as bundle_module, limits, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
@@ -425,6 +425,13 @@ def test_export_carries_the_whole_story(client):
|
||||
array in `test_export_keeps_retry_attempts`. That change reflects the
|
||||
same fact: a bundle that stores coordinates has no use for a repeating
|
||||
group. Everything else here still passes unmodified.
|
||||
|
||||
M9 changed the same one line again, for the same kind of reason — the
|
||||
version now says that the file can carry state events and historical
|
||||
prompts as well as a tree. It is asserted against `bundle.FORMAT` this
|
||||
time, so the next writer of a new version does not have to find this line:
|
||||
what the test is about is that an export declares its version, not which
|
||||
version this build happens to write.
|
||||
"""
|
||||
ScriptedProvider.replies = [gold_reply(t) for t in ["One.", "Two."]]
|
||||
_play(client, "go north")
|
||||
@@ -433,7 +440,7 @@ def test_export_carries_the_whole_story(client):
|
||||
r = client.get(f"/api/adventures/{client.adv_id}/export")
|
||||
assert r.status_code == 200, r.text
|
||||
bundle = r.json()
|
||||
assert bundle["format"] == "ai-dnd-adventure-v2"
|
||||
assert bundle["format"] == bundle_module.FORMAT
|
||||
assert bundle["title"] == "Cave"
|
||||
assert [a["text"] for a in bundle["actions"]] == [
|
||||
OPENING, "> You go north.", "One.", "> You go south.", "Two.",
|
||||
|
||||
@@ -0,0 +1,225 @@
|
||||
"""What M10 costs a campaign, measured rather than argued.
|
||||
|
||||
python -m tools.m10_media_cost [--turns 60]
|
||||
|
||||
Run from `backend/`. Plays a campaign of `--turns` turns with the real prompt
|
||||
builder and the real state pipeline, then reports the five numbers §21 of the
|
||||
M10 brief asks for.
|
||||
|
||||
Four of them are expected to be zero or near it, and that is the point: M10's
|
||||
central design decision was that **the scene snapshot already exists**, so the
|
||||
milestone persists nothing per scene and nothing per turn. A design claim like
|
||||
that is cheap to make and easy to get wrong by one accidental write, so it is
|
||||
measured here against a campaign long enough for a per-turn cost to show.
|
||||
|
||||
scene records written by M10 expected 0, and the scenes that do
|
||||
exist are M5's, counted for contrast
|
||||
bytes added to the database one row per profiled entity, once
|
||||
profile duplication what per-position profiles would have
|
||||
cost, against what campaign-scoped
|
||||
profiles do cost
|
||||
packet: persisted or constructed rows written while building one
|
||||
current-scene query behaviour statements per packet, at 10 turns and
|
||||
at N turns — a number that grows with
|
||||
the campaign is a scan
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
_HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(_HERE.parent / "tests"))
|
||||
|
||||
_DB = tempfile.NamedTemporaryFile(suffix="-m10-cost.db", delete=False)
|
||||
_DB.close()
|
||||
os.environ["AIDND_DB_PATH"] = _DB.name
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends # noqa: E402
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
|
||||
import m10_fixture # noqa: E402
|
||||
from app import auth, limits, memorybank, models # noqa: E402
|
||||
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
|
||||
from app.main import app # noqa: E402
|
||||
from app.routers import adventures # noqa: E402
|
||||
from fakes import ScriptedProvider, state_block # noqa: E402
|
||||
from tools import dbmeter # noqa: E402
|
||||
|
||||
PROSE = (
|
||||
"Roger pulled the whiteboard marker apart while he talked, which was how "
|
||||
"everyone knew the meeting had stopped being about the agenda. Alice wrote "
|
||||
"nothing down. Outside the glass, somebody wheeled a trolley of monitors "
|
||||
"past the door and did not look in."
|
||||
)
|
||||
|
||||
|
||||
class _Stub:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "The meeting went on for some time."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25, 0.125] for _ in texts]
|
||||
|
||||
|
||||
def _setup() -> tuple[TestClient, int]:
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: _Stub()
|
||||
memorybank.summary_provider = lambda s: _Stub()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
Base.metadata.create_all(bind=engine)
|
||||
with SessionLocal() as db:
|
||||
user = models.User(is_guest=False, email="m10cost@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="cost-model", embedding_model="stub",
|
||||
context_token_budget=8192, max_output_tokens=600,
|
||||
))
|
||||
adventure = models.Adventure(user_id=user.id, title="Cost",
|
||||
auto_summarize=True, memory_bank_enabled=True)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(adventure_id=adventure.id, type="start",
|
||||
text="Bill badges in on a Tuesday morning."))
|
||||
db.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
return TestClient(app), adv_id
|
||||
|
||||
|
||||
def _db_bytes() -> int:
|
||||
return Path(_DB.name).stat().st_size
|
||||
|
||||
|
||||
def _counts(adv_id: int) -> dict:
|
||||
with SessionLocal() as db:
|
||||
return {
|
||||
"actions": db.query(models.Action).filter(
|
||||
models.Action.adventure_id == adv_id).count(),
|
||||
"visual profiles": db.query(models.VisualProfile).filter(
|
||||
models.VisualProfile.adventure_id == adv_id).count(),
|
||||
"M5 per-position state snapshots": db.query(models.Action).filter(
|
||||
models.Action.adventure_id == adv_id,
|
||||
models.Action.narrative_state_after.isnot(None)).count(),
|
||||
}
|
||||
|
||||
|
||||
def _profile_bytes(adv_id: int) -> int:
|
||||
with SessionLocal() as db:
|
||||
rows = db.query(models.VisualProfile).filter(
|
||||
models.VisualProfile.adventure_id == adv_id).all()
|
||||
return sum(
|
||||
len(json.dumps({"entity_key": r.entity_key,
|
||||
"descriptors": r.descriptors,
|
||||
"features": r.features,
|
||||
"style_notes": r.style_notes}).encode("utf-8"))
|
||||
for r in rows
|
||||
)
|
||||
|
||||
|
||||
def _packet_statements(client, adv_id: int, meter: dbmeter.Meter, label: str):
|
||||
with meter.scope(label) as scope:
|
||||
started = time.perf_counter()
|
||||
response = client.get(f"/api/adventures/{adv_id}/scene-packet")
|
||||
seconds = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
return scope, seconds
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--turns", type=int, default=60)
|
||||
args = parser.parse_args()
|
||||
|
||||
client, adv_id = _setup()
|
||||
empty_bytes = _db_bytes()
|
||||
m10_fixture.build(client, adv_id)
|
||||
|
||||
meter = dbmeter.Meter()
|
||||
meter.attach(engine)
|
||||
try:
|
||||
early_scope, early_seconds = _packet_statements(
|
||||
client, adv_id, meter, "packet at 2 turns")
|
||||
|
||||
before_play = _db_bytes()
|
||||
for turn in range(1, args.turns + 1):
|
||||
ScriptedProvider.replies = [
|
||||
f"{PROSE} [{turn}]\n" + state_block([
|
||||
{"type": "set_scene",
|
||||
"summary": f"The meeting reaches item {turn}.",
|
||||
"location": "office",
|
||||
"present": ["bill", "alice", "roger"]},
|
||||
])
|
||||
]
|
||||
client.post(f"/api/adventures/{adv_id}/actions",
|
||||
json={"type": "do", "text": f"item {turn}"}
|
||||
).raise_for_status()
|
||||
|
||||
rows_before = _counts(adv_id)
|
||||
late_scope, late_seconds = _packet_statements(
|
||||
client, adv_id, meter, f"packet at {args.turns + 2} turns")
|
||||
rows_after = _counts(adv_id)
|
||||
finally:
|
||||
meter.detach()
|
||||
|
||||
played_bytes = _db_bytes()
|
||||
profile_bytes = _profile_bytes(adv_id)
|
||||
|
||||
print(f"\n{args.turns} turns, {rows_before['actions']} action rows\n")
|
||||
|
||||
print("scene records M10 wrote")
|
||||
print(f" visual_profiles rows {rows_before['visual profiles']:>8}"
|
||||
" (one per profiled entity, written once)")
|
||||
print(" scene rows 0"
|
||||
" M10 adds no scenes table")
|
||||
print(f" M5 per-position state snapshots "
|
||||
f"{rows_before['M5 per-position state snapshots']:>8}"
|
||||
" already there since M5; the scene lives here")
|
||||
|
||||
print("\nbytes added to the database")
|
||||
print(f" empty database {empty_bytes:>8} B")
|
||||
print(f" after the fixture campaign {before_play:>8} B")
|
||||
print(f" after {args.turns} more turns".ljust(36)
|
||||
+ f"{played_bytes:>8} B")
|
||||
print(f" visual profile content {profile_bytes:>8} B"
|
||||
f" {100 * profile_bytes / max(played_bytes, 1):.3f}% of the database")
|
||||
|
||||
per_position = profile_bytes * rows_before["M5 per-position state snapshots"]
|
||||
print("\nprofile duplication: campaign-scoped against per-position")
|
||||
print(f" as stored, once per entity {profile_bytes:>8} B")
|
||||
print(f" if snapshotted per position {per_position:>8} B"
|
||||
f" x{per_position / max(profile_bytes, 1):.0f}")
|
||||
|
||||
print("\npacket: persisted or constructed")
|
||||
print(f" rows written while building one "
|
||||
f"{rows_after['visual profiles'] - rows_before['visual profiles']:>8}")
|
||||
print(" packet rows in any table 0 built on read, never stored")
|
||||
print(f" build time, 2 turns {early_seconds * 1000:>8.1f} ms")
|
||||
print(f" build time, {args.turns + 2} turns".ljust(36)
|
||||
+ f"{late_seconds * 1000:>8.1f} ms")
|
||||
|
||||
print("\ncurrent-scene query behaviour")
|
||||
print(f" statements, 2 turns {early_scope.total.statements:>8}")
|
||||
print(f" statements, {args.turns + 2} turns".ljust(36)
|
||||
+ f"{late_scope.total.statements:>8}")
|
||||
verdict = ("does not grow with the campaign"
|
||||
if late_scope.total.statements <= early_scope.total.statements
|
||||
else "GROWS — the scene is being scanned, not read")
|
||||
print(f" {verdict}")
|
||||
print("\n" + dbmeter.render_scope(late_scope, statements=6))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,331 @@
|
||||
"""M9's migration claim, proved against a database M8's own code wrote.
|
||||
|
||||
# from the M8 worktree, using M8's interpreter:
|
||||
python -m tools.m9_migration_proof build <db_path>
|
||||
|
||||
# from the M9 tree, using M9's interpreter:
|
||||
python -m tools.m9_migration_proof open <db_path>
|
||||
|
||||
M9 claims to add no schema change. `git diff` proves that nothing in
|
||||
`migrations.py` or `models.py` moved, which is necessary and not sufficient: a
|
||||
migration can also be *missing*, and the failure then is a database that opens
|
||||
and quietly answers wrongly. The M8 report set the standard here — a database
|
||||
created by today's code and read by today's code proves nothing — so the
|
||||
campaign below is built by a server running the signed M8 commit, from a git
|
||||
worktree, and read back by M9.
|
||||
|
||||
`build` writes a campaign that touches every family M9 changed the handling of:
|
||||
story with an alternate take, a Save Point, a manual state correction, memories
|
||||
and a summary, imported knowledge including a disabled and a narrator-only
|
||||
source, and per-turn context snapshots. It prints what it wrote, as JSON.
|
||||
|
||||
`open` opens that file with the current code, runs the migration path, and
|
||||
checks every one of those against what `build` reported. It also asserts the
|
||||
schema version did not move and that a second open is a no-op, which is what
|
||||
"no migration" means in practice: the stamp is the same number before and after.
|
||||
|
||||
Neither half imports anything from the other. What crosses is the database file
|
||||
and one JSON report on stdout, which is the only way the two builds can be made
|
||||
to talk without one of them importing the other's code.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "tests"))
|
||||
|
||||
|
||||
def _app(db_path: str):
|
||||
"""Imports the application against `db_path`. Must run before any app import."""
|
||||
os.environ["AIDND_DB_PATH"] = db_path
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from app import auth, limits, memorybank, models
|
||||
from app.database import Base, SessionLocal, engine, get_db
|
||||
from app.main import app
|
||||
from app.routers import adventures
|
||||
from fakes import ScriptedProvider, state_block
|
||||
|
||||
class Stub:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "A memory of what had happened by then."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25] for _ in texts]
|
||||
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: Stub()
|
||||
memorybank.summary_provider = lambda s: Stub()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
return {
|
||||
"Base": Base, "SessionLocal": SessionLocal, "engine": engine,
|
||||
"models": models, "app": app, "auth": auth, "get_db": get_db,
|
||||
"Depends": Depends, "TestClient": TestClient,
|
||||
"ScriptedProvider": ScriptedProvider, "state_block": state_block,
|
||||
"memorybank": memorybank,
|
||||
}
|
||||
|
||||
|
||||
def _client(ctx, user_id: int):
|
||||
ctx["app"].dependency_overrides[ctx["auth"].get_current_user] = (
|
||||
lambda db=ctx["Depends"](ctx["get_db"]): db.get(ctx["models"].User, user_id)
|
||||
)
|
||||
return ctx["TestClient"](ctx["app"])
|
||||
|
||||
|
||||
# ------------------------------------------------------------------- building
|
||||
|
||||
def build(db_path: str) -> dict:
|
||||
ctx = _app(db_path)
|
||||
ctx["Base"].metadata.create_all(bind=ctx["engine"])
|
||||
models, SessionLocal = ctx["models"], ctx["SessionLocal"]
|
||||
|
||||
db = SessionLocal()
|
||||
try:
|
||||
user = models.User(is_guest=False, email="m9mig@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="m8-model", embedding_model="stub",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Built by M8", auto_summarize=True,
|
||||
memory_bank_enabled=True,
|
||||
campaign_canon={"rules": ["The dead do not return."]},
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=adventure.id, type="start",
|
||||
text="Aldric sits in the Crooked Lantern with Mara.",
|
||||
))
|
||||
db.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
finally:
|
||||
db.close()
|
||||
|
||||
client = _client(ctx, user_id)
|
||||
upload = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": ("canon.md", (
|
||||
"# Westhaven\n\n## The Old Abbey\n\nThe abbey crypt is sealed.\n"
|
||||
).encode(), "text/markdown")},
|
||||
data={"classification": "canon", "always_include": "true"},
|
||||
)
|
||||
assert upload.status_code == 201, upload.text[:300]
|
||||
hidden = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": ("secret.md", (
|
||||
"# The seal\n\nIt was broken once, sixty years ago.\n"
|
||||
).encode(), "text/markdown")},
|
||||
data={"classification": "canon", "visibility": "hidden"},
|
||||
)
|
||||
assert hidden.status_code == 201, hidden.text[:300]
|
||||
disabled = client.post(
|
||||
f"/api/adventures/{adv_id}/knowledge",
|
||||
files={"file": ("draft.md", b"# Draft\n\nAn earlier version.\n",
|
||||
"text/markdown")},
|
||||
data={"classification": "reference"},
|
||||
)
|
||||
assert disabled.status_code == 201, disabled.text[:300]
|
||||
assert client.patch(
|
||||
f"/api/adventures/{adv_id}/knowledge/{disabled.json()['id']}",
|
||||
json={"enabled": False},
|
||||
).status_code == 200
|
||||
|
||||
state_block = ctx["state_block"]
|
||||
for turn in range(1, 8):
|
||||
ctx["ScriptedProvider"].replies = [
|
||||
f"The rain keeps on, and Mara says nothing for a while. [{turn}]\n"
|
||||
+ state_block([{"type": "add_fact", "predicate": "tally",
|
||||
"value": turn * 10, "fact_id": f"tally-{turn * 10}"}])
|
||||
]
|
||||
response = client.post(f"/api/adventures/{adv_id}/actions",
|
||||
json={"type": "do", "text": f"ask about turn {turn}"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
if turn == 3:
|
||||
assert client.post(f"/api/adventures/{adv_id}/retry").status_code == 200
|
||||
point = client.post(f"/api/adventures/{adv_id}/checkpoints",
|
||||
json={"name": "Third turn", "note": "A position."})
|
||||
assert point.status_code == 201, point.text[:300]
|
||||
|
||||
correction = client.post(f"/api/adventures/{adv_id}/state/corrections", json={
|
||||
"events": [{"type": "add_fact", "predicate": "keeper", "value": "Mara",
|
||||
"fact_id": "keeper"}],
|
||||
"note": "Established in play.",
|
||||
})
|
||||
assert correction.status_code == 201, correction.text[:300]
|
||||
|
||||
import asyncio
|
||||
asyncio.run(ctx["memorybank"].run_post_turn(adv_id))
|
||||
|
||||
# Two Undos, so the head is behind the retained tip when M9 opens it.
|
||||
for _ in range(2):
|
||||
assert client.post(f"/api/adventures/{adv_id}/undo").status_code == 200
|
||||
|
||||
report = _describe(ctx, client, adv_id)
|
||||
ctx["app"].dependency_overrides.clear()
|
||||
return report
|
||||
|
||||
|
||||
# -------------------------------------------------------------------- reading
|
||||
|
||||
def _describe(ctx, client, adv_id: int) -> dict:
|
||||
"""Everything the other build has to agree with, read through the API."""
|
||||
models, SessionLocal = ctx["models"], ctx["SessionLocal"]
|
||||
page = client.get(f"/api/adventures/{adv_id}").json()
|
||||
with SessionLocal() as db:
|
||||
adventure = db.get(models.Adventure, adv_id)
|
||||
version = db.execute(_pragma()).scalar()
|
||||
counts = {
|
||||
name: db.query(model).filter(model.adventure_id == adv_id).count()
|
||||
for name, model in (
|
||||
("actions", models.Action), ("memories", models.Memory),
|
||||
("summaries", models.Summary), ("checkpoints", models.Checkpoint),
|
||||
("state_events", models.StateEvent),
|
||||
("state_proposals", models.StateProposal),
|
||||
("knowledge_sources", models.KnowledgeSource),
|
||||
("knowledge_chunks", models.KnowledgeChunk),
|
||||
)
|
||||
}
|
||||
head = {"branch_id": adventure.head_branch_id, "depth": adventure.head_depth}
|
||||
snapshots = db.query(models.Action).filter(
|
||||
models.Action.adventure_id == adv_id,
|
||||
models.Action.context_snapshot.isnot(None),
|
||||
).count()
|
||||
return {
|
||||
"adventure_id": adv_id,
|
||||
"schema_version": version,
|
||||
"title": page["title"],
|
||||
"canon_rules": page["canon_rules"],
|
||||
"transcript": [a["text"] for a in page["actions"]],
|
||||
"can_undo": page["can_undo"],
|
||||
"can_redo": page["can_redo"],
|
||||
"head": head,
|
||||
"counts": counts,
|
||||
"snapshot_rows": snapshots,
|
||||
"state": client.get(f"/api/adventures/{adv_id}/state").json()["document"],
|
||||
"checkpoints": sorted(
|
||||
(c["name"], c["depth"])
|
||||
for c in client.get(f"/api/adventures/{adv_id}/checkpoints").json()
|
||||
),
|
||||
"knowledge": sorted(
|
||||
(k["original_filename"], k["classification"], k["enabled"],
|
||||
k["visibility"], k["always_include"], k["content_hash"],
|
||||
k["index_state"])
|
||||
for k in client.get(f"/api/adventures/{adv_id}/knowledge").json()
|
||||
),
|
||||
"events": sorted(
|
||||
(e["event_type"], e["source"], json.dumps(e["payload"], sort_keys=True))
|
||||
for e in client.get(
|
||||
f"/api/adventures/{adv_id}/state/events?limit=500").json()
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def _comparable(value):
|
||||
"""`value` as it survives a JSON round trip, so the two builds compare like."""
|
||||
return json.loads(json.dumps(value, sort_keys=True, default=str))
|
||||
|
||||
|
||||
def _pragma():
|
||||
from sqlalchemy import text
|
||||
|
||||
return text("PRAGMA user_version")
|
||||
|
||||
|
||||
def open_it(db_path: str, expected: dict) -> dict:
|
||||
"""Opens an existing database with this build, and checks it against `expected`."""
|
||||
ctx = _app(db_path)
|
||||
from app import migrations
|
||||
|
||||
# This is the migration run. `main` already called `bootstrap` at import.
|
||||
with ctx["engine"].begin() as conn:
|
||||
after_first = conn.execute(_pragma()).scalar()
|
||||
# And again, to prove idempotence: a second run must change nothing.
|
||||
migrations.bootstrap(ctx["engine"])
|
||||
with ctx["engine"].begin() as conn:
|
||||
after_second = conn.execute(_pragma()).scalar()
|
||||
|
||||
adv_id = expected["adventure_id"]
|
||||
with ctx["SessionLocal"]() as db:
|
||||
user = db.query(ctx["models"].User).first()
|
||||
user_id = user.id
|
||||
client = _client(ctx, user_id)
|
||||
actual = _describe(ctx, client, adv_id)
|
||||
|
||||
problems = []
|
||||
for key in ("title", "canon_rules", "transcript", "head", "counts",
|
||||
"snapshot_rows", "state", "checkpoints", "knowledge", "events",
|
||||
"can_undo", "can_redo"):
|
||||
# Compared through JSON, because that is how the other build's answer
|
||||
# arrived: a tuple written by `_describe` comes back as a list, and a
|
||||
# comparison that called that a difference would report ten differences
|
||||
# in a database nothing had changed.
|
||||
if _comparable(actual[key]) != _comparable(expected[key]):
|
||||
problems.append(f"{key}: expected {expected[key]!r}, got {actual[key]!r}")
|
||||
if expected["schema_version"] != after_first:
|
||||
problems.append(
|
||||
f"the schema version moved: {expected['schema_version']} -> {after_first}"
|
||||
)
|
||||
if after_first != after_second:
|
||||
problems.append(
|
||||
f"a second open migrated again: {after_first} -> {after_second}"
|
||||
)
|
||||
|
||||
# And the campaign still works, rather than merely reading correctly.
|
||||
exported = client.get(f"/api/adventures/{adv_id}/export")
|
||||
if exported.status_code != 200:
|
||||
problems.append(f"export failed: {exported.status_code}")
|
||||
else:
|
||||
imported = client.post("/api/adventures/import", json=exported.json())
|
||||
if imported.status_code != 201:
|
||||
problems.append(f"round trip failed: {imported.text[:300]}")
|
||||
elif exported.json()["format"] != "ai-dnd-adventure-v3":
|
||||
problems.append("the M8 database did not export as v3")
|
||||
redo = client.post(f"/api/adventures/{adv_id}/redo")
|
||||
if redo.status_code != 200:
|
||||
problems.append(f"Redo failed on the migrated campaign: {redo.status_code}")
|
||||
|
||||
ctx["app"].dependency_overrides.clear()
|
||||
return {
|
||||
"schema_version_before": expected["schema_version"],
|
||||
"schema_version_after": after_first,
|
||||
"schema_version_second_open": after_second,
|
||||
"problems": problems,
|
||||
"checked": {
|
||||
"families": 12, "snapshot_rows": actual["snapshot_rows"],
|
||||
"counts": actual["counts"],
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("mode", choices=("build", "open"))
|
||||
parser.add_argument("db_path")
|
||||
parser.add_argument("--expected", help="the JSON `build` printed (open only)")
|
||||
args = parser.parse_args()
|
||||
|
||||
if args.mode == "build":
|
||||
print(json.dumps(build(args.db_path), sort_keys=True))
|
||||
return 0
|
||||
|
||||
expected = json.loads(Path(args.expected).read_text())
|
||||
result = open_it(args.db_path, expected)
|
||||
print(json.dumps(result, indent=2, sort_keys=True))
|
||||
return 1 if result["problems"] else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,354 @@
|
||||
"""Measures what a campaign bundle preserves, omits and rebuilds.
|
||||
|
||||
python -m tools.m9_portability_report # human-readable
|
||||
python -m tools.m9_portability_report --json # machine-readable
|
||||
|
||||
Run from `backend/`, with the virtualenv on the path. The script builds the M9
|
||||
portability fixture in a throwaway database, exports it, imports it into a second
|
||||
throwaway database, and then compares the two campaigns family by family.
|
||||
|
||||
It exists because the M9 brief asks for the baseline to be **measured** rather
|
||||
than assumed. Running it on the M8 commit produces the inventory M9 started from;
|
||||
running it on the M9 tree produces the one M9 finished with, and the difference
|
||||
between the two files is the milestone's portability claim in a form a reviewer
|
||||
can reproduce rather than take on trust.
|
||||
|
||||
The comparison is by data family rather than by row count. "12 actions in, 12
|
||||
actions out" is the check that misses a bundle carrying every turn and none of
|
||||
its state, so each family below reports what a reader could still see afterwards.
|
||||
|
||||
Nothing here touches the developer's own database: two temporary files are
|
||||
created and removed, and no network call is made — the narrator, the summariser
|
||||
and the embedder are all local fakes.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
# The test harness owns the fixture and the fakes. Both live under `tests/`,
|
||||
# which is not a package, so the path is extended rather than imported from.
|
||||
_HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(_HERE.parent / "tests"))
|
||||
|
||||
# `app.database` reads this at import and builds the engine once, exactly as
|
||||
# `tests/conftest.py` explains. It has to be set before the first `app` import.
|
||||
_SOURCE_DB = tempfile.NamedTemporaryFile(suffix="-m9-source.db", delete=False)
|
||||
_SOURCE_DB.close()
|
||||
os.environ["AIDND_DB_PATH"] = _SOURCE_DB.name
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends # noqa: E402
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
|
||||
import m9_fixture # noqa: E402
|
||||
from app import auth, limits, memorybank, models # noqa: E402
|
||||
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
|
||||
from app.main import app # noqa: E402
|
||||
from app.routers import adventures # noqa: E402
|
||||
from fakes import ScriptedProvider # noqa: E402
|
||||
|
||||
|
||||
class _StubDerivedProvider:
|
||||
"""Deterministic vectors and prose, so the report needs no model at all.
|
||||
|
||||
One object serves as both the embedder and the summariser, because the
|
||||
memory pass builds each from the same factory and stubbing only one of them
|
||||
is the M6 finding M6-F3 mistake: the unstubbed factory opens a socket
|
||||
against the default endpoint on every turn.
|
||||
"""
|
||||
|
||||
_written = 0
|
||||
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
_StubDerivedProvider._written += 1
|
||||
return (
|
||||
f"Memory {_StubDerivedProvider._written}: what the story had "
|
||||
f"established by this point."
|
||||
)
|
||||
|
||||
async def embed(self, texts):
|
||||
out = []
|
||||
for text in texts:
|
||||
lowered = text.lower()
|
||||
out.append([
|
||||
1.0,
|
||||
1.0 if "abbey" in lowered or "crypt" in lowered else 0.0,
|
||||
1.0 if "tavern" in lowered or "lantern" in lowered else 0.0,
|
||||
1.0 if "rain" in lowered else 0.0,
|
||||
])
|
||||
return out
|
||||
|
||||
|
||||
def _install_fakes() -> None:
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: _StubDerivedProvider()
|
||||
memorybank.summary_provider = lambda s: _StubDerivedProvider()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
|
||||
|
||||
def _new_user_and_campaign(title: str) -> tuple[int, int]:
|
||||
db = SessionLocal()
|
||||
try:
|
||||
user = models.User(is_guest=False, email=f"m9-{title}@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="report-model", embedding_model="stub-embed",
|
||||
context_token_budget=4000, max_output_tokens=400,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title=title,
|
||||
campaign_canon=m9_fixture.CAMPAIGN_CANON,
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
|
||||
))
|
||||
# A neighbour, so a bundle that reached past its own campaign would
|
||||
# bring back rows this report can see.
|
||||
neighbour = models.Adventure(user_id=user.id, title="Neighbour")
|
||||
db.add(neighbour)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=neighbour.id, type="start", text="A different story.",
|
||||
))
|
||||
db.commit()
|
||||
return adventure.id, user.id
|
||||
finally:
|
||||
db.close()
|
||||
|
||||
|
||||
def _client(user_id: int) -> TestClient:
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
return TestClient(app)
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- the families
|
||||
# One entry per data family the M9 brief asks the baseline to classify. Each
|
||||
# `present` function answers "did this survive into the file?" from the bundle
|
||||
# alone, because that is the question the classification is about.
|
||||
|
||||
def _actions(bundle: dict) -> list[dict]:
|
||||
return [a for a in (bundle.get("actions") or []) if isinstance(a, dict)]
|
||||
|
||||
|
||||
def _snapshots(bundle: dict) -> list[dict]:
|
||||
"""Every stored prompt in the file, decoded.
|
||||
|
||||
The export compresses them (`bundle._packed`), so a report that looked for a
|
||||
plain dict would say the evidence was omitted when it is merely encoded —
|
||||
which is the mistake this whole tool exists to avoid making about anything.
|
||||
"""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
out = []
|
||||
for action in _actions(bundle):
|
||||
snapshot = (
|
||||
action.get("contextSnapshot")
|
||||
if isinstance(action.get("contextSnapshot"), dict)
|
||||
else bundle_module._unpacked(action.get("contextSnapshotZ"))
|
||||
)
|
||||
if isinstance(snapshot, dict):
|
||||
out.append(snapshot)
|
||||
return out
|
||||
|
||||
|
||||
FAMILIES: list[tuple[str, str, callable]] = [
|
||||
("campaign identity",
|
||||
"title, instructions, persona, canon, the campaign's own settings",
|
||||
lambda b: bool(b.get("title"))),
|
||||
("transcript",
|
||||
"every accepted player and narrator action, live and superseded",
|
||||
lambda b: bool(_actions(b))),
|
||||
("branches",
|
||||
"the retained tree, its fork points and its names",
|
||||
lambda b: bool(b.get("branches"))),
|
||||
("branch disposition",
|
||||
"which lines the story left behind, and where",
|
||||
lambda b: any("supersededAt" in x for x in (b.get("branches") or []))),
|
||||
("active head",
|
||||
"the branch and depth the campaign is being read at",
|
||||
lambda b: b.get("headDepth") is not None),
|
||||
("alternate takes",
|
||||
"every attempt at a turn, and which one is the story",
|
||||
lambda b: any(not a.get("live", True) for a in _actions(b))),
|
||||
("take grouping",
|
||||
"which attempts belong to the same turn across a fork (SP9 parentage)",
|
||||
lambda b: any("parentId" in a for a in _actions(b))),
|
||||
("save points",
|
||||
"named coordinates, their notes and their positions",
|
||||
lambda b: bool(b.get("checkpoints"))),
|
||||
("narrative state (current)",
|
||||
"the authoritative document at the exported head",
|
||||
lambda b: b.get("narrativeState") is not None),
|
||||
("narrative state (per position)",
|
||||
"the snapshot every position restores from",
|
||||
lambda b: any("narrativeStateAfter" in a for a in _actions(b))),
|
||||
("state events",
|
||||
"the accepted typed events: the audit half of the hybrid",
|
||||
lambda b: bool(b.get("stateEvents"))),
|
||||
("state proposals",
|
||||
"what the model proposed and what the application did about it",
|
||||
lambda b: bool(b.get("stateProposals"))),
|
||||
("manual corrections",
|
||||
"state the user asserted, distinguishable from state the story did",
|
||||
lambda b: any(e.get("source") == "manual_correction"
|
||||
for e in (b.get("stateEvents") or []))),
|
||||
("historical prompt/context",
|
||||
"the exact prompt each turn was given",
|
||||
lambda b: bool(_snapshots(b))),
|
||||
("retrieval provenance",
|
||||
"which passages a historical turn was shown, and their text",
|
||||
lambda b: any((s.get("knowledge") or {}).get("used") for s in _snapshots(b))),
|
||||
("per-turn model settings",
|
||||
"the model and generation settings a historical turn ran under",
|
||||
lambda b: any(s.get("settings") for s in _snapshots(b))),
|
||||
("imported knowledge",
|
||||
"source content, class, lifecycle, visibility and hash",
|
||||
lambda b: bool(b.get("knowledge"))),
|
||||
("knowledge parser versions",
|
||||
"what produced the chunks the source last had",
|
||||
lambda b: any("parserVersion" in k for k in (b.get("knowledge") or []))),
|
||||
("summaries",
|
||||
"the generated rolling summaries and the story they cover",
|
||||
lambda b: bool(b.get("summaries"))),
|
||||
("memories",
|
||||
"long-term memories and the coordinate each hangs off",
|
||||
lambda b: bool(b.get("memories"))),
|
||||
("memory authority",
|
||||
"whether a memory is accepted story or a heuristic reading of it",
|
||||
lambda b: any("authority" in m for m in (b.get("memories") or []))),
|
||||
("scene metadata",
|
||||
"the scene section of the authoritative state document",
|
||||
lambda b: isinstance(b.get("narrativeState"), dict)
|
||||
and "scene" in b["narrativeState"]),
|
||||
("story cards (legacy)",
|
||||
"the inherited lore primitive, which has no v1 browser surface",
|
||||
lambda b: "storyCards" in b),
|
||||
]
|
||||
|
||||
#: Families that are deliberately rebuilt rather than carried, with the reason.
|
||||
REBUILDABLE = {
|
||||
"knowledge passages": "a deterministic function of the source content",
|
||||
"lexical (FTS) index": "rebuilt from the passages on import",
|
||||
"knowledge embeddings": "belong to the importing machine's embedding model",
|
||||
"memory embeddings": "the same, for the memory bank",
|
||||
"branch lineage cache": "computed from parent plus fork depth",
|
||||
"derived status": "describes the last run of a background pass, not the story",
|
||||
}
|
||||
|
||||
|
||||
def measure(json_out: bool) -> dict:
|
||||
_install_fakes()
|
||||
Base.metadata.create_all(bind=engine)
|
||||
adv_id, user_id = _new_user_and_campaign("M9 Portability Fixture")
|
||||
client = _client(user_id)
|
||||
|
||||
built = time.perf_counter()
|
||||
source = m9_fixture.build(client, adv_id)
|
||||
build_seconds = time.perf_counter() - built
|
||||
|
||||
started = time.perf_counter()
|
||||
response = client.get(f"/api/adventures/{adv_id}/export")
|
||||
export_seconds = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
bundle = response.json()
|
||||
encoded = json.dumps(bundle, ensure_ascii=False).encode("utf-8")
|
||||
|
||||
started = time.perf_counter()
|
||||
imported = client.post("/api/adventures/import", json=bundle)
|
||||
import_seconds = time.perf_counter() - started
|
||||
import_status = imported.status_code
|
||||
copy = (
|
||||
m9_fixture.snapshot_of(client, imported.json()["id"])
|
||||
if import_status == 201 else None
|
||||
)
|
||||
|
||||
source_db_bytes = Path(_SOURCE_DB.name).stat().st_size
|
||||
app.dependency_overrides.clear()
|
||||
|
||||
families = [
|
||||
{"family": name, "what": what,
|
||||
"verdict": "PRESERVED" if present(bundle) else "OMITTED"}
|
||||
for name, what, present in FAMILIES
|
||||
]
|
||||
report = {
|
||||
"format": bundle.get("format"),
|
||||
"families": families,
|
||||
"rebuildable": REBUILDABLE,
|
||||
"sizes": {
|
||||
"source_database_bytes": source_db_bytes,
|
||||
"bundle_bytes": len(encoded),
|
||||
"bundle_actions": len(_actions(bundle)),
|
||||
"bundle_keys": sorted(bundle),
|
||||
},
|
||||
"timings_seconds": {
|
||||
"fixture_build": round(build_seconds, 3),
|
||||
"export": round(export_seconds, 3),
|
||||
"import": round(import_seconds, 3),
|
||||
},
|
||||
"round_trip": {
|
||||
"import_status": import_status,
|
||||
"agrees": _agreement(source, copy) if copy else None,
|
||||
},
|
||||
}
|
||||
return report
|
||||
|
||||
|
||||
def _agreement(source: dict, copy: dict) -> dict:
|
||||
"""Which of the reader-visible families match between original and copy."""
|
||||
keys = ("title", "canon_rules", "transcript", "branch_count", "checkpoints",
|
||||
"knowledge", "state", "state_events", "memories", "summaries",
|
||||
"can_undo", "can_redo")
|
||||
return {key: source.get(key) == copy.get(key) for key in keys}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--json", action="store_true",
|
||||
help="print the report as JSON")
|
||||
args = parser.parse_args()
|
||||
try:
|
||||
report = measure(args.json)
|
||||
finally:
|
||||
for path in (_SOURCE_DB.name,):
|
||||
try:
|
||||
os.unlink(path)
|
||||
except OSError:
|
||||
pass
|
||||
if args.json:
|
||||
print(json.dumps(report, indent=2, sort_keys=True))
|
||||
return 0
|
||||
print(f"bundle format: {report['format']}")
|
||||
print(f"bundle size: {report['sizes']['bundle_bytes']:,} bytes "
|
||||
f"across {report['sizes']['bundle_actions']} actions")
|
||||
print(f"source db: {report['sizes']['source_database_bytes']:,} bytes")
|
||||
print(f"timings: {report['timings_seconds']}")
|
||||
print()
|
||||
width = max(len(name) for name, _, _ in FAMILIES)
|
||||
for row in report["families"]:
|
||||
print(f" {row['verdict']:<10} {row['family']:<{width}} {row['what']}")
|
||||
print()
|
||||
print(" DERIVED/REBUILDABLE (deliberately not carried)")
|
||||
for name, why in REBUILDABLE.items():
|
||||
print(f" {name:<24} {why}")
|
||||
print()
|
||||
print(f"round trip: HTTP {report['round_trip']['import_status']}")
|
||||
for key, agreed in (report["round_trip"]["agrees"] or {}).items():
|
||||
print(f" {'same' if agreed else 'DIFFERS':<8} {key}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,264 @@
|
||||
"""How a campaign bundle grows with the campaign, measured rather than reasoned.
|
||||
|
||||
python -m tools.m9_scale_report [--turns 120] [--budget 16384]
|
||||
|
||||
Run from `backend/`. Plays a campaign of `--turns` turns against a scripted
|
||||
narrator with the real prompt builder and a realistic context budget, then
|
||||
exports it and reports where the bytes are.
|
||||
|
||||
The question it exists to answer is the one M9's decision to carry historical
|
||||
prompts raises: **a per-turn prompt contains the story so far, so storing one per
|
||||
turn is quadratic in campaign length.** That is already true of the database —
|
||||
`compression.py` records the column as 89% of production storage — and M9 makes
|
||||
it true of the export as well. Reasoning about it gives the wrong number, because
|
||||
the prompt is bounded by the context budget rather than by the transcript: once
|
||||
the history window is full, each turn's snapshot stops growing and the total
|
||||
becomes linear again. Where that knee falls is a measurement.
|
||||
|
||||
It also watches for the accidental costs §26 names: a query per row, a
|
||||
duplicated body of knowledge content, or a snapshot written more than once per
|
||||
turn.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
_HERE = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(_HERE.parent / "tests"))
|
||||
|
||||
_DB = tempfile.NamedTemporaryFile(suffix="-m9-scale.db", delete=False)
|
||||
_DB.close()
|
||||
os.environ["AIDND_DB_PATH"] = _DB.name
|
||||
os.environ.pop("AIDND_DATABASE_URL", None)
|
||||
os.environ.pop("DATABASE_URL", None)
|
||||
|
||||
from fastapi import Depends # noqa: E402
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
|
||||
import m9_fixture # noqa: E402
|
||||
from app import auth, limits, memorybank, models # noqa: E402
|
||||
from app.database import Base, SessionLocal, engine, get_db # noqa: E402
|
||||
from app.main import app # noqa: E402
|
||||
from app.routers import adventures # noqa: E402
|
||||
from fakes import ScriptedProvider, state_block # noqa: E402
|
||||
|
||||
#: Prose long enough that a turn is a turn rather than a sentence. The history
|
||||
#: window is what fills the prompt, so a fixture of three-word replies would
|
||||
#: measure a campaign nobody plays.
|
||||
PROSE = (
|
||||
"The rain came harder off the fen and the lantern light shivered on the wet "
|
||||
"boards. Mara set down the cloth she had been folding and looked at him for "
|
||||
"a while without saying anything, the way she did when the answer was going "
|
||||
"to cost her something. Outside, somebody crossed the yard and did not stop."
|
||||
)
|
||||
|
||||
|
||||
class _Stub:
|
||||
async def complete(self, system, prompt, **kwargs):
|
||||
return "The story had established a good deal by this point."
|
||||
|
||||
async def embed(self, texts):
|
||||
return [[1.0, 0.5, 0.25, 0.125] for _ in texts]
|
||||
|
||||
|
||||
def _setup(budget: int) -> tuple[TestClient, int]:
|
||||
adventures.turns.OpenAICompatibleProvider = ScriptedProvider
|
||||
memorybank.embedding_provider = lambda s: _Stub()
|
||||
memorybank.summary_provider = lambda s: _Stub()
|
||||
limits.check_row_cap = lambda *a, **k: None
|
||||
Base.metadata.create_all(bind=engine)
|
||||
db = SessionLocal()
|
||||
try:
|
||||
user = models.User(is_guest=False, email="scale@example.com")
|
||||
db.add(user)
|
||||
db.flush()
|
||||
db.add(models.Settings(
|
||||
user_id=user.id, model="scale-model", embedding_model="stub",
|
||||
context_token_budget=budget, max_output_tokens=800,
|
||||
))
|
||||
adventure = models.Adventure(
|
||||
user_id=user.id, title="Scale", auto_summarize=True,
|
||||
memory_bank_enabled=True,
|
||||
campaign_canon=m9_fixture.CAMPAIGN_CANON,
|
||||
)
|
||||
db.add(adventure)
|
||||
db.flush()
|
||||
db.add(models.Action(
|
||||
adventure_id=adventure.id, type="start", text=m9_fixture.OPENING,
|
||||
))
|
||||
db.commit()
|
||||
adv_id, user_id = adventure.id, user.id
|
||||
finally:
|
||||
db.close()
|
||||
app.dependency_overrides[auth.get_current_user] = (
|
||||
lambda db=Depends(get_db): db.get(models.User, user_id)
|
||||
)
|
||||
return TestClient(app), adv_id
|
||||
|
||||
|
||||
def _bundle_bytes(client, adv_id) -> tuple[int, dict, float]:
|
||||
started = time.perf_counter()
|
||||
response = client.get(f"/api/adventures/{adv_id}/export")
|
||||
seconds = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
payload = response.json()
|
||||
return len(json.dumps(payload).encode("utf-8")), payload, seconds
|
||||
|
||||
|
||||
def _snapshot_bytes(payload: dict) -> int:
|
||||
"""What the stored prompts cost **in the file**, which is the encoded size.
|
||||
|
||||
Measured as they appear rather than decoded first: the question this report
|
||||
answers is how large the file gets and how close it comes to the import
|
||||
ceiling, so what counts is the bytes that actually travel.
|
||||
"""
|
||||
return sum(
|
||||
len(json.dumps(action[key]).encode("utf-8"))
|
||||
for action in payload["actions"]
|
||||
for key in ("contextSnapshotZ", "contextSnapshot")
|
||||
if action.get(key)
|
||||
)
|
||||
|
||||
|
||||
def _decoded_snapshot_bytes(payload: dict) -> int:
|
||||
"""What the same prompts would cost uncompressed, for the ratio."""
|
||||
from app import bundle as bundle_module
|
||||
|
||||
total = 0
|
||||
for action in payload["actions"]:
|
||||
snapshot = (
|
||||
action.get("contextSnapshot")
|
||||
if isinstance(action.get("contextSnapshot"), dict)
|
||||
else bundle_module._unpacked(action.get("contextSnapshotZ"))
|
||||
)
|
||||
if isinstance(snapshot, dict):
|
||||
total += len(json.dumps(snapshot).encode("utf-8"))
|
||||
return total
|
||||
|
||||
|
||||
def _section_bytes(payload: dict) -> dict:
|
||||
"""What each v3 addition costs in the file, separately.
|
||||
|
||||
Needed because "the snapshots are 21% of the file" does not answer "what did
|
||||
M9 add": the state events, the proposals and the summaries are v3 additions
|
||||
too, and a claim about M9's cost that counted only the prompts would be
|
||||
understating it.
|
||||
"""
|
||||
def size(value) -> int:
|
||||
return len(json.dumps(value).encode("utf-8"))
|
||||
|
||||
per_node = {"contextSnapshotZ": 0, "id": 0, "parentId": 0}
|
||||
for action in payload["actions"]:
|
||||
for key in per_node:
|
||||
if key in action:
|
||||
per_node[key] += size(action[key]) + len(key) + 4
|
||||
return {
|
||||
"prompts (contextSnapshotZ)": per_node["contextSnapshotZ"],
|
||||
"state events": size(payload.get("stateEvents") or []),
|
||||
"state proposals": size(payload.get("stateProposals") or []),
|
||||
"summaries": size(payload.get("summaries") or []),
|
||||
"node ids + parentage": per_node["id"] + per_node["parentId"],
|
||||
"per-position state (v2 already)": sum(
|
||||
size(a["narrativeStateAfter"]) for a in payload["actions"]
|
||||
if "narrativeStateAfter" in a
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--turns", type=int, default=120)
|
||||
parser.add_argument("--budget", type=int, default=16384)
|
||||
parser.add_argument("--every", type=int, default=20,
|
||||
help="report the running size every N turns")
|
||||
args = parser.parse_args()
|
||||
|
||||
client, adv_id = _setup(args.budget)
|
||||
for name, body, kind in (
|
||||
("canon.md", m9_fixture.CANON_MD, "canon"),
|
||||
("reference.md", m9_fixture.REFERENCE_MD, "reference"),
|
||||
("secret.md", m9_fixture.SECRET_MD, "canon"),
|
||||
):
|
||||
m9_fixture.upload(client, adv_id, name, body, kind)
|
||||
|
||||
print(f"budget {args.budget} tokens, {args.turns} turns\n")
|
||||
print(f"{'turns':>6} {'actions':>8} {'bundle B':>12} {'snapshots B':>13} "
|
||||
f"{'B/turn':>9} {'export s':>9} {'import s':>9}")
|
||||
rows = []
|
||||
sections: dict[int, dict] = {}
|
||||
for turn in range(1, args.turns + 1):
|
||||
ScriptedProvider.replies = [
|
||||
f"{PROSE} [{turn}]\n"
|
||||
+ state_block([{"type": "add_fact", "predicate": "tally",
|
||||
"value": turn * 10, "fact_id": f"tally-{turn}"}])
|
||||
]
|
||||
response = client.post(f"/api/adventures/{adv_id}/actions",
|
||||
json={"type": "do", "text": f"press on, {turn}"})
|
||||
assert response.status_code == 200, response.text[:300]
|
||||
if turn % args.every and turn != args.turns:
|
||||
continue
|
||||
m9_fixture.settle_derived(adv_id)
|
||||
size, payload, export_seconds = _bundle_bytes(client, adv_id)
|
||||
started = time.perf_counter()
|
||||
imported = client.post("/api/adventures/import", json=payload)
|
||||
import_seconds = time.perf_counter() - started
|
||||
assert imported.status_code == 201, imported.text[:300]
|
||||
client.delete(f"/api/adventures/{imported.json()['id']}")
|
||||
snapshots = _snapshot_bytes(payload)
|
||||
plain = _decoded_snapshot_bytes(payload)
|
||||
rows.append((turn, size, snapshots, plain))
|
||||
sections[turn] = _section_bytes(payload)
|
||||
print(f"{turn:>6} {len(payload['actions']):>8} {size:>12,} "
|
||||
f"{snapshots:>13,} {snapshots // turn:>9,} "
|
||||
f"{export_seconds:>9.3f} {import_seconds:>9.3f}")
|
||||
|
||||
db_bytes = Path(_DB.name).stat().st_size
|
||||
last_turn, last_size, last_snapshots, last_plain = rows[-1]
|
||||
per_turn = last_snapshots // last_turn
|
||||
cap = limits.MAX_IMPORT_BODY_BYTES
|
||||
print()
|
||||
print(f"database on disk: {db_bytes:,} bytes")
|
||||
print(f"snapshot share of file: {100 * last_snapshots // last_size}%")
|
||||
print(f"stored uncompressed: {last_plain:,} bytes "
|
||||
f"({last_plain / max(last_snapshots, 1):.1f}x the encoded size)")
|
||||
print(f"import body cap: {cap:,} bytes")
|
||||
print(f"turns before the cap: ~{cap // max(per_turn, 1):,} "
|
||||
f"at the marginal rate above")
|
||||
# Growth between the last two samples says whether the per-turn cost has
|
||||
# settled. It should: once the history window fills the budget, a prompt
|
||||
# stops growing with the transcript and the total becomes linear.
|
||||
if len(rows) >= 2:
|
||||
(t0, _, s0, _p0), (t1, _, s1, _p1) = rows[-2], rows[-1]
|
||||
print(f"marginal cost, last {t1 - t0} turns: "
|
||||
f"{(s1 - s0) // max(t1 - t0, 1):,} bytes/turn")
|
||||
print()
|
||||
print("where the bytes are, at the last sample:")
|
||||
last = sections[last_turn]
|
||||
added = sum(v for k, v in last.items() if not k.endswith("(v2 already)"))
|
||||
for name, value in sorted(last.items(), key=lambda kv: -kv[1]):
|
||||
print(f" {name:34} {value:>12,} {100 * value / last_size:5.1f}%")
|
||||
print(f" {'--- everything v3 added':34} {added:>12,} "
|
||||
f"{100 * added / last_size:5.1f}%")
|
||||
without = last_size - added
|
||||
print(f" a v2 file of the same campaign {without:>12,}")
|
||||
print(f" ceiling with v3 additions: ~{int((cap / (last_size / last_turn**2)) ** 0.5):,} turns")
|
||||
print(f" ceiling without them: ~{int((cap / (without / last_turn**2)) ** 0.5):,} turns")
|
||||
app.dependency_overrides.clear()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
try:
|
||||
raise SystemExit(main())
|
||||
finally:
|
||||
try:
|
||||
os.unlink(_DB.name)
|
||||
except OSError:
|
||||
pass
|
||||
@@ -153,6 +153,14 @@ export const api = {
|
||||
retry: (advId, handlers, signal) => streamSSE(`/adventures/${advId}/retry`, {}, handlers, signal),
|
||||
exportAdventure: (id) => request(`/adventures/${id}/export`),
|
||||
importAdventure: (bundle) => request('/adventures/import', { method: 'POST', body: JSON.stringify(bundle) }),
|
||||
|
||||
// M9. A verified copy of the whole database, which is a different tool from
|
||||
// exporting one campaign: the export moves a campaign between installations,
|
||||
// and this is a safety copy of everything on this machine. Neither takes a
|
||||
// path — the server derives the destination from the database it already has
|
||||
// open, so there is nothing here for a caller to point somewhere else.
|
||||
listBackups: () => request('/backups'),
|
||||
createBackup: () => request('/backups', { method: 'POST' }),
|
||||
undo: (advId) => request(`/adventures/${advId}/undo`, { method: 'POST' }),
|
||||
// Undo moves the story back without deleting it, so there is somewhere to
|
||||
// move forward to again (M3). Both answer with the newest window.
|
||||
|
||||
@@ -15,17 +15,53 @@ export function pickJSONFile() {
|
||||
const input = document.createElement('input')
|
||||
input.type = 'file'
|
||||
input.accept = '.json,application/json'
|
||||
// M9: in the document, and driveable, rather than detached.
|
||||
//
|
||||
// A detached input is what this was, and `.click()` on one opens the
|
||||
// browser's file dialog in Firefox and Chrome today — but it is not
|
||||
// something the HTML spec requires, and it made the one control that
|
||||
// recovers a campaign impossible to drive from a browser test: there is no
|
||||
// element for WebDriver to hand a path to, so the import workflow could
|
||||
// only ever be checked by calling the API underneath it.
|
||||
//
|
||||
// Taken out of layout rather than marked `hidden`, and the difference is
|
||||
// load-bearing. A `hidden` input is non-interactable, and WebDriver will
|
||||
// set `files` on one without dispatching `change` — so the file lands and
|
||||
// nothing happens, which is a worse failure than the detached input was
|
||||
// because it looks like it worked. This is the ordinary visually-hidden
|
||||
// file-input pattern: off-screen, zero-sized, out of the accessibility
|
||||
// tree and out of the tab order, so no reader meets a stray "Choose file"
|
||||
// control while the browser's own dialog is what they are looking at.
|
||||
input.setAttribute('aria-hidden', 'true')
|
||||
input.tabIndex = -1
|
||||
input.style.cssText =
|
||||
'position:fixed;left:-9999px;width:1px;height:1px;opacity:0;pointer-events:none'
|
||||
input.dataset.testid = 'import-file'
|
||||
const done = (settle) => (value) => { input.remove(); settle(value) }
|
||||
const ok = done(resolve)
|
||||
const bad = done(reject)
|
||||
// M9: a cancelled dialog settles the promise.
|
||||
//
|
||||
// It did not before. `onchange` does not fire when the reader closes the
|
||||
// picker without choosing anything, so the promise stayed pending forever
|
||||
// — and the Campaigns screen awaits it, so its `finally` never ran and the
|
||||
// Import button sat disabled reading "Importing…" until the page was
|
||||
// reloaded. The rejection carries an empty message, because that screen
|
||||
// already treats a message-less error as "they changed their mind" and
|
||||
// says nothing: a cancelled dialog is not a failure to report.
|
||||
input.oncancel = () => bad(Object.assign(new Error(), { message: '' }))
|
||||
input.onchange = () => {
|
||||
const file = input.files[0]
|
||||
if (!file) return reject(new Error('No file selected'))
|
||||
if (!file) return bad(new Error('No file selected'))
|
||||
const reader = new FileReader()
|
||||
reader.onload = () => {
|
||||
try { resolve(JSON.parse(reader.result)) }
|
||||
catch { reject(new Error('Not valid JSON')) }
|
||||
try { ok(JSON.parse(reader.result)) }
|
||||
catch { bad(new Error('Not valid JSON')) }
|
||||
}
|
||||
reader.onerror = () => reject(new Error('Could not read file'))
|
||||
reader.onerror = () => bad(new Error('Could not read file'))
|
||||
reader.readAsText(file)
|
||||
}
|
||||
document.body.appendChild(input)
|
||||
input.click()
|
||||
})
|
||||
}
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
/* M9: the file picker that recovers a campaign.
|
||||
*
|
||||
* `pickJSONFile` is four lines of DOM and was the only control in the product
|
||||
* with no test at all, for a structural reason: it built a detached
|
||||
* `<input type="file">` and clicked it, so there was no element for a test — or
|
||||
* for WebDriver — to hand a file to. The import workflow could therefore only
|
||||
* ever be checked by calling the API underneath it, which is not the workflow.
|
||||
*
|
||||
* Appending the input made it testable, and writing the test found a real bug
|
||||
* that had been there since the picker was written: closing the dialog without
|
||||
* choosing anything never settled the promise, so the Campaigns screen's
|
||||
* `finally` never ran and its Import button stayed disabled reading
|
||||
* "Importing…" until the page was reloaded. The screen's own comment says a
|
||||
* cancelled picker is not worth a message — it had just never received one.
|
||||
*/
|
||||
|
||||
import { fireEvent } from '@testing-library/react'
|
||||
import { beforeEach, describe, expect, it } from 'vitest'
|
||||
import { pickJSONFile } from './components'
|
||||
|
||||
function theInput() {
|
||||
return document.querySelector('input[type="file"]')
|
||||
}
|
||||
|
||||
/** A `File` the way the browser hands one to a change event. */
|
||||
function jsonFile(name, contents) {
|
||||
return new File([JSON.stringify(contents)], name, { type: 'application/json' })
|
||||
}
|
||||
|
||||
/** Puts `files` on the input, since `files` is read-only in jsdom. */
|
||||
function choose(input, files) {
|
||||
Object.defineProperty(input, 'files', { value: files, configurable: true })
|
||||
fireEvent.change(input)
|
||||
}
|
||||
|
||||
beforeEach(() => { document.body.innerHTML = '' })
|
||||
|
||||
describe('pickJSONFile', () => {
|
||||
it('puts a findable input in the document rather than a detached one', () => {
|
||||
pickJSONFile().catch(() => {})
|
||||
const input = theInput()
|
||||
expect(input).toBeInTheDocument()
|
||||
expect(input.dataset.testid).toBe('import-file')
|
||||
expect(input.accept).toContain('json')
|
||||
// Out of sight, out of the tab order and out of the accessibility tree,
|
||||
// because the reader is looking at the browser's own dialog — but **not**
|
||||
// `hidden`, which would make it non-interactable and stop the browser
|
||||
// dispatching `change` when a file is chosen programmatically.
|
||||
expect(input.hidden).toBe(false)
|
||||
expect(input.getAttribute('aria-hidden')).toBe('true')
|
||||
expect(input.tabIndex).toBe(-1)
|
||||
expect(input.style.position).toBe('fixed')
|
||||
})
|
||||
|
||||
it('resolves with the parsed bundle', async () => {
|
||||
const promise = pickJSONFile()
|
||||
choose(theInput(), [jsonFile('c.json', { format: 'ai-dnd-adventure-v3' })])
|
||||
await expect(promise).resolves.toEqual({ format: 'ai-dnd-adventure-v3' })
|
||||
})
|
||||
|
||||
it('rejects a file that is not JSON, with a message worth showing', async () => {
|
||||
const promise = pickJSONFile()
|
||||
const input = theInput()
|
||||
Object.defineProperty(input, 'files', {
|
||||
value: [new File(['not json at all'], 'c.json')], configurable: true,
|
||||
})
|
||||
fireEvent.change(input)
|
||||
await expect(promise).rejects.toThrow(/not valid json/i)
|
||||
})
|
||||
|
||||
it('settles even when the file cannot be read, rather than hanging', async () => {
|
||||
// A real case, not a defensive one. A browser can hand the page a `File`
|
||||
// whose contents it will not then let the page read — a sandboxed Firefox
|
||||
// does exactly that for a path outside its confinement, and reports
|
||||
// `NotFoundError` from the FileReader with the name and size intact.
|
||||
// Whatever happens, the promise must settle: leaving it pending is what
|
||||
// left the Import button disabled reading "Importing…".
|
||||
const promise = pickJSONFile()
|
||||
const input = theInput()
|
||||
Object.defineProperty(input, 'files', {
|
||||
value: [jsonFile('c.json', { ok: true })], configurable: true,
|
||||
})
|
||||
fireEvent.change(input)
|
||||
await expect(Promise.race([
|
||||
promise.then(() => 'settled', () => 'settled'),
|
||||
new Promise((r) => { setTimeout(() => r('hung'), 300) }),
|
||||
])).resolves.toBe('settled')
|
||||
})
|
||||
|
||||
it('settles when the dialog is cancelled, instead of hanging forever', async () => {
|
||||
const promise = pickJSONFile()
|
||||
fireEvent(theInput(), new Event('cancel'))
|
||||
// Rejected, so the caller's `finally` runs — and with no message, so the
|
||||
// caller shows nothing. Both halves matter: a hang leaves the button
|
||||
// disabled, and a message would report a decision as a failure.
|
||||
await expect(promise).rejects.toSatisfy((err) => err.message === '')
|
||||
})
|
||||
|
||||
it('takes the input back out of the document however it settles', async () => {
|
||||
const resolved = pickJSONFile()
|
||||
choose(theInput(), [jsonFile('c.json', { ok: true })])
|
||||
await resolved
|
||||
expect(theInput()).toBeNull()
|
||||
|
||||
const cancelled = pickJSONFile()
|
||||
fireEvent(theInput(), new Event('cancel'))
|
||||
await cancelled.catch(() => {})
|
||||
expect(theInput()).toBeNull()
|
||||
})
|
||||
})
|
||||
@@ -85,6 +85,105 @@ function DebugLog() {
|
||||
)
|
||||
}
|
||||
|
||||
|
||||
/* M9. A verified copy of the whole database, on this machine.
|
||||
*
|
||||
* Deliberately small, and deliberately not a second export. The campaign export
|
||||
* on the Campaigns screen is the tool for moving one campaign to another
|
||||
* installation; this is the tool for keeping a copy of everything before doing
|
||||
* something risky. Conflating them would leave a reader guessing which one
|
||||
* answers "how do I not lose my campaigns".
|
||||
*
|
||||
* There is no restore button and no download link, and both absences are
|
||||
* decisions rather than gaps:
|
||||
*
|
||||
* * **Restore** means replacing the file the running application has open,
|
||||
* which is how somebody loses both copies at once. `DEVELOPMENT.md` carries
|
||||
* the procedure — stop the app, move the file, start it — and it is a
|
||||
* procedure precisely because each step needs the app to be stopped.
|
||||
* * **Download** would put a copy of every campaign on the machine into the
|
||||
* browser's download directory and its cache. For an application whose
|
||||
* premise is that the story does not leave the machine, a path the reader
|
||||
* can copy is the better default.
|
||||
*/
|
||||
function DatabaseBackup() {
|
||||
const toast = useToast()
|
||||
const [state, setState] = useState(null)
|
||||
const [busy, setBusy] = useState(false)
|
||||
|
||||
const load = () => {
|
||||
api.listBackups()
|
||||
.then(setState)
|
||||
.catch((err) => toast(err.message, 'error'))
|
||||
}
|
||||
|
||||
const take = async () => {
|
||||
setBusy(true)
|
||||
try {
|
||||
const result = await api.createBackup()
|
||||
toast(`Backup written: ${result.filename} (${formatBytes(result.bytes)}).`)
|
||||
load()
|
||||
} catch (err) {
|
||||
toast(err.message, 'error')
|
||||
} finally {
|
||||
setBusy(false)
|
||||
}
|
||||
}
|
||||
|
||||
return (
|
||||
<details
|
||||
className="advanced-block"
|
||||
data-testid="database-backup"
|
||||
onToggle={(e) => { if (e.currentTarget.open && state === null) load() }}
|
||||
>
|
||||
<summary>Back up everything on this machine</summary>
|
||||
<p className="field-hint">
|
||||
Writes a verified copy of the whole database — every campaign, every
|
||||
imported file, every setting — beside the database itself. The copy is
|
||||
checked before it is kept, and an existing backup is never overwritten.
|
||||
To export a single campaign so it can be opened somewhere else, use
|
||||
Export on the campaign instead.
|
||||
</p>
|
||||
<div className="panel-actions">
|
||||
<button type="button" className="primary" onClick={take} disabled={busy}>
|
||||
{busy ? 'Backing up…' : 'Back up now'}
|
||||
</button>
|
||||
</div>
|
||||
{state && (
|
||||
<>
|
||||
<p className="field-hint">
|
||||
Backups are written to <code>{state.directory}</code>. To restore
|
||||
one, stop the application, put the file in place of the database, and
|
||||
start it again.
|
||||
</p>
|
||||
{state.backups.length === 0 ? (
|
||||
<p className="dim">No backups yet.</p>
|
||||
) : (
|
||||
<ul className="backup-list">
|
||||
{state.backups.map((b) => (
|
||||
<li key={b.filename}>
|
||||
<code>{b.filename}</code>
|
||||
<span className="dim">
|
||||
{' '}{formatBytes(b.bytes)} · {new Date(b.taken_at).toLocaleString()}
|
||||
</span>
|
||||
</li>
|
||||
))}
|
||||
</ul>
|
||||
)}
|
||||
</>
|
||||
)}
|
||||
</details>
|
||||
)
|
||||
}
|
||||
|
||||
function formatBytes(bytes) {
|
||||
if (!Number.isFinite(bytes)) return ''
|
||||
if (bytes < 1024) return `${bytes} B`
|
||||
if (bytes < 1024 * 1024) return `${Math.round(bytes / 1024)} kB`
|
||||
return `${(bytes / (1024 * 1024)).toFixed(1)} MB`
|
||||
}
|
||||
|
||||
|
||||
export default function Settings() {
|
||||
const [settings, setSettings] = useState(null)
|
||||
const [testResult, setTestResult] = useState(null)
|
||||
@@ -347,6 +446,8 @@ export default function Settings() {
|
||||
</div>
|
||||
</details>
|
||||
|
||||
<DatabaseBackup />
|
||||
|
||||
<DebugLog />
|
||||
</section>
|
||||
</div>
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
/* M9: the database backup control, and what it must not become.
|
||||
*
|
||||
* The control itself is three lines of state, so the interesting assertions are
|
||||
* about the boundaries around it rather than about the button:
|
||||
*
|
||||
* * it is **not** an export. The campaign export moves one campaign to
|
||||
* another installation; this copies everything on this machine. A reader
|
||||
* who cannot tell them apart has no way to answer "how do I not lose my
|
||||
* campaigns", so the panel says which is which.
|
||||
* * it offers **no restore and no download**, and both are decisions. Restore
|
||||
* means replacing the file the running application has open; download means
|
||||
* putting every campaign on the machine into the browser's cache.
|
||||
* * a failure **says so**. A backup that silently did not happen is worse
|
||||
* than no backup, because the reader believes they have one.
|
||||
*/
|
||||
|
||||
import { screen, waitFor } from '@testing-library/react'
|
||||
import userEvent from '@testing-library/user-event'
|
||||
import { beforeEach, describe, expect, it, vi } from 'vitest'
|
||||
import { api } from '../api'
|
||||
import Settings from './Settings'
|
||||
import { mockModelStatus, renderWith } from '../test/helpers'
|
||||
|
||||
const LISTING = {
|
||||
directory: '/home/reader/.adventure/backups',
|
||||
backups: [
|
||||
{ filename: 'adventure-storyteller-20260907-043000.db', bytes: 2_400_000,
|
||||
taken_at: '2026-09-07T04:30:00' },
|
||||
{ filename: 'adventure-storyteller-20260901-101500.db', bytes: 2_100_000,
|
||||
taken_at: '2026-09-01T10:15:00' },
|
||||
],
|
||||
}
|
||||
|
||||
beforeEach(() => { vi.restoreAllMocks() })
|
||||
|
||||
async function openTheBackupPanel() {
|
||||
mockModelStatus(api)
|
||||
vi.spyOn(api, 'listBackups').mockResolvedValue(LISTING)
|
||||
await renderWith(<Settings />)
|
||||
const panel = screen.getByTestId('database-backup')
|
||||
await userEvent.click(screen.getByText(/Back up everything on this machine/i))
|
||||
return panel
|
||||
}
|
||||
|
||||
describe('the backup control', () => {
|
||||
it('lists the backups already on disk, and where they are', async () => {
|
||||
await openTheBackupPanel()
|
||||
await waitFor(() => {
|
||||
expect(screen.getByText('adventure-storyteller-20260907-043000.db')).toBeInTheDocument()
|
||||
})
|
||||
expect(screen.getByText('adventure-storyteller-20260901-101500.db')).toBeInTheDocument()
|
||||
expect(screen.getByText('/home/reader/.adventure/backups')).toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('takes a backup and reports what was written', async () => {
|
||||
const panel = await openTheBackupPanel()
|
||||
const create = vi.spyOn(api, 'createBackup').mockResolvedValue({
|
||||
directory: LISTING.directory,
|
||||
filename: 'adventure-storyteller-20260907-050000.db',
|
||||
bytes: 2_500_000, pages: 610, seconds: 0.02, integrity: 'ok',
|
||||
})
|
||||
await userEvent.click(screen.getByRole('button', { name: /Back up now/i }))
|
||||
await waitFor(() => expect(create).toHaveBeenCalled())
|
||||
expect(await screen.findByText(/Backup written/)).toBeInTheDocument()
|
||||
expect(screen.getByText(/adventure-storyteller-20260907-050000\.db/)).toBeInTheDocument()
|
||||
expect(panel).toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('reloads the list afterwards, so the new file is visible', async () => {
|
||||
await openTheBackupPanel()
|
||||
vi.spyOn(api, 'createBackup').mockResolvedValue({
|
||||
directory: LISTING.directory, filename: 'new.db', bytes: 1, integrity: 'ok',
|
||||
})
|
||||
await userEvent.click(screen.getByRole('button', { name: /Back up now/i }))
|
||||
await waitFor(() => expect(api.listBackups).toHaveBeenCalledTimes(2))
|
||||
})
|
||||
|
||||
it('reports a failure rather than looking as though it worked', async () => {
|
||||
await openTheBackupPanel()
|
||||
vi.spyOn(api, 'createBackup').mockRejectedValue(
|
||||
new Error('The backup was written but did not verify: page 4 missing.'),
|
||||
)
|
||||
await userEvent.click(screen.getByRole('button', { name: /Back up now/i }))
|
||||
expect(await screen.findByText(/did not verify/)).toBeInTheDocument()
|
||||
expect(screen.queryByText(/Backup written/)).not.toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('says which tool this is, and which one moves one campaign', async () => {
|
||||
await openTheBackupPanel()
|
||||
expect(screen.getByText(/every campaign, every/i)).toBeInTheDocument()
|
||||
expect(
|
||||
screen.getByText(/To export a single campaign so it can be opened somewhere else/i),
|
||||
).toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('offers no restore button and no download link', async () => {
|
||||
const panel = await openTheBackupPanel()
|
||||
await waitFor(() => {
|
||||
expect(screen.getByText(LISTING.directory)).toBeInTheDocument()
|
||||
})
|
||||
const labels = [...panel.querySelectorAll('button')].map((b) => b.textContent)
|
||||
expect(labels.some((label) => /restore/i.test(label))).toBe(false)
|
||||
expect(labels.some((label) => /download/i.test(label))).toBe(false)
|
||||
expect(panel.querySelectorAll('a[download]')).toHaveLength(0)
|
||||
expect(panel.querySelectorAll('a[href]')).toHaveLength(0)
|
||||
// And the procedure is stated instead, so the absence is an answer.
|
||||
expect(screen.getByText(/stop the application/i)).toBeInTheDocument()
|
||||
})
|
||||
|
||||
it('does not fetch anything until the panel is opened', async () => {
|
||||
mockModelStatus(api)
|
||||
const list = vi.spyOn(api, 'listBackups').mockResolvedValue(LISTING)
|
||||
await renderWith(<Settings />)
|
||||
expect(list).not.toHaveBeenCalled()
|
||||
})
|
||||
})
|
||||
@@ -204,6 +204,21 @@
|
||||
}
|
||||
.advanced-block > summary:hover { color: var(--text); }
|
||||
.advanced-block h4 { margin: 10px 0 4px; font-size: 0.74rem; color: var(--text-dim); }
|
||||
/* M9: the list of database backups already on disk. A filename, a size and a
|
||||
time — enough to recognise one, and nothing that needs a table. */
|
||||
.backup-list {
|
||||
margin: 8px 0 0;
|
||||
padding: 0;
|
||||
list-style: none;
|
||||
font-size: 0.78rem;
|
||||
}
|
||||
.backup-list li {
|
||||
padding: 3px 0;
|
||||
border-top: 1px solid var(--border);
|
||||
/* A long filename wraps rather than pushing the panel sideways. */
|
||||
overflow-wrap: anywhere;
|
||||
}
|
||||
.backup-list li:first-child { border-top: none; }
|
||||
.advanced-block pre {
|
||||
max-height: 16em;
|
||||
overflow: auto;
|
||||
|
||||
@@ -128,6 +128,35 @@ The current endpoint should be clear.
|
||||
|
||||
The input box always continues from the currently active story head.
|
||||
|
||||
### 8A. The reader must be able to tell where they are (recorded 2026-09-07)
|
||||
|
||||
A hands-on session against accepted M8 found the sentence above **too weak to
|
||||
hold the behaviour it names**. Undo worked correctly, the input box did continue
|
||||
from the active head — so both statements above were satisfied — and the reader
|
||||
still could not tell which point in the story they had moved to.
|
||||
|
||||
The requirement, stated so that a working implementation cannot satisfy it while
|
||||
a reader is lost:
|
||||
|
||||
> After Undo, Redo, a Save Point restore, an edit to an earlier turn, or any
|
||||
> other movement of the active story position, the reader should be able to
|
||||
> identify **where they now are in the visible story** — and, where it matters,
|
||||
> whether later story remains available ahead of them.
|
||||
|
||||
Two constraints on any solution:
|
||||
|
||||
- **No implementation terminology.** `branch`, `fork`, `node`, `head` and
|
||||
`depth` stay off the reader-facing surface (§38), which is what makes this a
|
||||
presentation problem rather than a labelling one.
|
||||
- **It must be observable**, not merely inferable from the transcript scrolling,
|
||||
so that a release test can decide it.
|
||||
|
||||
**No wording is prescribed here, and none is ratified.** A lightweight named
|
||||
position — `Moment 8` becoming `Moment 7` after an Undo, optionally noting that
|
||||
later story is available — is one candidate among others. Ownership is M11
|
||||
release polish; `V1-ACCEPTANCE-TESTS.md` §P1 records what must be settled before
|
||||
this can become an acceptance test.
|
||||
|
||||
## 9. User Turn Presentation
|
||||
|
||||
User messages should support:
|
||||
|
||||
@@ -1118,6 +1118,187 @@ Make campaigns portable and recoverable without losing lineage, state, knowledge
|
||||
|
||||
A campaign can be safely exported, imported into a clean data directory, and reopened at the exact intended active position with authoritative history/state intact.
|
||||
|
||||
## Status: COMPLETE — 2026-09-07, pending independent review
|
||||
|
||||
Implemented on `m9-recovery` from the signed M8 commit `1ce9972`, measured
|
||||
before and after against the same fixture, and verified in a real browser
|
||||
against a real narrator. `planning/reports/M9-IMPLEMENTATION-REPORT.md` is the
|
||||
implementer's account, written for a reviewer.
|
||||
|
||||
**What it delivered, beyond the scope list above:**
|
||||
|
||||
- **The bundle became a version, and the reason is a rule.** `ai-dnd-adventure-v3`.
|
||||
Everything M9 adds could have been an optional key, the way four earlier
|
||||
additions were — and that mechanism fails exactly here, because a v2 file with
|
||||
no prompt provenance is ambiguous between "written before M9" and "written by
|
||||
M9 from a campaign with none". A version is how a recovery file states what it
|
||||
was capable of recording. v1 and v2 are still read, and every seam from
|
||||
pre-active-head onward is tested.
|
||||
- **A third category of data.** "Chosen travels, derived is recomputed" was
|
||||
enough until stored prompts had to be decided. They are derived and must
|
||||
travel, so the rule is now chosen / **evidence** / rebuildable, and the test
|
||||
separating the last two is not "could this be recomputed" but "would a
|
||||
recomputation answer the same question".
|
||||
- **The M8 handoff on prompt provenance is closed.** An old turn in a restored
|
||||
campaign shows what it was actually given, after the source has been deleted,
|
||||
the canon edited and the state moved on.
|
||||
- **State events and proposals travel**, so a moved campaign can still say why
|
||||
its state is what it is — and a manual correction is still identifiable as
|
||||
one, which it was not before.
|
||||
- **Summaries travel with their coordinates**, so an abandoned line's summary is
|
||||
still ineligible after the move, and a moved campaign resumes with its
|
||||
long-story continuity instead of behaving like a new one.
|
||||
- **A verified SQLite backup**, using the online backup API rather than a file
|
||||
copy, taken while the application is running, with a browser control in
|
||||
Settings.
|
||||
- **Take parentage**, memory authority and parser/chunking versions, each
|
||||
closing a smaller fidelity loss.
|
||||
|
||||
**Three defects found by running the milestone's own tests, and fixed here:**
|
||||
|
||||
1. **Deleting a campaign leaked its FTS index rows**, and SQLite then handed the
|
||||
freed ids to the *next* source imported into *any* campaign, which failed
|
||||
with an integrity error. Reindex could not repair it either. Pre-existing
|
||||
since M7; both ends are now closed and an already-damaged database repairs
|
||||
itself with no migration.
|
||||
2. **An imported node with no state snapshot was stamped with the campaign's
|
||||
head state**, so an Undo to turn 2 showed what the story knew at turn 20 —
|
||||
M5's review finding 3, arriving through the import.
|
||||
3. **A snapshot's `source_id` was not being translated on import** because the
|
||||
relink mutated a dict in place, which a non-`MutableDict` column does not
|
||||
notice. Found by a test asserting the outcome rather than the call.
|
||||
|
||||
**One decision the brief asked for, made and recorded:**
|
||||
|
||||
**Story cards are compatibility-only legacy data, and no longer enter the
|
||||
narrator's prompt.** They still travel in both directions, the rows and the API
|
||||
stay, and `memorybank.cast_brief` still reads them as the summariser's character
|
||||
roster. What stops is the injection: a keyword-matched card arrived in front of
|
||||
the narrator as `World Lore: …` with no class, no visibility, no source, no way
|
||||
to switch it off since M8 removed the editor, and no row in the context
|
||||
inspector — which is `IMPORTED-KNOWLEDGE-DESIGN.md` §73's "alternate untracked
|
||||
path around the new knowledge authority/provenance rules" in as many words.
|
||||
|
||||
**Debt carried forward, deliberately:**
|
||||
|
||||
- **A long campaign's bundle has a measured ceiling: ~279 turns** against the
|
||||
20 MB import limit. A per-turn prompt contains the story so far, so carrying
|
||||
one per turn is O(turns²); compressing them inside the file cut that to about
|
||||
an eighth of what it would have been. **M9's additions account for only 12% of
|
||||
the ceiling.** The other 88% is the per-position narrative state document,
|
||||
which is 74% of a bundle and which v2 already carried — so lifting the ceiling
|
||||
means addressing that, not the evidence. Far beyond M11's 100-turn
|
||||
certification (14% of the cap), and stated with its measurement rather than
|
||||
hidden. A streaming or chunked import is the fix if a later milestone needs
|
||||
one; the asymmetry to know about is that such a campaign can still be exported
|
||||
and would be refused on import.
|
||||
- **No discarded-history recovery screen** (§63). M9's job was that retained
|
||||
history survives correctly so a later screen can use it; it does.
|
||||
- **No whole-transcript copy and no story search** (§77, §78) — still M11 or
|
||||
later.
|
||||
|
||||
---
|
||||
|
||||
# Post-M8 hands-on playtest findings — recorded 2026-09-07, owned by M11
|
||||
|
||||
**Not M9 defects, not caused by M9, and they did not block M9 acceptance.**
|
||||
Recorded here rather than only in the M9 report because a milestone report is
|
||||
archived when the next one replaces it, and these must not go with it.
|
||||
|
||||
They come from a real play session against **accepted, signed M8**: real
|
||||
browser, real trusted-LAN Ollama, narrator `qwen2.5:3b-instruct-16k`, disposable
|
||||
isolated campaign database. **That database was deliberately destroyed
|
||||
afterwards**, so the stored context snapshot for finding D is gone and no root
|
||||
cause is claimed for it. Full write-up, with the verification behind each
|
||||
mechanism, is in the M9 report's **§Y** (in `archive/milestone-reports/` once
|
||||
M10's report replaces it).
|
||||
|
||||
**These are M11's, and explicitly not M10's.** M10 is media-readiness
|
||||
architecture and stays bounded; it inherits them as known carry-forward only.
|
||||
|
||||
### A. The browser still calls the product "AI D&D" — M11 release polish
|
||||
|
||||
`frontend/index.html` still carries the inherited `<title>AI D&D</title>`.
|
||||
M8 changed the navigation, the inspector and the screens, and never claimed the
|
||||
title — so this is an uncovered gap rather than a false claim.
|
||||
|
||||
**Do not fix it with a find-and-replace to "Adventure Storyteller".**
|
||||
`SPECIFICATION.md` requires a genre-agnostic engine, and *Adventure* is narrower
|
||||
than the product. The naming decision is the repository owner's; a neutral
|
||||
working name such as **Interactive Story** and a tab form such as
|
||||
`<Campaign Name> — Interactive Story` are candidates, not decisions.
|
||||
|
||||
### B. After Undo, the reader cannot tell where they are — M11 UX polish
|
||||
|
||||
Undo behaved correctly (M3 semantics; re-verified throughout M9). The reader
|
||||
could not tell **which point in the story** they had moved to.
|
||||
|
||||
`BROWSER-UX-SPEC.md` §8 said only *"The current endpoint should be clear"*, and
|
||||
its companion sentence was already true while the reader was lost — so the
|
||||
requirement could not hold the behaviour. §8 has been strengthened to state the
|
||||
orientation requirement; **no UI text is prescribed**, because none is ratified.
|
||||
A `Moment 8` → `Moment 7` style indicator is a candidate. Needs a browser
|
||||
regression scenario.
|
||||
|
||||
### C. Narration length has no measurable effect — M11 realistic-model behaviour
|
||||
|
||||
The setup choice becomes **one English sentence** in the campaign's
|
||||
`ai_instructions` and changes **no generation setting**. Independently,
|
||||
`length_hint()` derives a numeric word range from the **global**
|
||||
`Settings.max_output_tokens` and places it after the history — and at the default
|
||||
800 it reads *"must not exceed 506 words, and it should not stop short of about
|
||||
177"* **identically for brief, medium and long**.
|
||||
|
||||
That is a mechanism, verified by reading and running the code — **not a proven
|
||||
cause** of what the reader saw; the narrator's instruction following is also in
|
||||
play. A reproduction must measure what enters the stored prompt, whether the
|
||||
setting moves any generation budget, and actual word/paragraph counts across
|
||||
repeated turns, on the reference 3B narrator **and** a stronger local one.
|
||||
Clearer numeric targets (`Brief ~100-200 words` and so on) are a design
|
||||
candidate, not ratified. **Do not hard-truncate prose** — the state block is
|
||||
emitted last and truncation removes it.
|
||||
|
||||
### D. Character identity / coreference confusion — M11 diagnostic
|
||||
|
||||
Four people in one scene — Bill (protagonist), Roger, John, Alice — and later
|
||||
narration treated Alice as two different Alices.
|
||||
|
||||
**Root cause UNKNOWN and no longer establishable.** Candidates: a model
|
||||
coreference failure on a correct prompt; duplicate/conflicting state; a
|
||||
context/summary/memory assembly failure; or a context that is not contradictory
|
||||
but too implicit for a small model.
|
||||
|
||||
**One structural fact to check first**, verified by reading the code: the
|
||||
narrative state **permits two entities to share a display name and reports
|
||||
nothing**. Entities are keyed by the model-supplied id; `DUPLICATE_ENTITY`
|
||||
rejects only a repeated *key*; no check exists on `name`. That is one of this
|
||||
finding's failure modes, and establishes nothing about what happened.
|
||||
|
||||
**M11 must run an explicit diagnostic** with a protagonist and three same-scene
|
||||
supporting characters, stressing pronouns, dialogue attribution, entrances and
|
||||
exits, reference by name and by role, and one character speaking about another.
|
||||
It must detect duplicate creation, same-name duplication, protagonist drift,
|
||||
misattributed dialogue, self-as-other reference, and state/context disagreement
|
||||
— and on any failure preserve the pre-generation state, the exact stored prompt
|
||||
snapshot, history, summaries, memories, imported knowledge, narrator output and
|
||||
model settings, then classify:
|
||||
|
||||
```text
|
||||
STATE DEFECT / CONTEXT ASSEMBLY DEFECT / DERIVED MEMORY-SUMMARY DEFECT /
|
||||
MODEL FAILURE WITH CORRECT CONTEXT / AMBIGUOUS
|
||||
```
|
||||
|
||||
**Do not "fix" a model failure by changing authoritative state, and do not blame
|
||||
the model if the prompt already contained the error.** M9 made all of that
|
||||
evidence portable, so a failing campaign can be exported and handed over intact.
|
||||
|
||||
**The standard fixture does not cover this class.** `TEST-CAMPAIGN-FIXTURE.md`'s
|
||||
seven traps are knowledge, authority, branch leakage and possession; there is no
|
||||
identity trap and its on-stage cast is effectively two people. The established
|
||||
fixture was **not modified** — it is the deterministic baseline earlier results
|
||||
are compared against. A companion fixture, `Multi-Character Identity Test`, is
|
||||
proposed in an appendix to that document.
|
||||
|
||||
---
|
||||
|
||||
# M10 — Future Media Extension Hooks Only
|
||||
@@ -1126,6 +1307,13 @@ A campaign can be safely exported, imported into a clean data directory, and reo
|
||||
|
||||
Preserve the approved future media interfaces without adding a media-generation dependency to v1.
|
||||
|
||||
## Note — the post-M8 playtest findings are **not** M10 scope
|
||||
|
||||
The four findings recorded above are owned by M11. M10 inherits them as known
|
||||
carry-forward items only: it should neither implement nor test them, and its
|
||||
scope below is unchanged by them. They are listed before this milestone rather
|
||||
than after it only because they were recorded during M9's closeout.
|
||||
|
||||
## Scope
|
||||
|
||||
- scene snapshots/packets suitable for future providers,
|
||||
@@ -1157,6 +1345,67 @@ Do not implement:
|
||||
|
||||
Future media providers can be added through defined local interfaces without redesigning core story authority/history.
|
||||
|
||||
## Status: COMPLETE — 2026-09-07, pending independent review
|
||||
|
||||
Implemented on `m10-media-hooks` from the signed M9 commit `44edece`.
|
||||
`planning/reports/M10-IMPLEMENTATION-REPORT.md` is the implementer's account,
|
||||
written for a reviewer.
|
||||
|
||||
**The finding that shaped the milestone: the scene snapshot already existed.**
|
||||
|
||||
The media contract's §5 asks for a persisted or derived scene snapshot, and M5
|
||||
built one three milestones ago. `narrative_state["scene"]` holds the summary, the
|
||||
location, who is present and the `(branch_id, depth)` coordinate; it is written
|
||||
by a validated `set_scene` event, snapshotted per position, and restored on every
|
||||
head move. That was verified with a probe — a campaign played, diverged, undone
|
||||
and exported — rather than taken from M9's report.
|
||||
|
||||
So M10 built **no scenes table**, and the Scene Packet is derived on read with a
|
||||
computed identity (`c<adventure>:b<branch>:<start>-<end>`) rather than an
|
||||
allocated one. A second scene store would have been a duplicate representation of
|
||||
the same fact, with its own lineage rules to get wrong; the lineage rules are the
|
||||
hard part, which is precisely the argument for reusing the ones that already
|
||||
work.
|
||||
|
||||
**What it delivered:**
|
||||
|
||||
- **One table, `visual_profiles`** — the only field in the contract's scene list
|
||||
that nothing already stored. Campaign-scoped rather than per-position, because
|
||||
a character does not change appearance when the story forks, and because
|
||||
per-position profiles would have cost 245 copies of the same 367 bytes in a
|
||||
120-turn campaign to say something that never varies.
|
||||
- **A scene packet built on read**, excluding the raw transcript, all imported
|
||||
knowledge, memories and summaries. Excluding imported knowledge *as a class* is
|
||||
what keeps a hidden Canon source out of a future depiction without a filter
|
||||
anyone has to remember to extend.
|
||||
- **Provider contracts as `typing.Protocol` structural types**, with an empty
|
||||
registry, no adapter, and no dependency added. Nothing imports a media library
|
||||
because none is installed.
|
||||
- **The STT asymmetry in the type**: a `DraftTranscription` is editable and has
|
||||
no commit method, so a transcriber structurally cannot bypass the authoritative
|
||||
commit path.
|
||||
- **A stricter endpoint policy than narration uses** — loopback only, reusing
|
||||
`endpoints.py`'s resolved-address check rather than trusting a hostname. No
|
||||
media configuration setting exists, because one that exists can be pointed at a
|
||||
cloud by mistake.
|
||||
- **Profiles travel in the bundle** with no format bump. M9's own semantic test
|
||||
decides it: an absent `visualProfiles` key is unambiguous, because a campaign
|
||||
with no profiles is the ordinary case. Older v3 files still import.
|
||||
|
||||
**A defect found by the milestone's own tests, and fixed here:**
|
||||
|
||||
- **A redundant index migration made two databases disagree.** M10 first shipped
|
||||
migration 93 creating `ix_visual_profiles_adventure`. `create_all` already
|
||||
builds `ix_visual_profiles_adventure_id` from the column's `index=True`, on
|
||||
fresh installs and existing databases alike — so an *upgraded* database ended
|
||||
up with both indexes and a fresh one with only the second. The comparison of a
|
||||
fresh schema against an upgraded schema is what caught it. **M10 adds no
|
||||
migration at all**; `LATEST_VERSION` stays 92.
|
||||
|
||||
**Carry-forward, unchanged by this milestone:** the four post-M8 playtest
|
||||
findings recorded above (browser title, post-Undo orientation, narration length,
|
||||
character identity) remain **M11's**. M10 neither implemented nor tested them.
|
||||
|
||||
---
|
||||
|
||||
# M11 — v1 Security, Long-Run, and Release Validation
|
||||
@@ -1178,7 +1427,11 @@ Validate the full product against the release contract after all functional mile
|
||||
- fantasy and science-fiction fixtures,
|
||||
- export/import/recovery tests,
|
||||
- migration tests,
|
||||
- documentation and packaging.
|
||||
- documentation and packaging,
|
||||
- **the four post-M8 hands-on playtest findings above**: the browser product
|
||||
name, reader orientation after history movement, a narration-length setting
|
||||
with a measurable effect, and the multi-character identity diagnostic with its
|
||||
companion fixture.
|
||||
|
||||
## Tests / Acceptance
|
||||
|
||||
|
||||
@@ -829,6 +829,61 @@ media_asset:
|
||||
|
||||
Potential metadata includes prompt, seed, model, workflow, dimensions, duration, character references, and source turn range.
|
||||
|
||||
## 28A. Media Extension Points (M10, as implemented)
|
||||
|
||||
M10 implemented the seams the two sections above describe, and the implementation
|
||||
is mostly an account of what it did **not** build.
|
||||
|
||||
```text
|
||||
visual_profiles the one thing M10 persists
|
||||
id, adventure_id campaign-scoped; no branch coordinate, deliberately
|
||||
entity_key the M5 narrative-state key: "mara", "the_office"
|
||||
descriptors open map of trait -> value
|
||||
features list of distinctive visible things
|
||||
style_notes free text about how it should be rendered
|
||||
created_at, updated_at
|
||||
UNIQUE (adventure_id, entity_key)
|
||||
```
|
||||
|
||||
No `scenes` table. No `media_jobs` table. No `media_assets` table.
|
||||
|
||||
**The scene snapshot of §20 already exists**, and has since M5. It is
|
||||
`narrative_state["scene"]` — `summary`, `location`, `present[]`, and the
|
||||
`at: {branch_id, depth}` coordinate that says where it was written. It is
|
||||
produced by the validated `set_scene` event, snapshotted per position in
|
||||
`actions.narrative_state_after`, restored by the head move on every Undo, Redo,
|
||||
Retry and Save Point restore, and carried in the v3 bundle. Building a second
|
||||
scene record beside it would have been a duplicate representation of the same
|
||||
fact with its own lineage rules to get wrong — and the lineage rules are the hard
|
||||
part, which is exactly why the answer is to reuse the one that already works.
|
||||
|
||||
So the **Scene Packet** (`app/media/packet.py`) is *derived on read* and stored
|
||||
nowhere. Its identity is `c<adventure>:b<branch>:<start>-<end>`, computed from
|
||||
the campaign and the position rather than allocated, so the same position yields
|
||||
the same id in any process and after any restart without a row to keep in step.
|
||||
|
||||
`media_job` and `media_asset` (§27, §28) remain **unbuilt**. M10 defines their
|
||||
contracts as `typing.Protocol` structural types in `app/media/providers.py` —
|
||||
`MediaProvider`, `SpeechProvider`, `TranscriptionProvider`, and the
|
||||
`MediaRequest` / `MediaResult` / `DraftTranscription` shapes — with an empty
|
||||
registry. A queue with no producer and no consumer would be speculative
|
||||
architecture, and this codebase has already declined that once: M6's
|
||||
`derived_status` carries the note "not a job queue".
|
||||
|
||||
**Why a visual profile has no branch coordinate.** Every other derived record in
|
||||
the schema carries `(branch_id, depth)` because it describes a *moment*. A
|
||||
profile describes none: a character does not change appearance because the story
|
||||
forked. Making it per-position would have hidden a reader's cast from them the
|
||||
moment they diverged, and would have put a descriptor document into every
|
||||
per-position snapshot — measured at 245 copies of 367 bytes in a 120-turn
|
||||
campaign, to say something that never varies (`tools/m10_media_cost.py`).
|
||||
|
||||
**A profile is not a fact.** Nothing here is state the story established.
|
||||
Writing one cannot change `narrative_state`, and the guarantee is structural
|
||||
rather than remembered: nothing in `app/media/` imports the code that writes it.
|
||||
The reverse direction is the same rule seen from the other side — a depiction
|
||||
never becomes canon (`MEDIA-EXTENSION-CONTRACT.md` §35, §37).
|
||||
|
||||
## 29. Export Package
|
||||
|
||||
A campaign export should be capable of preserving:
|
||||
@@ -887,6 +942,101 @@ has stopped being told the rules, with nothing to notice.
|
||||
A bundle written before M7 has no knowledge section and imports with an empty
|
||||
library, which is what such a campaign had.
|
||||
|
||||
### M9: the format became a version, and the evidence started travelling
|
||||
|
||||
**M9 bumped the format to `ai-dnd-adventure-v3`**, and the reason is a rule
|
||||
rather than a preference. Everything M9 added *could* have been an optional key
|
||||
read with `.get`, the way `persona`, `checkpoints`, `narrativeState` and
|
||||
`knowledge` each were. That mechanism stops working at exactly this addition:
|
||||
a v2 file carrying no prompt provenance is **ambiguous** — written before M9,
|
||||
when no file could carry one, or by M9 from a campaign whose turns predate the
|
||||
column? Those are different facts about the campaign and a reader has to be able
|
||||
to tell them apart. It is the same distinction the head rule above draws when it
|
||||
says a pre-M3 file opens at its tip *because that is the position such a file
|
||||
recorded*. A version number is how a recovery file states what it was capable of
|
||||
recording. The reader keeps every older version; only the writer moved.
|
||||
|
||||
What each version can be trusted to say:
|
||||
|
||||
```text
|
||||
v1 a linear story, its turns, and its retries as a repeating group
|
||||
v2 + the tree, the live flags, the after-snapshots, the chosen head,
|
||||
Save Points, the narrative state document, imported knowledge
|
||||
v3 + state events and proposals, historical prompt/context provenance,
|
||||
lineage-anchored summaries, take parentage, memory authority
|
||||
```
|
||||
|
||||
**Three categories, not two.** §31 below distinguishes authoritative from
|
||||
derived, which was sufficient until M9 had to decide about stored prompts. They
|
||||
are derived — a machine assembled them — and they must travel anyway, so the
|
||||
rule the bundle applies has a middle category:
|
||||
|
||||
```text
|
||||
chosen what a person decided: the story, the head, the takes, the
|
||||
Save Points, the classifications, the canon. travels
|
||||
evidence what happened, and what the application was told at the time:
|
||||
the state events and proposals, the per-turn prompt and the
|
||||
passages it was shown, the model and generation settings that
|
||||
turn ran under. travels
|
||||
rebuildable a deterministic function of what travels: knowledge passages,
|
||||
the FTS index, embeddings, the branch lineage cache.
|
||||
rebuilt on import
|
||||
```
|
||||
|
||||
The test that separates evidence from rebuildable is **not** "could this be
|
||||
recomputed" but "would a recomputation answer the same question". Rebuilding the
|
||||
FTS index answers the same question it answered before. Rebuilding an old turn's
|
||||
prompt does not — it would say what that turn *would be told now*, from today's
|
||||
canon, today's sources and today's state, which is the opposite of what the
|
||||
context inspector is for. Historical evidence is not a cache.
|
||||
|
||||
So v3 additionally carries, all restored verbatim:
|
||||
|
||||
- **`stateEvents` and `stateProposals`.** §17's hybrid keeps the events for
|
||||
audit and the snapshots for restore; v2 carried only the snapshots, so a moved
|
||||
campaign could be read at any position and could no longer say what changed
|
||||
there, who asserted it, or what the value was before. A **manual correction**
|
||||
was the worst case: the one state change no narration explains, and with the
|
||||
events gone nothing distinguished it from something the story established.
|
||||
Both tables travel, because the inspector reads both — the event says what was
|
||||
accepted and the proposal says what the model asked for and what was refused.
|
||||
- **A per-node context snapshot**, which is `SPECIFICATION.md` §6.1's exact
|
||||
prompt/context record. It carries the assembled prompt section by section, the
|
||||
passages retrieved with the text each supplied, which summary was eligible,
|
||||
and the model and generation settings the call ran under — so an old turn can
|
||||
still say what it was told after the source was deleted, the canon edited, the
|
||||
chunker changed and the state moved on. Stored once per turn on the live
|
||||
attempt, so a retried turn is not a multiplier.
|
||||
- **`summaries`**, with the coordinate that decides eligibility. v2 carried only
|
||||
the `storySummary` mirror, which has no lineage of its own, so a restored
|
||||
campaign resumed with no usable long-story continuity — and a summary
|
||||
belonging to an abandoned line stays ineligible after the move for the same
|
||||
reason it was before it: eligibility is the coordinate lying on the active
|
||||
capped lineage, not a stored flag.
|
||||
- **Take parentage**, so attempts under two different takes of one turn stay two
|
||||
pagers rather than merging into one.
|
||||
- **Memory `authority`**, so a heuristic memory is not promoted to accepted
|
||||
story by being moved (F07).
|
||||
- **Per-source `parserVersion`/`chunkingVersion`**, recording what produced the
|
||||
passages a historical retrieval record describes.
|
||||
|
||||
**Encoding.** A per-turn prompt contains the story so far, so one per turn is
|
||||
O(turns²) in campaign length — measured at 20,797 bytes per turn at turn 20 and
|
||||
55,291 at turn 120, 68% of a 9.7 MB file. The snapshot therefore travels as
|
||||
`contextSnapshotZ`: the same JSON, zlib-compressed and base64-encoded, using the
|
||||
same pack/unpack the database column already uses. Nothing is dropped or
|
||||
summarised; the file is still JSON, and every other section of it is still plain
|
||||
text. The plain `contextSnapshot` key is still read and takes precedence, so a
|
||||
hand-edited file keeps importing.
|
||||
|
||||
**Two pointers are translated on import, and nothing else is.** Branch numbers
|
||||
already were. M9 adds the `source_id` inside a restored retrieval record: it
|
||||
names a row on the machine that wrote the file, so left alone it would point the
|
||||
inspector's "open this source" at whatever holds that id here. Where the file's
|
||||
own knowledge section contains the source it is repointed; where it does not — a
|
||||
source deleted before the export — it becomes `null`, and the record keeps its
|
||||
text and filename. The evidence is never rewritten; only the pointer is.
|
||||
|
||||
## 30. Deletion vs Archival
|
||||
|
||||
The system must distinguish:
|
||||
@@ -916,6 +1066,23 @@ Retry, Undo, Restore, and branch switching must not silently perform permanent d
|
||||
|
||||
Derived data should be rebuildable where practical.
|
||||
|
||||
**"Where practical" does real work in that sentence, and M9 had to split this
|
||||
list to act on it** (§29). Embeddings and the lexical and semantic indexes are
|
||||
deterministic functions of content that travels, so a rebuild answers the same
|
||||
question and they are not exported. A **summary** is not: it took a model call,
|
||||
it describes a stretch of story that may since have been abandoned, and
|
||||
regenerating one on another machine produces different prose about a different
|
||||
reading — so it is derived, not practically rebuildable, and it travels with the
|
||||
coordinate that decides whether it still applies.
|
||||
|
||||
The same reasoning puts **stored prompt/context snapshots** on the travelling
|
||||
side, and they are not in either list above because they are neither: they are
|
||||
not authoritative — nothing decides anything from them — and calling them
|
||||
derived would invite a rebuild. They are *evidence*: a record of what the
|
||||
application was told at the time, which a regeneration would not reproduce
|
||||
because it would use today's canon, today's sources and today's state. §29 states
|
||||
the three-way rule the export applies.
|
||||
|
||||
## 32. Provenance
|
||||
|
||||
Important information should answer:
|
||||
|
||||
@@ -1110,6 +1110,47 @@ Normal imported files are campaign-level source material and need not inherit st
|
||||
|
||||
Story Cards may remain as an inherited authored-rule/lore primitive during migration if useful, but they must not become an alternate untracked path around the new knowledge authority/provenance rules.
|
||||
|
||||
### Settled in M9 (2026-09-07): compatibility-only, and out of the prompt
|
||||
|
||||
M8 removed the Story Card browser editor and left the question open; the M9
|
||||
brief asked for it to be decided. The finding was that story cards **were** the
|
||||
alternate untracked path the paragraph above forbids, and not in principle: a
|
||||
keyword-matched card was injected into the narrator's prompt as
|
||||
`World Lore: <entry>`, taking up to 40% of what was left after the imported
|
||||
knowledge had been placed, with
|
||||
|
||||
- no class, so nothing framed how far the narrator could rely on it;
|
||||
- no visibility, so no narrator-only distinction existed;
|
||||
- no source, no hash and no lifecycle, so there was nothing to disable;
|
||||
- no browser surface after M8, so a reader could neither see nor switch it off;
|
||||
- no row in the context inspector, which renders `knowledge` and never rendered
|
||||
`cards`;
|
||||
|
||||
and competing with imported Canon for one budget, which is the arrangement M7
|
||||
spent a milestone separating.
|
||||
|
||||
**The decision, and it is the smallest change that closes it:**
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| New export | carries them, unchanged, under `storyCards` |
|
||||
| Legacy import | accepted, unchanged, from every format version |
|
||||
| Normal narration | **no longer reached.** The `world_lore` section is gone |
|
||||
| Re-export | carries them again, so a round trip destroys nothing |
|
||||
|
||||
Nothing is deleted. The rows stay, the `/api/story-cards` endpoints stay, and
|
||||
`memorybank.cast_brief` still reads them as the **summariser's character
|
||||
roster** — that names who is on stage so a memory says "Aldric" rather than
|
||||
"he", never reaches the narrator, and every memory written from it is
|
||||
authority-classified by the application afterwards. The `cards` key stays in the
|
||||
context report and is now always empty for a new turn, because M9 made
|
||||
historical snapshots portable and an old turn's record must go on saying that
|
||||
story cards were included.
|
||||
|
||||
A campaign that wants the narrator to know something imports it as Canon,
|
||||
Reference or Inspiration, where it is classified, inspectable, disableable and
|
||||
attributable — which is what §73 asks for.
|
||||
|
||||
### Retrieval implementation direction
|
||||
|
||||
Use:
|
||||
|
||||
@@ -1420,3 +1420,122 @@ Optional Media Coordinator
|
||||
```
|
||||
|
||||
No production media provider is required for v1.
|
||||
|
||||
## 90. As Implemented (M10)
|
||||
|
||||
The contract above is Phase 0B design. This section records what M10 built
|
||||
against it, what it deliberately left unbuilt, and the three places where
|
||||
implementation answered a question the contract left open. It is appended rather
|
||||
than woven in, so the original contract stays readable as the document it is.
|
||||
|
||||
### 90.1 What was built
|
||||
|
||||
```text
|
||||
app/media/packet.py §10-11 the scene packet, derived on read
|
||||
app/media/profiles.py §7-9 visual profiles for any entity
|
||||
app/media/providers.py §13, §21-28, §55 the contracts and the endpoint policy
|
||||
app/models.py VisualProfile — the only table M10 adds
|
||||
app/routers/adventures/visuals.py six endpoints, all read/write of the above
|
||||
```
|
||||
|
||||
Four HTTP endpoints for profiles (list, read, write, delete) and one for the
|
||||
packet. No coordinator, no queue, no worker, no provider adapter, no dependency
|
||||
added.
|
||||
|
||||
### 90.2 §5 was already satisfied — the scene snapshot exists
|
||||
|
||||
The single most consequential finding of the milestone. §5 requires a persisted
|
||||
or derived scene snapshot; **M5 had already built it**, and it has been carrying
|
||||
lineage correctly for three milestones. `narrative_state["scene"]` holds the
|
||||
summary, the location, who is present, and the `(branch_id, depth)` coordinate;
|
||||
it is written by a validated `set_scene` event, snapshotted per position, and
|
||||
restored on every head move.
|
||||
|
||||
That was verified rather than assumed — a probe played a campaign, diverged it,
|
||||
and checked that the scene at each position was the scene that position had, that
|
||||
Undo cleared it back to the state before, and that a bundle carried both
|
||||
branches' scenes.
|
||||
|
||||
So M10 built **no scenes table**. §6 says the scene snapshot must never be
|
||||
authoritative over the story; deriving it from the authoritative state on read is
|
||||
the strongest available form of that guarantee, because there is no second copy
|
||||
that could disagree.
|
||||
|
||||
### 90.3 §12 answered: the packet excludes more than the transcript
|
||||
|
||||
§12 says the packet must not require raw transcript access. The implementation
|
||||
draws the line wider, and this is a deliberate reading rather than an omission.
|
||||
|
||||
The packet carries **what the story established at this position**: location,
|
||||
present characters with their profiles, significant objects, an action summary,
|
||||
continuity constraints, ambience, turn range, lineage. It excludes the raw
|
||||
transcript, **all imported knowledge** (§7 of `IMPORTED-KNOWLEDGE-DESIGN.md`),
|
||||
memories and summaries.
|
||||
|
||||
Excluding imported knowledge as a *class* is what makes the hidden-information
|
||||
rule hold. A narrator-only Canon source — the mechanism a reader uses to keep a
|
||||
secret from themselves — never reaches a depiction, and does not need a filter
|
||||
that someone must remember to apply to each new secret. Once the story
|
||||
*establishes* something through a validated event it is no longer narrator-only,
|
||||
and it appears in the packet, because at that point it is something that
|
||||
happened rather than something the narrator was told.
|
||||
|
||||
### 90.4 §14-15 left unbuilt, and why
|
||||
|
||||
`MediaJob` and `MediaAsset` are defined as contracts (`MediaRequest`,
|
||||
`MediaResult`, `ProviderCapabilities`) and not as tables. A job queue with no
|
||||
producer and no consumer would be speculative architecture whose shape would be
|
||||
decided by a provider nobody has chosen yet; the codebase declined the same thing
|
||||
once already, in M6's `derived_status` ("not a job queue"). §16-20, §45-47 —
|
||||
provenance, cleanup, retries, lineage — are therefore also deferred, and they
|
||||
should be designed against a real coordinator.
|
||||
|
||||
What M10 does guarantee for them is the part that would be expensive to retrofit:
|
||||
the scene identity a future asset must reference (`c<adventure>:b<branch>:<start>-<end>`)
|
||||
is derived from the campaign and position rather than allocated, so it is stable
|
||||
across processes, restarts and re-derivation without a row to keep in step.
|
||||
|
||||
### 90.5 §7-9 collapsed into one table, deliberately
|
||||
|
||||
The contract describes character, location and item profiles in three sections.
|
||||
The implementation has one `visual_profiles` table keyed by the M5 entity key,
|
||||
because M5's entity model is genre-neutral by design and a character, a location,
|
||||
an item, a vehicle and a spaceship are all entities with a `type`. Three tables —
|
||||
or one table with a `kind` column duplicating the entity's own `type` — would
|
||||
have reintroduced the genre shape M5 spent a milestone removing.
|
||||
|
||||
The fields are open by construction: `descriptors` is a trait map, `features` a
|
||||
list, `style_notes` free text. `{"hair": "dark auburn"}` and
|
||||
`{"hull": "pitted white composite"}` are the same shape. The contract's examples
|
||||
are fantasy-shaped and the test fixture is deliberately not
|
||||
(`backend/tests/m10_fixture.py`: four people in an office), because a schema
|
||||
written while looking at hair and oil lamps acquires that shape without anyone
|
||||
choosing it.
|
||||
|
||||
### 90.6 §27-28 implemented strictly
|
||||
|
||||
A media provider endpoint must be **loopback**. `providers.endpoint_rejection_reason`
|
||||
reuses the local-only policy in `app/endpoints.py` — which resolves the address
|
||||
rather than trusting the hostname — and then requires loopback in addition. This
|
||||
is stricter than narrator inference, which permits a trusted LAN host: a GPU
|
||||
rendering a reader's campaign is a machine that reader is sitting at. No TLS
|
||||
verification bypass exists anywhere in the path.
|
||||
|
||||
There is no provider configuration setting, because none is needed yet, and a
|
||||
setting that exists can be pointed at a cloud by mistake.
|
||||
|
||||
### 90.7 §24A implemented as an asymmetry in the type
|
||||
|
||||
`TranscriptionProvider.transcribe` returns a `DraftTranscription` carrying
|
||||
`editable: bool = True` and **no commit method**. A transcriber can produce a
|
||||
draft and structurally cannot submit one. §24A's rule — STT never bypasses the
|
||||
authoritative commit path — is therefore enforced by the shape of the interface
|
||||
rather than by a caller remembering it.
|
||||
|
||||
### 90.8 §35 and §37 hold structurally
|
||||
|
||||
Nothing in `app/media/` imports the code that writes narrative state, no media
|
||||
event type exists in the state vocabulary, and every M10 test that touches the
|
||||
media layer compares the authoritative document before and after and requires it
|
||||
to be identical (`backend/tests/test_m10_authority.py`). A depiction cannot
|
||||
become canon because there is no path by which it could.
|
||||
|
||||
+69
-38
@@ -3,7 +3,8 @@
|
||||
**This file is the index. Start here.**
|
||||
|
||||
**Current state:** Phase 0 complete; AI-DnD forked as the production base;
|
||||
milestones **M1 through M6 implemented and accepted** (M3 and M4: 2026-09-03;
|
||||
milestones **M1 through M8 implemented and accepted**, and **M9 and M10
|
||||
implemented and awaiting review**. M1-M6 were accepted on the dates below (M3 and M4: 2026-09-03;
|
||||
M5: 2026-09-04; M6: 2026-09-06). M5 and M6 were each accepted only after an
|
||||
independent review found a real defect and a corrective pass fixed it.
|
||||
|
||||
@@ -19,8 +20,23 @@ carries the closeout: the build-evidence classification in its §P, finding 14's
|
||||
operational resolution, and the acceptance record in its §V. The M8 tree is
|
||||
staged and awaits the repository owner's signed commit.
|
||||
|
||||
**Next: M9 — Export, Backup, Recovery, and Migration Hardening.** It has not
|
||||
been started.
|
||||
**M9 — Export, Backup, Recovery, and Migration Hardening — is implemented and
|
||||
awaiting independent review** (2026-09-07).
|
||||
`reports/M9-IMPLEMENTATION-REPORT.md` is the implementer's account, written for
|
||||
a reviewer: a set of claims with the measurements attached, not yet a record of
|
||||
acceptance. M8's report has moved to `archive/milestone-reports/`, which is
|
||||
where a milestone report goes once the next milestone's report replaces it.
|
||||
|
||||
**M10 — Future Media Extension Hooks Only — is implemented and awaiting
|
||||
independent review** (2026-09-07). `reports/M10-IMPLEMENTATION-REPORT.md` is the
|
||||
implementer's account. It built the seam and no media: one `visual_profiles`
|
||||
table, a scene packet derived on read, provider contracts with an empty
|
||||
registry, and no dependency added. Its central finding is that the scene
|
||||
snapshot the media contract asks for **already existed**, built by M5.
|
||||
|
||||
**Next: M11 — v1 Security, Long-Run, and Release Validation.** It has not been
|
||||
started, and no brief for it exists. It also owns the four post-M8 hands-on
|
||||
playtest findings recorded in `BUILD-MILESTONES.md`.
|
||||
|
||||
**Package version:** see `VERSION.md`, which records what each revision changed
|
||||
and why.
|
||||
@@ -82,8 +98,8 @@ Two standing qualifications:
|
||||
| Document | What it is for |
|
||||
| --- | --- |
|
||||
| `SPECIFICATION.md` | What the product must do. The top of the authority order. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M7 built, recorded as fact. |
|
||||
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the export shape. |
|
||||
| `TECHNICAL-DESIGN.md` | The selected architecture, including what M1-M10 built, recorded as fact. |
|
||||
| `DATA-MODEL.md` | Entities, the stored head, branch disposition, and the v3 export contract. |
|
||||
| `STORY-BRANCH-SEMANTICS.md` | Undo/Redo/Retry/branch/take behavior, including the M3 ratifications. |
|
||||
| `CONTEXT-AND-MEMORY.md` | Prompt assembly, summarization, branch-safe memory. |
|
||||
| `IMPORTED-KNOWLEDGE-DESIGN.md` | Canon / Reference / Inspiration knowledge as a first-class subsystem. |
|
||||
@@ -110,9 +126,10 @@ Two standing qualifications:
|
||||
10. `BROWSER-UX-SPEC.md`
|
||||
11. `V1-ACCEPTANCE-TESTS.md`
|
||||
12. `DECISIONS/` — all of them; they are short.
|
||||
13. `reports/M8-IMPLEMENTATION-REPORT.md`, for what the most recent milestone
|
||||
actually left behind — reading it as a claim to check, not a record, until
|
||||
it is reviewed. Nothing in `planning/archive/` unless sent there.
|
||||
13. `reports/M10-IMPLEMENTATION-REPORT.md` and
|
||||
`reports/M9-IMPLEMENTATION-REPORT.md`, for what the most recent milestones
|
||||
actually left behind — read as claims to check, not records, until they are
|
||||
reviewed. Nothing in `planning/archive/` unless sent there.
|
||||
|
||||
## Architectural decisions
|
||||
|
||||
@@ -143,20 +160,23 @@ work until Phase 0 closes — which Phase 0 satisfied on 2026-09-01. It is in
|
||||
`reports/` holds the report for the milestone most recently completed, because
|
||||
that is the one the next milestone's planning has to consult:
|
||||
|
||||
- `reports/M8-IMPLEMENTATION-REPORT.md` — the M8 implementation, its baseline
|
||||
UX measurement, and the browser evidence for every acceptance test it claims.
|
||||
Written by the implementer for an independent reviewer, and completed at
|
||||
closeout after that review accepted the milestone: it is a set of claims with
|
||||
the measurements attached **and** the record of the acceptance. Its §U carries
|
||||
the M9 handoff — the four questions the next brief has to decide.
|
||||
- `reports/M10-IMPLEMENTATION-REPORT.md` — the M10 implementation: the media
|
||||
seam, everything it deliberately did not build, and the evidence for K01-K04.
|
||||
Written by the implementer for an independent reviewer, so it is a set of
|
||||
claims with the measurements attached and **not** a record of acceptance.
|
||||
- `reports/M9-IMPLEMENTATION-REPORT.md` — the M9 implementation: the measured M8
|
||||
portability baseline it started from, the final bundle contract, and the
|
||||
evidence for every acceptance test it claims. Its §W carries the M10-M11
|
||||
handoff and its §Y holds the post-M8 playtest findings.
|
||||
|
||||
**It stays here until M9's report replaces it.** A milestone report is useful
|
||||
during the immediately following milestone; M8's is not archived merely
|
||||
because M8 is accepted.
|
||||
**It stays here rather than moving to the archive**, against the usual
|
||||
rotation, because M9 has not been accepted yet: a reviewer of either milestone
|
||||
needs it, since M10 built on M9 and its baseline is M9's. It moves once M9 is
|
||||
accepted.
|
||||
|
||||
Completed earlier milestones are in `archive/milestone-reports/`, which M7's
|
||||
report joined when M8's was written: a milestone report is useful during the
|
||||
immediate next milestone and historical afterwards. M1-M7 are all there,
|
||||
Completed earlier milestones are in `archive/milestone-reports/`, which M8's
|
||||
report joined when M9's was written: a milestone report is useful during the
|
||||
immediate next milestone and historical afterwards. M1-M8 are all there,
|
||||
unedited.
|
||||
|
||||
## The decision this package rests on
|
||||
@@ -282,33 +302,44 @@ Milestone M8 COMPLETE / ACCEPTED (2026-09-06)
|
||||
story operations review + closeout, in sequence
|
||||
|
|
||||
v
|
||||
Milestone M9 NEXT — not started
|
||||
export, backup, recovery, see BUILD-MILESTONES.md
|
||||
migration hardening
|
||||
Milestone M9 COMPLETE — awaiting review (2026-09-07)
|
||||
export, backup, recovery, reports/M9-IMPLEMENTATION-REPORT.md
|
||||
migration hardening bundle format v3; SQLite online backup
|
||||
|
|
||||
v
|
||||
M10-M11, one at a time see BUILD-MILESTONES.md
|
||||
Milestone M10 COMPLETE — awaiting review (2026-09-07)
|
||||
future media extension hooks reports/M10-IMPLEMENTATION-REPORT.md
|
||||
scene packet derived, not stored; no media
|
||||
|
|
||||
v
|
||||
Milestone M11 NEXT — not started
|
||||
v1 security, long-run, release see BUILD-MILESTONES.md
|
||||
validation also owns the post-M8 playtest findings
|
||||
```
|
||||
|
||||
## Stop Rule
|
||||
|
||||
**One milestone at a time. Do not begin a milestone before its brief exists.**
|
||||
|
||||
**No M9 brief has been prepared.** Writing one is the current action, informed
|
||||
by the M8 report and by the debt `BUILD-MILESTONES.md` records against M8 — in
|
||||
particular that the campaign bundle still carries no context snapshots, so an
|
||||
imported campaign has no historical prompt provenance; that story cards survive
|
||||
in the backend and the bundle with no browser surface, and M9 should decide
|
||||
deliberately whether the bundle keeps carrying them; and that a deployment whose
|
||||
Ollama enforces a small context window truncates an imported long campaign
|
||||
immediately unless the `DEVELOPMENT.md` procedure or a matching
|
||||
`context_token_budget` is applied.
|
||||
**No M11 brief has been prepared**, and neither M9 nor M10 is accepted — both
|
||||
are implemented and awaiting independent review. Writing the M11 brief is the
|
||||
action after those reviews close, informed by the M9 report's §W, the M10
|
||||
report's handoff, and the four post-M8 playtest findings in
|
||||
`BUILD-MILESTONES.md`.
|
||||
|
||||
**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`: story
|
||||
cards have no browser editor, the RPG world state is read-only, copy is
|
||||
per-message only, there is no discarded-history recovery screen, and the tablet
|
||||
layout is usable but untuned. Each names the milestone that owns it; none is an
|
||||
open M8 condition.
|
||||
All three questions the M8 debt raised against M9 are settled and recorded:
|
||||
the bundle carries historical context snapshots (`DATA-MODEL.md` §29); story
|
||||
cards are compatibility-only legacy data and no longer reach the narrator
|
||||
(`IMPORTED-KNOWLEDGE-DESIGN.md` §73); and the deployment context ceiling is
|
||||
documented in `DEVELOPMENT.md` with the note that an imported long campaign
|
||||
meets it on its first turn rather than gradually. The window itself stays M11's.
|
||||
|
||||
**M8's own carried debt** is recorded under M8 in `BUILD-MILESTONES.md`. Two of
|
||||
its five items are now closed by M9 — story cards have a settled policy, and the
|
||||
bundle carries the provenance. The RPG world state is still read-only, copy is
|
||||
still per-message only, there is still no discarded-history recovery screen, and
|
||||
the tablet layout is still untuned. Each names the milestone that owns it; none
|
||||
is an open M8 condition.
|
||||
|
||||
The M6 retrieval debt this milestone was warned about is partly addressed and
|
||||
partly still open. Imported material does **not** compete with story memory for
|
||||
|
||||
@@ -766,6 +766,43 @@ The browser should access them through:
|
||||
|
||||
Do not expose arbitrary filesystem browsing.
|
||||
|
||||
## 42A. Media Endpoints As Implemented (M10) — narrower than this document allows
|
||||
|
||||
M10 built the media seam and **did not widen the trust boundary**. Two notes,
|
||||
because in one place the implementation is deliberately stricter than the text
|
||||
above, and a stricter implementation than the threat model describes is still a
|
||||
discrepancy worth writing down.
|
||||
|
||||
**Media endpoints are loopback only.** §73's pass condition permits "explicitly
|
||||
configured trusted-LAN Ollama/media endpoints". `providers.endpoint_rejection_reason`
|
||||
allows that for narration and refuses it for media: it applies the shared
|
||||
local-only policy in `app/endpoints.py` — which resolves the address rather than
|
||||
trusting the hostname, so `localhost.evil.example` does not pass — and then
|
||||
requires loopback in addition. The reasoning is that a GPU rendering someone's
|
||||
campaign is a machine that person is sitting at, and that a picture of a scene
|
||||
carries the scene with it. If a later milestone finds a real trusted-LAN media
|
||||
use, that is a decision to make explicitly, not a limit to relax quietly.
|
||||
|
||||
**Nothing is contacted, and nothing can be configured to be.** M10 adds no
|
||||
provider adapter, no HTTP client, and no media endpoint *setting* — a setting
|
||||
that exists is a setting that can be pointed at a cloud by mistake. The registry
|
||||
ships empty; no module under `app/media/` imports `httpx`, `requests`,
|
||||
`urllib.request`, `socket`, `aiohttp` or `subprocess`, and that is asserted by
|
||||
test rather than by inspection (`backend/tests/test_m10_no_media.py`). No TLS
|
||||
verification bypass exists anywhere in the path.
|
||||
|
||||
**§41's media-metadata concern is unreached**, because no media is generated and
|
||||
no asset is stored. It stays open for whichever milestone builds a coordinator.
|
||||
|
||||
**One new boundary that is not a network one.** The Scene Packet is the input a
|
||||
future provider would receive, so what it carries is a disclosure decision. It
|
||||
carries what the *story* established at a position and excludes the raw
|
||||
transcript, all imported knowledge, memories and summaries — so a hidden Canon
|
||||
source (§56's story secrets) cannot reach a depiction. The exclusion is by class
|
||||
rather than by filtering marked secrets, which is what makes it hold for a secret
|
||||
nobody thought to mark. Tested with a sentinel in a hidden source, alongside a
|
||||
positive control proving the narrator did receive it.
|
||||
|
||||
## 43. Logging
|
||||
|
||||
Logs should minimize story-content exposure.
|
||||
|
||||
@@ -608,6 +608,89 @@ stated reason. A misplaced head affects every read in the file; a bookmark
|
||||
pointing outside the story affects only itself, and rejecting a whole campaign
|
||||
to protect one bookmark would lose the story to save the pointer.
|
||||
|
||||
### 9.3 As implemented in M9 — the version, and the third category
|
||||
|
||||
**The format is now `ai-dnd-adventure-v3`, and the bump is the design.** §9.1 and
|
||||
§9.2 each declined one, correctly: an absent `headDepth` or `checkpoints` key is
|
||||
unambiguous, because a file either states a position or it does not. That
|
||||
property fails for what M9 adds. A v2 file with no prompt provenance may have
|
||||
been written before M9, when no file could carry any, or by M9 from a campaign
|
||||
whose turns predate the column — different facts about the campaign, and a
|
||||
reader has to be able to tell them apart. A version number is how a recovery
|
||||
file states what it was capable of recording, which is exactly the reasoning
|
||||
§9.1 uses to justify opening a pre-M3 file at its tip. The reader keeps every
|
||||
version; only the writer moved.
|
||||
|
||||
**§9.1's two categories became three.** "Chosen travels, derived is recomputed"
|
||||
was sufficient until M9 had to decide about stored prompts, which are derived —
|
||||
a machine assembled them — and must travel anyway:
|
||||
|
||||
```text
|
||||
chosen the story, the head, the takes, the Save Points, the
|
||||
classifications, the canon travels
|
||||
evidence the state events and proposals, the per-turn prompt and the
|
||||
passages it was shown, the model and generation settings that
|
||||
turn ran under travels
|
||||
rebuildable knowledge passages, the FTS index, embeddings, the branch
|
||||
lineage cache rebuilt on import
|
||||
```
|
||||
|
||||
The test separating the last two is not "could this be recomputed" but "would a
|
||||
recomputation answer the same question". A rebuilt FTS index answers the same
|
||||
question. A rebuilt prompt does not — it says what the turn *would be told now*,
|
||||
from today's canon, today's sources and today's state, which is the opposite of
|
||||
what the inspector is for. Historical evidence is not a cache, so M9 does not
|
||||
regenerate one on import at any point.
|
||||
|
||||
`DATA-MODEL.md` §29 lists what v3 carries. Three implementation facts belong
|
||||
here rather than there:
|
||||
|
||||
1. **The snapshots are encoded, not summarised.** A per-turn prompt contains the
|
||||
story so far, so one per turn is O(turns²) — measured at 68% of a 9.7 MB file
|
||||
at 120 turns, against a 20 MB import ceiling. The snapshot therefore travels
|
||||
as `contextSnapshotZ`, zlib-compressed and base64-encoded through the same
|
||||
`compression.pack`/`unpack` the database column already uses. The file is
|
||||
still JSON and every other section of it is still plain text. The readable
|
||||
`contextSnapshot` key is still accepted and wins when both are present, so a
|
||||
hand-edited file keeps importing. A residual ceiling remains and is stated in
|
||||
the M9 report rather than hidden.
|
||||
2. **One more pointer is translated, and only pointers ever are.** Branch
|
||||
numbers already were. A restored retrieval record's `source_id` names a row
|
||||
on the machine that wrote the file, so it is repointed at the source that
|
||||
landed here, or set to `null` when the file carries no such source. The text
|
||||
the record holds — the evidence — is never rewritten.
|
||||
3. **Import stays two-phase inside one transaction.** `plan` refuses everything
|
||||
a hand-edited file can get wrong before a row exists; `materialize` writes,
|
||||
and the endpoint commits once and rolls back explicitly otherwise. A
|
||||
*rebuildable* index failing after that does not roll the campaign back: it is
|
||||
reported on the response as a warning, shown per source in the Knowledge
|
||||
panel, and repaired by Reindex. So a caller sees either "the campaign is not
|
||||
there" or "the campaign is complete", never a third thing.
|
||||
|
||||
### 9.4 The database backup, as implemented in M9
|
||||
|
||||
A second recovery tool, deliberately not merged with the first. The bundle is a
|
||||
logical, portable, human-readable copy of **one campaign** and is the supported
|
||||
way to move a campaign between installations; the backup is a physical copy of
|
||||
**this machine's whole database** and is what you take before an upgrade.
|
||||
|
||||
`backend/app/backup.py` uses SQLite's online backup API rather than a file copy,
|
||||
because a copy taken while the application runs can read one page before a
|
||||
transaction and another after it and produce a file that opens, reports a schema
|
||||
and is quietly missing rows. It writes to a temporary name beside the
|
||||
destination, runs `PRAGMA quick_check` against the finished file, and only then
|
||||
renames it into place; it opens the source read-only, never overwrites an
|
||||
existing backup, and leaves nothing behind on failure.
|
||||
|
||||
No path comes from a caller: the destination is derived from the database the
|
||||
application already has open and the filename from the clock, so the endpoints
|
||||
accept no body at all (H08).
|
||||
|
||||
**There is no restore endpoint, and that is a decision.** Restoring means
|
||||
replacing the file the running process has open, which is how both copies are
|
||||
lost at once. The procedure is in `DEVELOPMENT.md` and is a procedure precisely
|
||||
because each step needs the application stopped.
|
||||
|
||||
## 10. Authoritative Narrative State
|
||||
|
||||
### 10.1 Do not retain the RPG state protocol as the product model
|
||||
@@ -1015,6 +1098,64 @@ microphone/audio -> local STT -> editable draft -> normal user submission
|
||||
|
||||
STT never bypasses the ordinary authoritative story commit path.
|
||||
|
||||
### 15.1 As implemented (M10)
|
||||
|
||||
The boundary above is built. What follows is what it turned out to be.
|
||||
|
||||
**The scene snapshot was already there.** §15 says "persist or derive", and the
|
||||
answer is *derive*, because M5 had persisted it three milestones earlier:
|
||||
`narrative_state["scene"]` holds `summary`, `location`, `present[]` and the
|
||||
`at: {branch_id, depth}` coordinate, written by the validated `set_scene` event
|
||||
and snapshotted per position. It already restores correctly through Undo, Redo,
|
||||
Retry, Save Point restore and divergence, because the head move restores the
|
||||
whole state document and the scene is part of it. A second scene store would
|
||||
have had to reimplement all of that, and would have been a second answer to
|
||||
"where is the story now".
|
||||
|
||||
So the seam is three modules under `app/media/`, and only one of them has a
|
||||
table:
|
||||
|
||||
```text
|
||||
app/media/packet.py the normalized scene packet §15 names, built on read
|
||||
app/media/profiles.py visual character/location profiles — the one field in
|
||||
§15's list that nothing already stored
|
||||
app/media/providers.py the contracts a future coordinator implements
|
||||
```
|
||||
|
||||
`GET /api/adventures/{id}/scene-packet` returns the packet; the four
|
||||
`visual-profiles` endpoints read and write profiles. Nothing else in the
|
||||
application calls either — no turn, no prompt, no context section.
|
||||
|
||||
**What the packet contains, and the rule behind it.** Location, characters
|
||||
present with their profiles, significant objects, an action summary, continuity
|
||||
constraints, ambience, the source turn range and the lineage coordinate. What it
|
||||
deliberately excludes is the more interesting half: the raw transcript, all
|
||||
imported knowledge, memories and summaries. The rule is *what the story
|
||||
established at this position*, not *everything the narrator was told* — which is
|
||||
what keeps a hidden Canon source out of a depiction without needing a filter
|
||||
that someone has to remember to apply to each new secret.
|
||||
|
||||
**Story engine functional with media disabled** is the milestone's central
|
||||
acceptance condition rather than a footnote, and it is tested as one
|
||||
(`backend/tests/test_m10_no_media.py`): a whole campaign — turns, state
|
||||
extraction, memory and summary activity, knowledge retrieval, Undo, Redo, Retry,
|
||||
Save Point restore, restart — with an empty provider registry, no media setting
|
||||
in existence, and no media row written. Media readiness is inert until something
|
||||
uses it, and nothing does yet.
|
||||
|
||||
**STT asymmetry.** The `TranscriptionProvider` contract returns a
|
||||
`DraftTranscription` with `editable: bool = True` and no commit method, so the
|
||||
diagram above is enforced by the shape of the interface: a transcriber can
|
||||
produce a draft and cannot submit one. The ordinary authoritative commit path is
|
||||
the only way in.
|
||||
|
||||
**Endpoint policy.** A future media provider endpoint is checked by
|
||||
`providers.endpoint_rejection_reason`, which reuses the local-only policy in
|
||||
`app/endpoints.py` and then requires loopback in addition — stricter than
|
||||
narrator inference, which permits a trusted LAN host. A GPU that renders a
|
||||
reader's campaign is a machine that reader is sitting at. No provider
|
||||
configuration setting exists to point anywhere, because none is needed yet.
|
||||
|
||||
## 16. Database Direction
|
||||
|
||||
SQLite remains the selected v1 authoritative store.
|
||||
|
||||
@@ -996,3 +996,117 @@ actual persistence / memory / canon bugs
|
||||
The prose may vary.
|
||||
|
||||
The expected state, authority, and lineage rules should not.
|
||||
|
||||
---
|
||||
|
||||
# Appendix A — Proposed companion fixture: Multi-Character Identity Test
|
||||
|
||||
**Status: proposed, not built. This appendix changes nothing above it.**
|
||||
|
||||
## Why a companion rather than an extension
|
||||
|
||||
The Continuity Test above is the deterministic baseline that several milestones'
|
||||
results are compared against. Adding characters or turns to it would invalidate
|
||||
those comparisons, so **it is deliberately left exactly as it is**.
|
||||
|
||||
## The gap this fills
|
||||
|
||||
Reviewed on 2026-09-07 against a hands-on finding. The Continuity Test's seven
|
||||
deliberate traps (§13) are:
|
||||
|
||||
```text
|
||||
1-3 knowledge boundaries and secrets
|
||||
4 reference authority
|
||||
5 canon precedence
|
||||
6 branch leakage
|
||||
7 possession
|
||||
```
|
||||
|
||||
There is **no identity trap**, the word *coreference* does not appear, and the
|
||||
on-stage cast is effectively two people — Aldric and Mara, with Edrin
|
||||
established as missing rather than present. So **same-scene multi-character
|
||||
identity continuity is not exercised anywhere in the standard fixture.**
|
||||
|
||||
A play session against accepted M8 produced exactly that failure: four people in
|
||||
one office, and narration that treated one of them as two different people
|
||||
sharing a name. Root cause is unknown and no longer establishable — the playtest
|
||||
database was destroyed — which is itself part of why a *deterministic* fixture
|
||||
for this class is worth having.
|
||||
|
||||
## Shape
|
||||
|
||||
Four people, all present in one ordinary scene, with no fantasy vocabulary — the
|
||||
point is identity, not genre:
|
||||
|
||||
```yaml
|
||||
bill: { type: character, role: protagonist, controlled_by: reader }
|
||||
alice: { type: character, role: coworker }
|
||||
roger: { type: character, role: coworker }
|
||||
john: { type: character, role: coworker }
|
||||
location: { type: location, name: the office }
|
||||
```
|
||||
|
||||
Identities and roles established unambiguously before the first test turn, so
|
||||
that any later ambiguity is the system's and not the setup's.
|
||||
|
||||
## What the sequence must stress
|
||||
|
||||
- pronouns with more than one plausible referent in scene;
|
||||
- dialogue attribution across three speakers;
|
||||
- characters entering and leaving;
|
||||
- reference by name **and** by role, for the same person;
|
||||
- one character speaking *about* another;
|
||||
- one character speaking about **themself in the third person**, which is the
|
||||
shape the observed failure took.
|
||||
|
||||
## Traps
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **I1 — one Alice** | No turn may produce a second entity whose display name is `Alice`. The state model currently permits this and reports nothing, so the trap is real rather than theoretical |
|
||||
| **I2 — the protagonist stays the protagonist** | Bill must not drift into being narrated as a third party, or acquire a second entity |
|
||||
| **I3 — attribution** | A line spoken by Roger must not be attributed to John |
|
||||
| **I4 — self-reference** | No character may refer to themself as a separate same-named person |
|
||||
| **I5 — state and context agree** | The authoritative state and the assembled prompt must not disagree about who is present or who anyone is |
|
||||
| **I6 — exit and return** | A character who leaves and returns is the same entity, not a new one |
|
||||
|
||||
## Required evidence on failure
|
||||
|
||||
Unlike the fixture above, this one exists to **classify** a failure rather than
|
||||
only to detect it, because its failure modes are split between the application
|
||||
and the model. Any failing turn must capture:
|
||||
|
||||
```text
|
||||
authoritative state immediately before generation
|
||||
the exact stored context/prompt snapshot
|
||||
the recent-history section
|
||||
summaries
|
||||
retrieved memories
|
||||
imported knowledge, if any
|
||||
narrator output
|
||||
model identifier and generation settings
|
||||
```
|
||||
|
||||
then classify:
|
||||
|
||||
```text
|
||||
STATE DEFECT
|
||||
CONTEXT ASSEMBLY DEFECT
|
||||
DERIVED MEMORY/SUMMARY DEFECT
|
||||
MODEL FAILURE WITH CORRECT CONTEXT
|
||||
AMBIGUOUS / MULTIPLE CONTRIBUTORS
|
||||
```
|
||||
|
||||
Two rules for whoever runs it: **do not "fix" a model failure by editing
|
||||
authoritative state**, and **do not blame the model when the prompt already
|
||||
contained the identity error.**
|
||||
|
||||
Since M9 the whole of that evidence is portable in one campaign bundle, so a
|
||||
failing run can be exported intact and investigated elsewhere.
|
||||
|
||||
## Ownership
|
||||
|
||||
M11, alongside the realistic-model review. `V1-ACCEPTANCE-TESTS.md` §P3 records
|
||||
which parts of this are candidate **acceptance** criteria — the state-level
|
||||
traps, which are decidable — and which are model-quality observations that
|
||||
belong in a recorded review rather than in the pass/fail contract.
|
||||
|
||||
@@ -1800,6 +1800,16 @@ Export standard campaign.
|
||||
### Pass
|
||||
Export completes locally and contains enough data to restore story.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
`ai-dnd-adventure-v3`, checked section by section rather than by file size:
|
||||
the story and its whole retained tree, the branches and their disposition, the
|
||||
chosen head, Save Points, the authoritative state and its per-position
|
||||
snapshots, the state events and proposals, the imported library, the
|
||||
lineage-anchored summaries, the memories, and a stored prompt for every narrator
|
||||
turn that has one. `test_m9_portability.py::test_i01_*`, and reproducible with
|
||||
`python -m tools.m9_portability_report`, which classifies every data family as
|
||||
PRESERVED, OMITTED or DERIVED/REBUILDABLE.
|
||||
|
||||
---
|
||||
|
||||
## I02 — Import Exported Campaign
|
||||
@@ -1814,6 +1824,14 @@ Export completes locally and contains enough data to restore story.
|
||||
### Pass
|
||||
Active transcript and state are restored.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Across a **genuine clean data directory**: a second server process, in a second
|
||||
directory, against a database file that has never existed, with the exporting
|
||||
process stopped. Transcript, authoritative state, canon, Save Points, knowledge,
|
||||
state events and the size of the retained tree all match the source campaign
|
||||
(`test_m9_clean_import.py`). The same round trip inside one process is in
|
||||
`test_m9_portability.py`, and is labelled there as the weaker of the two.
|
||||
|
||||
---
|
||||
|
||||
## I03 — Branch/Disposable History Export
|
||||
@@ -1823,6 +1841,19 @@ Active transcript and state are restored.
|
||||
### Pass
|
||||
Export preserves retained alternate/disposable history needed for recovery, unless user explicitly chooses a trimmed export.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Both futures come back and stay distinguishable: the abandoned line's turns are
|
||||
present and readable, the branch that was left carries the depth it was left at,
|
||||
and a superseded take is still at its coordinate with the selected take still
|
||||
selected. There is no trimmed export — M9 offers no option to drop history, so
|
||||
the clause about one does not arise.
|
||||
|
||||
**M9 also closed a fidelity gap here.** Take parentage was not exported, so every
|
||||
imported node landed parentless and the pager grouped on the coordinate instead.
|
||||
That is right for a plain retry and wrong once two takes of one turn each have
|
||||
takes of their own beneath them; the copy read `5/5` where the source read `2/2`
|
||||
and `3/3`. v3 carries the parentage.
|
||||
|
||||
---
|
||||
|
||||
## I04 — Checkpoint Export
|
||||
@@ -1841,6 +1872,15 @@ campaign. Importing Save Points does **not** move the active head — the head
|
||||
still comes from the bundle's `headDepth`. Bundles written before M4 carry no
|
||||
`checkpoints` key, import cleanly, and create none.
|
||||
|
||||
### Re-verified — PASS (M9, 2026-09-07)
|
||||
Unchanged by the format bump, and extended in two directions. Every restored
|
||||
Save Point resolves, restores to the position it names through M3's head
|
||||
movement, leaves the retained history it moved back over intact, and the two
|
||||
in the M9 fixture restore to *different* states. A Save Point whose coordinate
|
||||
is not in the file is **dropped with the rest of the campaign kept**, never
|
||||
retargeted to a nearby turn: the reader named a position, and if that position
|
||||
is not in the file then no other position is the one they named.
|
||||
|
||||
---
|
||||
|
||||
## I05 — Knowledge Provenance Export
|
||||
@@ -1870,13 +1910,29 @@ hand-edited knowledge block with an unknown classification or empty content
|
||||
refuses the import rather than half-landing in it; and an edited content hash is
|
||||
recomputed from what actually arrived and the discrepancy recorded on the source.
|
||||
|
||||
**Limit, unchanged from before M7 and owned by M9.** The bundle carries no
|
||||
context snapshots at all, so an imported campaign has no historical prompt
|
||||
provenance — for imported knowledge or for any other component. Nothing M7
|
||||
creates is turned into a dangling id by a round trip, because no ids are
|
||||
exported; the evidence simply is not in the file.
|
||||
`test_historical_prompt_evidence_survives_an_export_round_trip` pins that
|
||||
behaviour so it cannot regress silently.
|
||||
**That limit is closed (M9, 2026-09-07).** The paragraph below is kept as
|
||||
written because it records what was true through M7 and M8, and the M9 decision
|
||||
is only legible against it.
|
||||
|
||||
> *The bundle carries no context snapshots at all, so an imported campaign has no
|
||||
> historical prompt provenance — for imported knowledge or for any other
|
||||
> component. Nothing M7 creates is turned into a dangling id by a round trip,
|
||||
> because no ids are exported; the evidence simply is not in the file.*
|
||||
|
||||
M9 carries the snapshots. An old turn in a restored campaign shows the prompt it
|
||||
was actually assembled from, the passages it was shown and the text each
|
||||
supplied — after the source has been deleted, the canon edited and the state
|
||||
moved on. `test_historical_prompt_evidence_survives_an_export_round_trip` was
|
||||
inverted rather than deleted: it now pins the thing M7 was worried about and
|
||||
could not check, which is that the provenance arriving on the other side names
|
||||
*this* campaign's sources rather than the ids they had where the file was
|
||||
written. Only that pointer is translated; the evidence is restored verbatim, and
|
||||
a source the file does not carry becomes `null` rather than pointing at a
|
||||
different file.
|
||||
|
||||
Also added in M9: `parserVersion` and `chunkingVersion` per source, recording
|
||||
what produced the passages a historical retrieval record describes, and a
|
||||
`sourceId` that exists only so the translation above can be made.
|
||||
|
||||
---
|
||||
|
||||
@@ -1887,6 +1943,28 @@ behaviour so it cannot regress silently.
|
||||
### Pass
|
||||
No external API credentials are embedded in campaign export.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Tested rather than assumed, and tested against the file's **text** rather than
|
||||
against a list of columns — a field added to a model the exporter walks would
|
||||
otherwise reach the bundle with no test of a column noticing. The inert
|
||||
`api_key` column is written with a recognisable value first, so its absence is
|
||||
evidence rather than a tautology.
|
||||
|
||||
Nothing matching `api_key`, `apiKey`, the written secret, the inference endpoint
|
||||
or its port, or an absolute filesystem path appears anywhere in an export —
|
||||
including inside the compressed snapshots, which the test decodes rather than
|
||||
skipping. Checked in both suites, so the clean-directory run covers the same
|
||||
ground across a real process boundary
|
||||
(`test_m9_portability.py::test_i06_*`, `test_m9_clean_import.py`).
|
||||
|
||||
**What deliberately does not travel**, and why it is not an omission: the
|
||||
inference endpoint, the model name, the context budget and every other row of
|
||||
`settings`. Those describe the machine, not the campaign, and importing a
|
||||
campaign must not silently repoint the destination's inference at the source's.
|
||||
Per-turn model and generation settings *do* travel, inside the historical
|
||||
snapshot, because there they are a record of what happened rather than a
|
||||
configuration to apply.
|
||||
|
||||
---
|
||||
|
||||
## I07 — Export/Import Preserves an Undone Active Head
|
||||
@@ -1921,6 +1999,28 @@ An export whose stated head lies beyond the story it contains is a file
|
||||
disagreeing with itself and must be refused rather than opened at a guessed
|
||||
position.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
All three clauses, and the main one across a genuine machine boundary.
|
||||
|
||||
- **The exact head.** The M9 fixture ends two Undos behind its own branch's
|
||||
retained tip and behind the abandoned line's, and the last thing it does is an
|
||||
Undo — so the head is not the newest row written, not the deepest row, not the
|
||||
tip, and not on the branch holding the most story. An importer guessing any one
|
||||
of those lands somewhere else. The copy opens exactly where the source was,
|
||||
the later turns are still in the database as retained future, and Redo is
|
||||
offered rather than the story having silently been redone. Redo then walks to
|
||||
the same next turn in both.
|
||||
- **The legacy clause.** A bundle with its `headDepth` removed opens at the tip
|
||||
of its head branch, offers no Redo, and offers Undo — which is the position
|
||||
such a file recorded, because at the time it was written the head could not be
|
||||
anywhere else. Checked at every seam in `test_m9_legacy_bundles.py`.
|
||||
- **The self-disagreeing file.** A head past the retained story is refused with
|
||||
a message naming where the branch actually ends, and nothing is written.
|
||||
|
||||
Evidence: `test_m9_clean_import.py::test_it_opens_at_the_exact_head_it_was_exported_at`
|
||||
(second process, empty directory), plus `test_m9_portability.py::test_i07_*` and
|
||||
the head cases in `test_m9_corrupt_bundles.py`.
|
||||
|
||||
---
|
||||
|
||||
# J. Genre Independence
|
||||
@@ -1984,6 +2084,24 @@ Reach a scene involving multiple characters and a clear location.
|
||||
### Pass
|
||||
Application can persist a structured scene representation sufficient for future media use.
|
||||
|
||||
### Result — PASS, and it was already passing (M10, 2026-09-07)
|
||||
|
||||
The scene representation is `narrative_state["scene"]` and has existed since M5:
|
||||
`summary`, `location`, `present[]`, and the `at: {branch_id, depth}` coordinate
|
||||
that says where it was written. It is produced by the validated `set_scene`
|
||||
event, snapshotted per position in `actions.narrative_state_after`, restored on
|
||||
every head move, and carried in the v3 bundle.
|
||||
|
||||
M10 verified this with a probe rather than trusting the earlier report — a
|
||||
campaign played to a scene, diverged, undone and exported — and then built the
|
||||
normalized packet on top of it (`app/media/packet.py`,
|
||||
`GET /api/adventures/{id}/scene-packet`) instead of a second scene store.
|
||||
|
||||
Tests: `backend/tests/test_m10_media_hooks.py` (the packet's contents and
|
||||
bounds), `test_m10_lineage.py` (the scene follows the active lineage through
|
||||
Undo, Redo, Retry, divergence, Save Point restore and two genuine process
|
||||
restarts).
|
||||
|
||||
---
|
||||
|
||||
## K02 — Visual Character Profile
|
||||
@@ -1993,6 +2111,22 @@ Application can persist a structured scene representation sufficient for future
|
||||
### Pass
|
||||
Character can retain optional stable visual descriptors.
|
||||
|
||||
### Result — PASS (M10, 2026-09-07)
|
||||
|
||||
`visual_profiles`, keyed by `(adventure_id, entity_key)` — the M5 entity key, not
|
||||
a new identity namespace — with an open `descriptors` map, a `features` list and
|
||||
free `style_notes`. Read and written through
|
||||
`/api/adventures/{id}/visual-profiles[/{entity_key}]`.
|
||||
|
||||
**Optional** is tested as well as stated: the fixture leaves one character
|
||||
deliberately unprofiled, and the packet reports `visual_profile: null` for them
|
||||
rather than an empty profile, because "nobody has decided what Roger looks like"
|
||||
and "Roger looks like nothing" are different answers to a future provider.
|
||||
|
||||
**Stable** means campaign-scoped rather than per-position: a character does not
|
||||
change appearance because the story forked, so a profile survives Undo, Redo,
|
||||
divergence and Save Point restore unchanged, and travels in the bundle.
|
||||
|
||||
---
|
||||
|
||||
## K03 — Visual Location Profile
|
||||
@@ -2002,6 +2136,19 @@ Character can retain optional stable visual descriptors.
|
||||
### Pass
|
||||
Location can retain optional visual continuity descriptors.
|
||||
|
||||
### Result — PASS, through the same mechanism as K02 (M10, 2026-09-07)
|
||||
|
||||
There is no separate location table. M5's entity model is genre-neutral and a
|
||||
location is an entity with a `type`, so one `visual_profiles` table serves
|
||||
characters, locations and items alike. Splitting them would have reintroduced the
|
||||
genre shape M5 spent a milestone removing, and a `kind` column would have been a
|
||||
second copy of the entity's own type.
|
||||
|
||||
The fixture is deliberately non-fantasy — an open-plan office under flat
|
||||
fluorescent light — because the contract's examples are fantasy-shaped and a
|
||||
schema written while looking at them acquires that shape without anyone choosing
|
||||
it.
|
||||
|
||||
---
|
||||
|
||||
## K04 — Attach Media Asset to Scene
|
||||
@@ -2016,6 +2163,27 @@ A local dummy/test image can be associated with a scene/turn without altering st
|
||||
If media tables are deferred:
|
||||
- architecture/types should demonstrate equivalent extension point.
|
||||
|
||||
### Result — PASS on the deferred branch (M10, 2026-09-07)
|
||||
|
||||
Media tables are **deferred deliberately**, so this is reported against the
|
||||
acceptance text's second clause rather than its first.
|
||||
|
||||
`app/media/providers.py` defines the extension point as structural contracts:
|
||||
`MediaRequest`, `MediaResult`, `ProviderCapabilities`, `DraftTranscription`, and
|
||||
the `MediaProvider` / `SpeechProvider` / `TranscriptionProvider` protocols, with
|
||||
a registry that is empty and stays empty. A test registers a dummy provider,
|
||||
builds a scene packet, generates a fake PNG carrying the packet's `scene_id` as
|
||||
provenance, and confirms the story model is byte-for-byte unchanged — which is
|
||||
the equivalence the clause asks for.
|
||||
|
||||
`media_jobs` and `media_assets` were not built because a queue with no producer
|
||||
and no consumer would be speculative architecture, shaped by a provider nobody
|
||||
has chosen; M6 declined the same thing (`derived_status`: "not a job queue").
|
||||
The part that would be expensive to retrofit is guaranteed now: the scene
|
||||
identity a future asset must reference is *derived* from campaign and position
|
||||
(`c<adventure>:b<branch>:<start>-<end>`), so it is stable across processes and
|
||||
restarts without a row to keep in step.
|
||||
|
||||
---
|
||||
|
||||
## K05 — Generate Local Image
|
||||
@@ -2082,6 +2250,17 @@ turn.
|
||||
### Pass
|
||||
State at each position matches original accepted state.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07)
|
||||
Measured **after a round trip**, which is the M9 form of it: the copy and the
|
||||
source are walked back three turns and forward three turns in step, and the
|
||||
authoritative document is compared at every position. They agree throughout.
|
||||
|
||||
That this stays a snapshot read rather than a replay is the point. Undo, Redo
|
||||
and Save Point restore all resolve a coordinate and read the state recorded
|
||||
there (`TECHNICAL-DESIGN.md` §10.4), so an import that carried the events and
|
||||
dropped the per-position snapshots would have made every one of them
|
||||
proportional to campaign length. Both halves of §17's hybrid travel.
|
||||
|
||||
---
|
||||
|
||||
## L03 — Checkpoint Reconstruction After Restart
|
||||
@@ -2105,6 +2284,15 @@ exactly that value after the campaign had been advanced past it
|
||||
`TestClient` restart, which could not distinguish durable state from a live
|
||||
object.
|
||||
|
||||
### Re-verified after a move — PASS (M9, 2026-09-07)
|
||||
The same claim with a machine boundary in front of it. A campaign is exported
|
||||
from one server process, imported into a **second process against a database
|
||||
file that has never existed**, a Save Point is restored there, that process is
|
||||
killed, and a **third** process against the same file is asked again. The
|
||||
transcript, the authoritative state and the size of the retained tree all match
|
||||
what the second process had after restoring
|
||||
(`test_m9_clean_import.py::test_l03_*`).
|
||||
|
||||
---
|
||||
|
||||
## L04 — Derived Data Can Be Rebuilt
|
||||
@@ -2121,6 +2309,34 @@ using a safe test copy.
|
||||
### Pass
|
||||
Authoritative campaign history remains intact and derived structures can be recreated.
|
||||
|
||||
### Result — PASS (M9, 2026-09-07), and one defect found by running it
|
||||
On a safe copy — an imported campaign, not the original. Every physically
|
||||
derived structure is destroyed and rebuilt from the source content the bundle
|
||||
carried: passages, the FTS rows and the vectors. Afterwards every source is
|
||||
`ready` with passages again, retrieval works, and the transcript, the
|
||||
authoritative state, the classifications and the lifecycle flags are identical
|
||||
either side. Deleting the vectors alone leaves lexical retrieval working, which
|
||||
is M7's rule that the lexical half is a production path and not a fallback. A
|
||||
rebuild does not make an abandoned line's summary eligible.
|
||||
|
||||
**Running it found a real defect, which is fixed here.** The FTS5 index is a
|
||||
virtual table, so no foreign key reaches it and no `ON DELETE CASCADE` covers
|
||||
it: deleting a campaign dropped its passages and left one index row per passage
|
||||
behind. Nothing read them — every search joins through `knowledge_chunks` — so
|
||||
the leak was invisible until SQLite handed the freed primary key out again, at
|
||||
which point the **next source imported into any campaign** failed with an
|
||||
integrity error. Reindex could not repair it either, because `clear_index` finds
|
||||
index rows *through* the chunks, and there were none. Both ends are closed: a
|
||||
campaign's index rows are removed before it is deleted, and the index insert now
|
||||
replaces a stale row rather than colliding with it — so a database already
|
||||
carrying the leak repairs itself and needs no migration. See the M9 report's
|
||||
findings.
|
||||
|
||||
The rebuild path M9 depends on is therefore implemented and tested rather than
|
||||
assumed, which the SHOULD priority above did not require but the milestone did:
|
||||
knowledge passages and indexes are omitted from the bundle precisely because
|
||||
they can be rebuilt.
|
||||
|
||||
---
|
||||
|
||||
# M. Long-Run Test
|
||||
@@ -2327,3 +2543,92 @@ L01-L03
|
||||
```
|
||||
|
||||
This keeps implementation work tied to observable behavior rather than repository-specific architecture.
|
||||
|
||||
---
|
||||
|
||||
# P. M11 Test-Design Tasks — Not Yet Acceptance Tests
|
||||
|
||||
**Nothing in this section is part of the pass/fail contract.** These are three
|
||||
behaviours a real play session against accepted M8 showed are worth testing, for
|
||||
which **the pass criterion is not yet settled**. They are recorded here so the
|
||||
release tester finds them where they will look, and they are deliberately not
|
||||
written as A-through-O items: giving them IDs would imply a criterion has been
|
||||
ratified when it has not, and would destabilise a document whose value is that
|
||||
every item in it is decidable.
|
||||
|
||||
Each names **what must be settled** before it can become an acceptance test.
|
||||
Full context is in the M9 report's §Y and, durably, in `BUILD-MILESTONES.md`
|
||||
under the post-M8 playtest findings.
|
||||
|
||||
## P1 — The reader can tell where they are after history movement
|
||||
|
||||
**From:** a real session in which Undo worked correctly and the reader could not
|
||||
tell which point in the story they had reached.
|
||||
|
||||
**Behaviour to test:** after Undo, Redo, a Save Point restore, or an edit to an
|
||||
earlier turn, the reader can identify their current position in the visible
|
||||
story without implementation terminology (`branch`, `head`, `node`, `depth`
|
||||
remain forbidden at the surface).
|
||||
|
||||
**Settle first:** what the indicator *is*. "The reader can tell" is not
|
||||
decidable as written — it needs an observable artifact, such as a named position
|
||||
that changes with movement and is present in the DOM. `BROWSER-UX-SPEC.md` §8
|
||||
now carries the requirement; **no wording is ratified**, and this cannot become
|
||||
an acceptance test before one is.
|
||||
|
||||
## P2 — Narration length has a measurable directional effect
|
||||
|
||||
**From:** a reader who chose *2-4 paragraphs* and received substantially longer
|
||||
replies.
|
||||
|
||||
**Behaviour to test:** the narration-length setting produces a **measurable
|
||||
directional difference** in output length across repeated realistic turns —
|
||||
brief shorter than standard, standard shorter than detailed — on at least the
|
||||
reference 3B narrator and one stronger local narrator.
|
||||
|
||||
**Settle first:** the numbers, and the mechanism. Today the setting adds one
|
||||
English sentence to the campaign instructions and changes **no** generation
|
||||
budget, while a separate numeric hint derived from the global
|
||||
`max_output_tokens` is identical for every setting (M9 report §Y). Until the
|
||||
product decides what each setting *means* — and whether it moves the budget —
|
||||
there is no threshold to test against. A directional test is stateable; an
|
||||
absolute one is not, and this should not become an acceptance test that asserts
|
||||
word counts nobody has ratified.
|
||||
|
||||
**Do not** turn this into a truncation test: the state block is emitted last and
|
||||
hard truncation removes it.
|
||||
|
||||
## P3 — Multi-character identity continuity
|
||||
|
||||
**From:** four people in one scene, and narration that treated one of them as
|
||||
two different people of the same name. **Root cause unknown** — the playtest
|
||||
database was destroyed, so no evidence survives.
|
||||
|
||||
**Behaviour to test:** across a multi-turn scene with a protagonist and three
|
||||
supporting characters, the story does not create duplicate characters, does not
|
||||
duplicate a display name across two entities, does not drift the protagonist's
|
||||
identity, does not misattribute dialogue, and does not have a character refer to
|
||||
themself as a separate same-named character — and the authoritative state and
|
||||
the assembled context do not disagree about who anyone is.
|
||||
|
||||
**Settle first:** which of those are **product** guarantees and which are
|
||||
**model-quality** observations. They are not the same kind of claim and must not
|
||||
share one verdict:
|
||||
|
||||
- *"The state never holds two entities with the same display name"* is
|
||||
decidable and enforceable, and is a candidate acceptance test today. The
|
||||
implementation currently permits it and reports nothing (M9 report §Y).
|
||||
- *"The narrator never confuses two same-named characters"* is not a pass/fail
|
||||
property of this application — it depends on the model — and belongs in M11's
|
||||
realistic-model review with a recorded classification, not in this contract.
|
||||
|
||||
**A run of this must capture**, on any failure: pre-generation state, the exact
|
||||
stored prompt snapshot, history, summaries, retrieved memories, imported
|
||||
knowledge, narrator output, and model settings — then classify as a state,
|
||||
context-assembly, derived-data, or model failure. M9 made all of that portable,
|
||||
so a failing campaign can be exported whole and investigated elsewhere.
|
||||
|
||||
**Fixture:** the standard Continuity Test does not exercise this — its traps are
|
||||
knowledge, authority, branch leakage and possession, and its on-stage cast is
|
||||
effectively two people. A companion fixture is proposed in
|
||||
`TEST-CAMPAIGN-FIXTURE.md`; the established fixture is deliberately unchanged.
|
||||
|
||||
+133
-3
@@ -1,8 +1,138 @@
|
||||
# Planning Package Version
|
||||
|
||||
- **Package:** Adventure Storyteller Planning Package v3.3
|
||||
- **Revision date:** 2026-09-06
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted** (M8 closed out 2026-09-06). M9 is next and has not been started.
|
||||
- **Package:** Adventure Storyteller Planning Package v3.6
|
||||
- **Revision date:** 2026-09-07
|
||||
- **Status:** Phase 0 complete; architecture selected; **Milestones M1-M8 implemented and accepted**; **M9 and M10 implemented and awaiting independent review** (2026-09-07). M11 has not been started.
|
||||
|
||||
## v3.6 — M10 implemented: future media extension hooks only (2026-09-07)
|
||||
|
||||
M10 built the media seam and no media. The documentation change is mostly a
|
||||
record of what was deliberately **not** built, because that is the part a later
|
||||
reader will otherwise re-litigate.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `DATA-MODEL.md` | **New §28A**, the media extension points as implemented: one `visual_profiles` table, no scenes table, no job or asset tables, and why each. | as-implemented record |
|
||||
| `TECHNICAL-DESIGN.md` | **New §15.1** under the Scene and Future Media Boundary: what the seam turned out to be, what the packet carries and excludes, the STT asymmetry, the endpoint policy. | as-implemented record |
|
||||
| `MEDIA-EXTENSION-CONTRACT.md` | **New §90**, appended rather than woven in so the Phase 0B contract stays readable. Records the three places implementation answered an open question, and the sections left unbuilt. | as-implemented record |
|
||||
| `BUILD-MILESTONES.md` | **M10 status block**: the finding that shaped the milestone, what shipped, and the migration defect its own tests caught. | milestone status |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | **K01-K04 results.** K01 passes and *was already passing*; K04 is reported on the acceptance text's deferred branch, as its own wording provides for. | acceptance evidence |
|
||||
| `SECURITY-THREAT-MODEL.md` | **New §42A.** The trust boundary did not widen; in one place the implementation is deliberately **narrower** than §73 permits, which is still a discrepancy worth recording. | boundary note |
|
||||
| `README.md`, `DEVELOPMENT.md` | The `media/` package in the architecture map, the backend test count (1,189), and the stricter rule a future media endpoint will meet. | developer docs |
|
||||
| `reports/M10-IMPLEMENTATION-REPORT.md` | New. The implementer's account, written for a reviewer. | milestone report |
|
||||
|
||||
**The decision behind the whole milestone**: the scene snapshot the media
|
||||
contract asks for **already existed** as `narrative_state["scene"]`, built by M5
|
||||
and carrying lineage correctly since. Verified with a probe rather than taken
|
||||
from an earlier report. So M10 added no scenes table, and derives the scene
|
||||
packet on read.
|
||||
|
||||
**Two decisions recorded rather than assumed:**
|
||||
|
||||
- **The bundle format stays `ai-dnd-adventure-v3`.** M9's own semantic test —
|
||||
does omission create ambiguity about what an older file *could* have recorded?
|
||||
— says no: a campaign with no visual profiles is the ordinary case, so an
|
||||
absent key unambiguously means "none". Older v3 files still import.
|
||||
- **M10 adds no migration.** A `CREATE INDEX` migration was written first and
|
||||
removed: `create_all` already builds the table and the index declared on its
|
||||
column, so the migration left an upgraded database holding an index a fresh
|
||||
install did not have. `LATEST_VERSION` stays 92.
|
||||
|
||||
**The four post-M8 playtest findings below remain M11's** and are untouched by
|
||||
this entry.
|
||||
|
||||
## v3.5 — Post-M8 hands-on playtest findings recorded (2026-09-07)
|
||||
|
||||
**Documentation only. No application code changed, and M9's verified result is
|
||||
untouched** — see the note at the end of this entry.
|
||||
|
||||
A real play session against **accepted, signed M8** (real browser, trusted-LAN
|
||||
Ollama, `qwen2.5:3b-instruct-16k`, disposable database since destroyed) surfaced
|
||||
four product-quality observations. **None is an M9 defect, none was caused by
|
||||
M9, and none blocks M9 acceptance.** They are recorded so they cannot be lost
|
||||
when the M9 report is archived.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `reports/M9-IMPLEMENTATION-REPORT.md` | **New §Y**, the full write-up with the verification behind each mechanism, plus a pointer from §W. Explicitly labelled as not-M9. | observation record |
|
||||
| `BUILD-MILESTONES.md` | **New section before M10**, the durable copy owned by M11; a note bounding M10 out of it; the four items added to M11's scope list. | milestone sequencing |
|
||||
| `V1-ACCEPTANCE-TESTS.md` | **New §P**, three items as **test-design tasks, explicitly not acceptance tests**, each naming what must be settled before it could become one. No existing test changed or weakened. | test design |
|
||||
| `BROWSER-UX-SPEC.md` | **New §8A.** §8's *"The current endpoint should be clear"* was satisfied while a real reader was lost, so it could not hold the behaviour. States the orientation requirement; **prescribes no wording**. | requirement clarification |
|
||||
| `TEST-CAMPAIGN-FIXTURE.md` | **New Appendix A** proposing a companion `Multi-Character Identity Test`. The established deterministic fixture is **unchanged** — altering it would invalidate earlier milestones' comparisons. | test design |
|
||||
|
||||
**What was verified rather than assumed**, since the playtest campaign no longer
|
||||
exists and root causes largely cannot be proven:
|
||||
|
||||
- The browser title genuinely is `AI D&D` (`frontend/index.html`), and **no
|
||||
accepted document ever claimed otherwise** — so this is an uncovered gap, not
|
||||
documentation needing correction. Nothing was corrected and nothing was
|
||||
renamed: *Adventure Storyteller* is itself narrower than the genre-agnostic
|
||||
engine `SPECIFICATION.md` requires, and the naming decision is the owner's.
|
||||
- The narration-length setting adds **one English sentence** and changes **no
|
||||
generation budget**, while the numeric hint derived from the global
|
||||
`max_output_tokens` is **identical for every setting** — measured at the
|
||||
default as *"must not exceed 506 words, and it should not stop short of about
|
||||
177."* Recorded as a mechanism to check first, **not** as the proven cause.
|
||||
- The narrative state **permits two entities to share a display name and
|
||||
reports nothing** — `DUPLICATE_ENTITY` rejects a repeated key only. That is
|
||||
one of the identity finding's failure modes; it establishes nothing about what
|
||||
actually happened.
|
||||
|
||||
**M9 is unaffected.** No application file changed in this pass, so M9's final
|
||||
verified result stands exactly as recorded: **1,102 backend passed / 14 skipped
|
||||
/ 0 failed**, 145 frontend, 36/36 browser. The expensive M9 suites were
|
||||
deliberately **not** re-run, because only Markdown changed.
|
||||
|
||||
## v3.4 — M9 Implementation (2026-09-07)
|
||||
|
||||
M9 — Export, Backup, Recovery, and Migration Hardening — is implemented on
|
||||
`m9-recovery` from the signed M8 commit `1ce9972`. This revision records what the
|
||||
implementation settled. It is **not** an acceptance: the milestone report is
|
||||
written for a reviewer and the tree is staged for the repository owner's signed
|
||||
commit.
|
||||
|
||||
**Requirement corrections:** none. M9 altered no product requirement.
|
||||
`SPECIFICATION.md` and `SECURITY-THREAT-MODEL.md` are unchanged — §16 already
|
||||
required the export to preserve the exact active position, §6.1 already required
|
||||
an exact prompt/context snapshot per turn, and M9 implements both rather than
|
||||
redefining either.
|
||||
|
||||
| Document | Change | Kind |
|
||||
| --- | --- | --- |
|
||||
| `DATA-MODEL.md` §29 | The v3 format, the three-category rule (chosen / evidence / rebuildable), what each version can be trusted to say, the encoding of the snapshots, and the two pointers the import translates. | implementation fact |
|
||||
| `TECHNICAL-DESIGN.md` §9.3 | Why the version was bumped when §9.1 and §9.2 each correctly declined one; the third data category; the two-phase transaction and the warning path for a failed derived rebuild. | implementation fact |
|
||||
| `TECHNICAL-DESIGN.md` §9.4 | **New.** The SQLite backup: the online backup API rather than a file copy, the verify-then-rename order, and why there is no restore endpoint. | implementation fact |
|
||||
| `IMPORTED-KNOWLEDGE-DESIGN.md` §73 | **New subsection.** Story Cards settled as compatibility-only legacy data and removed from the narrator's prompt, with the evidence that they were the "alternate untracked path" §73 already forbade. | newly settled design decision |
|
||||
| `V1-ACCEPTANCE-TESTS.md` I01-I07, L02-L04 | Results recorded. I05's M7-era limit is marked closed with the original paragraph kept, because the M9 decision is only legible against it. L04 records the defect running it found. | implementation fact |
|
||||
| `BUILD-MILESTONES.md` M9 | Marked complete, with what it delivered, the three defects it found, the story-card decision, and the debt carried forward. | implementation fact |
|
||||
| `README.md`, `VERSION.md` | Status. | implementation fact |
|
||||
| `DEVELOPMENT.md` | **New section**: the two recovery tools and when each applies, taking a backup, and the stop-move-start restore procedure. Plus a note that an imported long campaign meets a small context ceiling on its first turn rather than gradually. | implementation fact |
|
||||
|
||||
**The four M8 handoff questions, answered**
|
||||
|
||||
| | Answer |
|
||||
| --- | --- |
|
||||
| **A. Complete campaign portability** | Every family travels and is measured family by family, before and after, by a tool a reviewer can rerun. |
|
||||
| **B. Historical prompt provenance** | **It belongs in the bundle, and it is in it.** An old turn in a restored campaign shows what it was actually given, after the source has been deleted and the canon edited. |
|
||||
| **C. Legacy story cards** | Compatibility-only. Carried in both directions; removed from the narrator's prompt; still the summariser's character roster. |
|
||||
| **D. Context-window portability** | The campaign travels; the machine's model configuration does not. Importing changes no setting of the destination's, and a campaign imports whether or not any model is installed. The window itself remains M11's. |
|
||||
|
||||
**What the implementation found rather than assumed**
|
||||
|
||||
Three defects, all found by running the milestone's own tests rather than by
|
||||
reading: an FTS index leak that made an ordinary import fail in an unrelated
|
||||
campaign and that Reindex could not repair; an imported node with no state
|
||||
snapshot being stamped with the campaign's *head* state; and a snapshot pointer
|
||||
that was not being translated because the code mutated a dict in place. The
|
||||
first two predate M9.
|
||||
|
||||
One measurement changed a plan, and then corrected the conclusion drawn from it.
|
||||
Carrying per-turn prompts looked like it would halve the length of campaign that
|
||||
can be restored. Measured, and after compressing them inside the file, everything
|
||||
M9 added costs **12%** of reachable campaign length — the import ceiling moves
|
||||
from about 318 turns to about 279, against a 100-turn certification target. The
|
||||
dominant cost is not M9's at all: the **per-position narrative state document is
|
||||
74% of a bundle**, and v2 already carried it.
|
||||
|
||||
## v3.3 — M8 Closeout (2026-09-06)
|
||||
|
||||
|
||||
@@ -0,0 +1,855 @@
|
||||
# M10 — Future Media Extension Hooks Only
|
||||
|
||||
**Implementation report, written for an independent reviewer.**
|
||||
|
||||
Branch `m10-media-hooks`, from the signed M9 commit `44edece`. Implemented
|
||||
2026-09-07. This is a set of claims with the evidence attached; it is not a
|
||||
record of acceptance.
|
||||
|
||||
**The one-sentence version:** the scene snapshot the media contract asks for
|
||||
already existed, built by M5, so M10 built the seam around it and no media —
|
||||
one table, a packet derived on read, provider contracts with an empty registry,
|
||||
no dependency, no network, and no change to a single reader-facing surface.
|
||||
|
||||
---
|
||||
|
||||
## A. Repository baseline
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| **M9 base commit** | `44edece67e7f65bacf78010ef56f98c3c8864073` — *"M9: a campaign you can actually get back"* |
|
||||
| **Signature** | `git verify-commit 44edece` → **Good signature**, RSA key `02C9BF7D8A4A77DF7A8905617D8AE19DB5C68569`, "JesseMarkowitz", trust `[ultimate]`. `%G?` = `G`. |
|
||||
| **Branch** | `m10-media-hooks`, created from that commit. The working tree was clean at the start. |
|
||||
| **Upstream ancestry** | `upstream` = `https://github.com/parththakkar106/AI-DnD.git`. `git merge-base --is-ancestor d72f7c1b HEAD` → true: the fork point is still an ancestor, so this remains a fork rather than a rewrite. |
|
||||
| **License / provenance** | `LICENSE` unchanged — md5 `07fde30437134836e2ee875e82a7cd31`, still MIT, still "Copyright (c) 2026 Parth Thakkar". `PROVENANCE.md` unchanged: M10 adds only new files under `backend/app/media/`, one router, one model and tests. **No dependency was added** to `requirements.txt` or `package.json`. |
|
||||
|
||||
---
|
||||
|
||||
## B. Architecture implemented
|
||||
|
||||
### B.1 The finding that decided the shape of the milestone
|
||||
|
||||
`MEDIA-EXTENSION-CONTRACT.md` §5 asks for a persisted or derived scene snapshot
|
||||
carrying location, participants, objects, actions, ambience, profiles,
|
||||
continuity constraints, and a turn range with lineage.
|
||||
|
||||
**Most of it already existed, and had since M5.** `narrative_state["scene"]`
|
||||
holds:
|
||||
|
||||
```json
|
||||
{"summary": "...", "location": "office", "present": ["bill", "alice", "roger"],
|
||||
"at": {"branch_id": 3, "depth": 4}}
|
||||
```
|
||||
|
||||
written by the validated `set_scene` typed event, snapshotted per position in
|
||||
`actions.narrative_state_after`, restored by `attempts.restore_state` on every
|
||||
head move, and carried in the M9 v3 bundle.
|
||||
|
||||
That was **verified rather than inherited from M9's report**: a probe played a
|
||||
campaign to a scene ("Mara enters the cellar" at branch 1, depth 4), undid it and
|
||||
confirmed the scene cleared to `{}`, diverged to a second continuation ("Mara
|
||||
remains upstairs" at branch 3, depth 4), confirmed both were retained and
|
||||
distinguishable, and confirmed a bundle carried both.
|
||||
|
||||
So **M10 created no scenes table.** A second scene store would have been a
|
||||
duplicate representation of the same fact, with its own lineage rules to get
|
||||
wrong — and the lineage rules are the hard part, which is the argument for
|
||||
reusing the ones that already work rather than against it.
|
||||
|
||||
### B.2 What exists now
|
||||
|
||||
```text
|
||||
backend/app/media/__init__.py 66 lines the finding, recorded where a
|
||||
future implementer will hit it
|
||||
backend/app/media/packet.py 328 lines the Scene Packet, built on read
|
||||
backend/app/media/profiles.py 220 lines visual profiles: the one thing
|
||||
§5 asks for that nothing stored
|
||||
backend/app/media/providers.py 332 lines contracts + endpoint policy
|
||||
backend/app/routers/adventures/visuals.py 131 four profile endpoints + the packet
|
||||
backend/app/models.py +1 model VisualProfile
|
||||
```
|
||||
|
||||
**Scene representation** — derived, not stored. `packet.build(db, adventure,
|
||||
start=, end=)` reads `narrative_store.current(adventure)`, which is the
|
||||
authoritative document at the active head, and returns:
|
||||
|
||||
```text
|
||||
scene_id derived identity, below
|
||||
campaign {id, title}
|
||||
turn_range {branch_id, start, end}
|
||||
lineage the head-capped branch lineage, as coordinates
|
||||
location entity view + its visual profile, or null
|
||||
characters present entities, each with its profile or null
|
||||
objects significant items, held or in the location
|
||||
action_summary the scene's own summary text
|
||||
continuity_constraints what must stay true in a depiction
|
||||
ambience {time_of_day, lighting, mood} — present, empty, §B.5
|
||||
source {packet_version: 1, head_depth}
|
||||
```
|
||||
|
||||
**Scene identity** — `c<adventure>:b<branch>:<start>-<end>`, e.g.
|
||||
`c7:b3:4-4`. **Derived, not allocated.** The same position yields the same id in
|
||||
any process, after any restart, and after the packet is thrown away and rebuilt,
|
||||
with no row to keep in step. `parse_scene_id()` reads it back. This is the part
|
||||
of a future `media_assets` table that would be expensive to retrofit, so it is
|
||||
guaranteed now even though the table is not built.
|
||||
|
||||
**Turn-range representation** — depths on the branch the scene was set on.
|
||||
Defaults to the scene's own single position; a caller passing `start` and `end`
|
||||
describes a stretch, which is what a future video provider would ask for
|
||||
(§31 of the contract, multi-turn packets).
|
||||
|
||||
**Lineage association** — the packet carries `turn_range.branch_id` and the
|
||||
head-capped `lineage` array. The scene's coordinate comes from the scene's own
|
||||
`at`, not from the head, and the difference is deliberate: a story can move on
|
||||
without re-establishing the scene, and a picture belongs to the moment the scene
|
||||
was set rather than to a later turn that did not change it.
|
||||
|
||||
**Visual profiles** — `visual_profiles`, keyed `(adventure_id, entity_key)`,
|
||||
with an open `descriptors` map, a `features` list and free `style_notes`. Four
|
||||
endpoints under `/api/adventures/{id}/visual-profiles` — list, read, write,
|
||||
delete. Campaign-scoped, not
|
||||
per-position (§D, §C.3).
|
||||
|
||||
**Provider-neutral interfaces** — `typing.Protocol` structural types, so a
|
||||
future adapter satisfies them by shape and imports nothing from here:
|
||||
|
||||
```text
|
||||
MediaProvider capabilities() -> ProviderCapabilities
|
||||
generate(MediaRequest) -> MediaResult
|
||||
SpeechProvider speak(...) -> MediaResult
|
||||
TranscriptionProvider transcribe(...) -> DraftTranscription
|
||||
```
|
||||
|
||||
with `MediaRequest`, `MediaResult`, `ProviderCapabilities`, `DraftTranscription`
|
||||
and `MediaProviderError` beside them, and a registry (`register`, `unregister`,
|
||||
`registered`, `for_kind`) that **is empty and ships empty**.
|
||||
|
||||
**STT draft-input contract** — `DraftTranscription` carries
|
||||
`editable: bool = True` and **has no commit method**. A transcriber can produce a
|
||||
draft and structurally cannot submit one; the ordinary authoritative commit path
|
||||
is the only way in. §24A's rule is enforced by the shape of the type rather than
|
||||
by a caller remembering it.
|
||||
|
||||
**Request/job/asset persistence** — **does not exist.** `MediaRequest` and
|
||||
`MediaResult` are contracts; there are no `media_jobs` or `media_assets` tables.
|
||||
See §F/K04 for the reasoning and how it is reported.
|
||||
|
||||
### B.3 Nothing calls any of it
|
||||
|
||||
No turn, prompt, context section, health check or startup path touches
|
||||
`app/media/`. Asserted structurally: `test_m10_no_media.py` parses the import
|
||||
statements of `turns.py`, `context/builder.py`, `narrative/apply.py`,
|
||||
`narrative/store.py`, `tree.py`, `head.py` and `memorybank.py` and requires that
|
||||
none imports the media package.
|
||||
|
||||
### B.4 What the packet excludes, which is the more interesting half
|
||||
|
||||
Excluded: the raw transcript, **all imported knowledge**, memories, summaries,
|
||||
and the state document's facts, relationships and threads.
|
||||
|
||||
The rule is *what the story established at this position*, not *everything the
|
||||
narrator was told*. Excluding imported knowledge **as a class** rather than
|
||||
filtering marked secrets is what makes §H hold for a secret nobody thought to
|
||||
mark: there is no filter to forget to extend.
|
||||
|
||||
### B.5 One deliberate gap
|
||||
|
||||
`ambience` returns `{time_of_day: null, lighting: null, mood: null}`. The fields
|
||||
are in the shape because a provider adapter should not have to branch on their
|
||||
absence; they are empty because filling them would mean extending the `set_scene`
|
||||
event, which is on the prompt path — a change to what the narrator is asked for,
|
||||
which is not M10's to make. Stated here rather than left to be discovered.
|
||||
|
||||
---
|
||||
|
||||
## C. Authority analysis
|
||||
|
||||
```text
|
||||
AUTHORITATIVE DERIVED
|
||||
┌────────────────────────────────┐ ┌──────────────────────────┐
|
||||
│ actions (the transcript) │ │ Scene Packet │
|
||||
│ state_events / state_proposals │──►│ built on read │
|
||||
│ narrative_state │ │ stored nowhere │
|
||||
│ narrative_state_after (per pos)│ │ identity computed │
|
||||
│ branches, head_branch/depth │ │ │
|
||||
│ checkpoints │ │ (future: media assets) │
|
||||
└────────────────────────────────┘ └──────────────────────────┘
|
||||
▲ │
|
||||
└────────── NO PATH ◄──────────┘
|
||||
|
||||
visual_profiles ── presentation metadata, campaign-scoped.
|
||||
Written by the reader, read by the packet, never by the story.
|
||||
```
|
||||
|
||||
### C.1 The reverse path does not exist, proved three ways
|
||||
|
||||
**By structure.** `test_m10_authority.py::test_the_media_package_imports_nothing_that_writes_state`
|
||||
requires that no file under `app/media/` references `narrative.apply`,
|
||||
`set_current`, `head.move_to` or `tree.place_action`. `narrative.model` and
|
||||
`narrative_store.current` are reads and are used.
|
||||
|
||||
**By vocabulary.** `test_no_state_event_type_was_added_for_media` requires that
|
||||
`events.ALLOWED` contains no name beginning `media` or containing `visual` or
|
||||
`asset`. There is no way for the media layer to speak in the story's language,
|
||||
so there is nothing for the validator to accept.
|
||||
|
||||
**By behaviour, which is the one that would catch a mistake nobody predicted.**
|
||||
Every test in `test_m10_authority.py` records the authoritative document —
|
||||
`narrative_state`, `head_branch_id`, `head_depth`, and the counts of state
|
||||
events, proposals and actions — before and after a media operation, and requires
|
||||
them to be **identical**:
|
||||
|
||||
| Operation | Result |
|
||||
| --- | --- |
|
||||
| Write a visual profile | authoritative document unchanged |
|
||||
| Update Alice's profile to "blue coat" | unchanged, and `"blue coat"` appears in no fact and in no entity |
|
||||
| Delete a profile | unchanged |
|
||||
| Build five scene packets | unchanged |
|
||||
| Build a packet with an explicit range | head does not move |
|
||||
| A dummy provider returns "Alice in a red coat in a corridor" | unchanged; neither string is anywhere in the state |
|
||||
| A provider raises `MediaProviderError` | unchanged; head does not advance |
|
||||
| Packet derivation raises inside `build` | unchanged, and the campaign still plays |
|
||||
| Delete every profile | campaign intact; packet still builds with `visual_profile: null` |
|
||||
| A profile naming an entity that does not exist | inert; packet unaffected; play continues |
|
||||
|
||||
### C.2 The claim in the other direction
|
||||
|
||||
§35 and §37 of the contract — a depiction never becomes canon, and promoting a
|
||||
visual detail into canon must be a deliberate act by the reader. M10 makes that
|
||||
structural: there is no code path from a `MediaResult` to a state event, because
|
||||
there is no code that consumes a `MediaResult` at all.
|
||||
|
||||
### C.3 Why a profile carries no branch coordinate
|
||||
|
||||
Every other derived record in the schema carries `(branch_id, depth)` because it
|
||||
describes a *moment*. A profile describes none: a character does not change
|
||||
appearance because the story forked. Per-position profiles would have been wrong
|
||||
twice — a reader who diverged would lose their cast's appearance, which is the
|
||||
opposite of the continuity a profile exists for, and a descriptor document would
|
||||
land in every per-position snapshot (measured: 245 copies of the same 367 bytes
|
||||
in a 120-turn campaign, §K).
|
||||
|
||||
---
|
||||
|
||||
## D. Schema and migration
|
||||
|
||||
**Tables added:** one.
|
||||
|
||||
```text
|
||||
visual_profiles
|
||||
id INTEGER PRIMARY KEY
|
||||
adventure_id INTEGER NOT NULL FK adventures(id) ON DELETE CASCADE, indexed
|
||||
entity_key VARCHAR(200) NOT NULL
|
||||
descriptors JSON
|
||||
features JSON
|
||||
style_notes TEXT
|
||||
created_at DATETIME
|
||||
updated_at DATETIME
|
||||
UNIQUE (adventure_id, entity_key) -- uq_visual_entity
|
||||
INDEX ix_visual_profiles_adventure_id
|
||||
```
|
||||
|
||||
**Columns added to existing tables:** none.
|
||||
**Columns changed or dropped:** none.
|
||||
**Story tables touched:** none.
|
||||
|
||||
**Migrations added: none.** `LATEST_VERSION` is **92**, exactly as M9 left it.
|
||||
|
||||
`create_all` builds a new table on every path — fresh install, existing database,
|
||||
test setup — as it did for `memories`, `branches`, `checkpoints`, `summaries` and
|
||||
the M7 knowledge tables, and it builds the index too, because the index is
|
||||
declared on the column rather than in `__table_args__`. Migration 92's own
|
||||
comment states this rule for the M7 tables; M10 follows it.
|
||||
|
||||
**Proof that none is needed**, and that the two paths converge
|
||||
(`test_m10_bundle.py`, §15 group):
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| An M9-era database (schema at 92, no `visual_profiles`, a campaign already in it) opened by this build | table present, empty, `ix_visual_profiles_adventure_id` present, version still 92 |
|
||||
| The campaign that was already there | title, action text and `PRAGMA foreign_key_check` unchanged |
|
||||
| Opening the same database three times | index set identical after each; no error; no accumulation |
|
||||
| Fresh install vs upgraded M9 file | `sqlite_master` DDL for `visual_profiles` **identical**; index sets **identical** |
|
||||
| `backup.create()` on the upgraded file | `integrity == "ok"`, `quick_check` ok, `foreign_key_check` empty, opens independently, carries the new table and the pre-M10 campaign, `user_version` 92 |
|
||||
| A profile written after the upgrade, then backed up | present in the backup with its descriptors |
|
||||
|
||||
**A defect this found — see §M.1.** M10 first shipped migration 93 creating
|
||||
`ix_visual_profiles_adventure`. Because `create_all` had already built
|
||||
`ix_visual_profiles_adventure_id`, an *upgraded* database ended up with both and
|
||||
a fresh install with one. The fresh-versus-upgraded comparison caught it; the
|
||||
migration was removed rather than renamed, because the right number of
|
||||
migrations here is zero.
|
||||
|
||||
---
|
||||
|
||||
## E. Bundle implications
|
||||
|
||||
**Did M9's v3 bundle change?** Yes — it gained one optional key:
|
||||
|
||||
```json
|
||||
"visualProfiles": [
|
||||
{"entityKey": "alice",
|
||||
"descriptors": {"build": "tall", "hair": "short black", "clothing": "grey blazer"},
|
||||
"features": ["tortoiseshell glasses"],
|
||||
"styleNotes": "photographic, natural light",
|
||||
"createdAt": "2026-09-07T…"}
|
||||
]
|
||||
```
|
||||
|
||||
**Did the format version change?** **No. It stays `ai-dnd-adventure-v3`.**
|
||||
|
||||
**Why.** M9 introduced a version because a v2 file with no prompt provenance was
|
||||
ambiguous between "written before M9" and "written by M9 from a campaign that
|
||||
had none". The test is therefore not "did the format gain a key" but *does
|
||||
omission create ambiguity about what an older file could have recorded*. It does
|
||||
not: a campaign with no visual profiles is the ordinary case — appearance is
|
||||
something a reader adds, not something a campaign has by default — so an absent
|
||||
key unambiguously means "none", exactly as `checkpoints` did before M4 and
|
||||
`knowledge` before M7. Bumping to v4 for a key whose absence is unambiguous would
|
||||
spend the mechanism M9 built and make it mean less next time.
|
||||
|
||||
**Legacy import behaviour.** Unchanged, and re-checked:
|
||||
|
||||
| File | Result |
|
||||
| --- | --- |
|
||||
| A v3 file with `visualProfiles` deleted (an M9-written file) | imports; campaign intact; zero profiles |
|
||||
| v2, v1 | unchanged — M9's readers are untouched |
|
||||
| A malformed profile in an otherwise good file | **dropped, campaign still imports.** A story that would not import because a description of somebody's coat is malformed would be the wrong trade |
|
||||
| A profile for an entity that no longer exists | imported and inert |
|
||||
|
||||
**Round trip.** Export → import → export produces the same profiles, in the same
|
||||
shape (`test_a_profile_survives_a_second_round_trip_unchanged`). The copy's
|
||||
profiles are its own rows — editing the copy does not reach the original — and
|
||||
the copy's scene packet is populated from them, which is the point of carrying
|
||||
them at all. A neighbouring campaign's profiles do not travel. The planner
|
||||
(`bundle.plan`, M9's before-anything-is-written checkpoint) checks profiles
|
||||
there rather than partway through a write.
|
||||
|
||||
**Clean-directory round trip, with profiles, across two real processes.**
|
||||
`test_profiles_reach_a_clean_data_directory_on_another_machine` follows M9's
|
||||
shape — two directories, two databases, two server processes, nothing crossing
|
||||
but the file, and machine B's database a file that never existed before, so its
|
||||
migrations run from nothing. Machine A profiles Alice and the office and exports;
|
||||
machine B imports and builds a scene packet whose `action_summary` matches,
|
||||
whose Alice carries her descriptors, whose office carries its lighting, and whose
|
||||
Roger is still `visual_profile: null`. Only the campaign id differs, which is
|
||||
what a new machine's id space means.
|
||||
|
||||
This exists because M9's own `test_m9_clean_import.py` predates visual profiles
|
||||
and carries none — it passes unchanged (§L), but it could not have caught a
|
||||
profile that failed to cross. M10 added no cross-machine coupling: `entity_key`
|
||||
is a key inside the campaign's own state document, which travels in the same
|
||||
file, so unlike a branch number or a knowledge source id it needs no translation
|
||||
on import.
|
||||
|
||||
---
|
||||
|
||||
## F. Acceptance matrix
|
||||
|
||||
| Test | Verdict | Evidence |
|
||||
| --- | --- | --- |
|
||||
| **K01 — Scene Snapshot Exists** | **PASS** *(and was already passing)* | `narrative_state["scene"]` since M5; normalized packet at `GET /api/adventures/{id}/scene-packet`. `test_m10_media_hooks.py` (contents, bounds, identity), `test_m10_lineage.py` (position correctness through every history operation and two process restarts). |
|
||||
| **K02 — Visual Character Profile** | **PASS** | `visual_profiles` + four endpoints (list, read, write, delete). Optional is tested, not just stated: Roger is deliberately unprofiled and the packet reports `visual_profile: null` rather than an empty profile. Stability is tested per operation — Undo and divergence, Redo and Save Point restore, a genuine process restart, and a two-process move to a clean data directory. |
|
||||
| **K03 — Visual Location Profile** | **PASS** | Same table and same code path — a location is an entity with a `type`. `the office` carries a profile; the packet's `location.visual_profile` returns it. |
|
||||
| **K04 — Attach Media Asset to Scene** | **PASS on the deferred branch** | Media tables are deliberately deferred, so this is reported against the acceptance text's own second clause ("if media tables are deferred: architecture/types should demonstrate equivalent extension point"). A test registers a dummy provider, builds a packet, generates a fake PNG carrying the packet's `scene_id` as provenance, and shows the story model byte-for-byte unchanged. **It is not PASS on the first clause**, and a reviewer who requires physical media tables in v1 should read this as PARTIAL. |
|
||||
|
||||
**K04's exact status, stated plainly.** Physically implementing `media_jobs` and
|
||||
`media_assets` now would mean designing a queue with no producer and no consumer,
|
||||
whose shape would be decided by a provider nobody has chosen; the codebase
|
||||
declined the same thing once already (M6's `derived_status`, commented "not a job
|
||||
queue"). What is guaranteed instead is the part that would be expensive to
|
||||
retrofit: a scene identity that is *derived* from campaign and position, so a
|
||||
future asset can reference a scene without a scenes table existing to reference.
|
||||
|
||||
---
|
||||
|
||||
## G. History and lineage evidence
|
||||
|
||||
`test_m10_lineage.py` — 8 tests, all passing, including the brief's own §4
|
||||
example and its §17 sequence, and two **genuine spawned-process restarts**
|
||||
(reusing `test_process_restart.Server`, so the process really goes away).
|
||||
|
||||
| Operation | What was checked | Result |
|
||||
| --- | --- | --- |
|
||||
| Set a scene, continue | packet describes the scene's position, not the head's | pass |
|
||||
| **Undo** | packet follows the state back; a scene set after the undone point is gone from it | pass |
|
||||
| **Redo** | packet returns to the later scene, with the same `scene_id` it had before | pass |
|
||||
| **Retry** | the take that is live decides the scene; the superseded take's scene does not leak | pass |
|
||||
| **Divergence** | Path A's scene and Path B's scene are different packets with different ids; both retained; the abandoned one is not current | pass |
|
||||
| **Save Point restore** | packet matches the position the Save Point names | pass |
|
||||
| **Redo, and a Save Point restore** | the profile is untouched by either; the scene set after the Save Point is correctly gone from the packet while the profile remains — which is the difference between story state and presentation metadata | pass |
|
||||
| **Restart (real process)** | the same position yields the same `scene_id` and the same packet contents in a new process | pass |
|
||||
| **Restart after divergence (real process)** | the campaign reopens on the branch it was left on, and the packet is that branch's | pass |
|
||||
|
||||
The mechanism behind all of it is M5's, not M10's: the head move restores the
|
||||
whole state document and the scene is part of it. What M10 adds is the test that
|
||||
pins it for the media seam, plus the derived identity that makes the restart
|
||||
comparison meaningful — a stored id would have been trivially stable and would
|
||||
have proved nothing.
|
||||
|
||||
---
|
||||
|
||||
## H. Hidden-information evidence
|
||||
|
||||
The test uses a **hidden M7 knowledge source**, because that is the product's
|
||||
real narrator-only mechanism, rather than an invented marker. Each check carries
|
||||
a **positive control**, so a pass cannot be a campaign where the secret was never
|
||||
established.
|
||||
|
||||
**The sentinel.** `ZARQUON-CONCEALED-OBSERVER-7731`, in a hidden Canon source
|
||||
describing a concealed observer behind the office's north wall.
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| **Control:** does the narrator actually receive it? | **Yes** — the sentinel appears in the assembled prompt for a turn about the north wall panelling |
|
||||
| Does it reach the Scene Packet? | **No** — neither the sentinel nor "concealed observer" is anywhere in the packet |
|
||||
| **Control:** does a *visible* reference source reach the narrator? | **Yes** — the handbook appears in the context report's used-knowledge list |
|
||||
| Does that visible source reach the packet? | **No** — imported knowledge is excluded as a class, which is what makes the rule hold for a secret nobody thought to mark |
|
||||
| Do memories and summaries reach the packet? | **No** — after eight turns and a summary pass, the recurring "printer incident" text is absent |
|
||||
| Does something the *story* established reach the packet? | **Yes**, and it should — once a validated `set_scene` event puts the observer in the room, the observer is in the packet. It is no longer narrator-only knowledge; it is something that happened. The sentinel is still absent, because the story never said it. |
|
||||
|
||||
That last row is the reason the boundary is drawn where it is: a packet that hid
|
||||
established story from a depiction would be hiding the story from itself.
|
||||
|
||||
---
|
||||
|
||||
## I. No-media operation
|
||||
|
||||
`test_m10_no_media.py` — 12 tests, all passing. §20's list, run in one campaign
|
||||
with an empty provider registry:
|
||||
|
||||
```text
|
||||
several story turns ✓ Undo ✓
|
||||
state extraction ✓ Redo ✓
|
||||
memory + summary activity ✓ Retry ✓
|
||||
knowledge retrieval ✓ Save Point restore ✓
|
||||
restart ✓ *
|
||||
```
|
||||
|
||||
\* In this suite the restart is a fresh session reading what was written, not a
|
||||
new process. The **genuine spawned-process restarts are in
|
||||
`test_m10_lineage.py`** (§G), where they carry more weight: they run with
|
||||
profiles written and packets built, and check the derived `scene_id` is the same
|
||||
in a process that never saw the first one.
|
||||
|
||||
with **no** media warning, **no** media connection attempt, **no**
|
||||
missing-provider error, and no media schema requirement reaching narration.
|
||||
|
||||
Also checked:
|
||||
|
||||
- **The prompt is unchanged.** No context section labelled `media`,
|
||||
`scene_packet` or `visual_profile*` exists; the strings `visual_profile` and
|
||||
`scene_id` appear nowhere in the assembled system or story prompt.
|
||||
- **Ordinary play writes no media row.**
|
||||
- **No media setting exists** — neither in the settings API response nor as a
|
||||
column on `Settings`. A setting that exists is a setting that can be pointed at
|
||||
a cloud by mistake.
|
||||
- **The turn path cannot reach the media package**, checked by parsing imports
|
||||
rather than by grepping text.
|
||||
- The app serves, plays and passes health checks with `providers.registered() == {}`.
|
||||
|
||||
---
|
||||
|
||||
## J. Security and locality
|
||||
|
||||
**Outbound destinations added: none.** No module under `app/media/` references
|
||||
`httpx`, `requests`, `urllib.request`, `socket`, `aiohttp` or `subprocess`, and
|
||||
that is a test, not an inspection. No provider adapter ships, so there is nothing
|
||||
to connect *to*; the registry is empty at import and stays empty.
|
||||
|
||||
**Endpoint architecture — stricter than narration.**
|
||||
`providers.endpoint_rejection_reason(url)` applies `endpoints.rejection_reason`
|
||||
first (the shared policy: every address the hostname resolves to must be
|
||||
loopback, RFC1918, link-local, IPv6 ULA or CGNAT; known cloud inference hosts are
|
||||
refused by name; the check is on the **resolved address**, so
|
||||
`localhost.evil.example` does not pass) and **then requires loopback in
|
||||
addition**. Verified:
|
||||
|
||||
| Endpoint | Verdict |
|
||||
| --- | --- |
|
||||
| `http://127.0.0.1:8188/` | allowed |
|
||||
| `http://192.168.x.x:8188/` (trusted LAN — allowed for *narration*) | **refused** for media |
|
||||
| `https://api.openai.com/v1`, `https://replicate.com`, `http://8.8.8.8:8188`, `http://example.com` | refused |
|
||||
|
||||
This is deliberately narrower than `SECURITY-THREAT-MODEL.md` §73 permits, and
|
||||
§42A now records the discrepancy rather than leaving it to be found. The
|
||||
reasoning: a picture of a scene carries the scene with it, and a GPU rendering
|
||||
someone's campaign is a machine that person is sitting at.
|
||||
|
||||
**No TLS verification bypass** was introduced. There is no `verify=False`, no
|
||||
`-k`, and no new HTTP client at all; `tlstrust.py` is untouched and
|
||||
`test_tls_trust.py` passes.
|
||||
|
||||
**Filesystem/media exposure: none.** No asset is stored, no directory is served,
|
||||
no path comes from a caller. M10 adds no file-serving route.
|
||||
|
||||
**CSP/CORS: unchanged.** No frontend file was modified, no new origin is
|
||||
contacted, and `test_offline_assets.py` (which reads the built SPA and the CSP)
|
||||
passes.
|
||||
|
||||
**Cloud dependency: none.** Nothing was installed to demonstrate an interface —
|
||||
no ComfyUI, no diffusers, no Whisper, no Kokoro, no model download. The
|
||||
`requirements.txt` diff is empty.
|
||||
|
||||
**One new disclosure boundary, and it is not a network one.** The Scene Packet is
|
||||
the input a future provider would receive, so its contents are a disclosure
|
||||
decision — covered in §H.
|
||||
|
||||
---
|
||||
|
||||
## K. Performance and storage
|
||||
|
||||
Measured with `backend/tools/m10_media_cost.py`, a 120-turn campaign played
|
||||
through the real turn engine and the real state pipeline, with SQL statements
|
||||
counted by `tools/dbmeter.py`.
|
||||
|
||||
```text
|
||||
120 turns, 245 action rows
|
||||
|
||||
scene records M10 wrote
|
||||
visual_profiles rows 2 one per profiled entity, written once
|
||||
scene rows 0 M10 adds no scenes table
|
||||
M5 per-position state snapshots 245 already there; the scene lives here
|
||||
|
||||
bytes added to the database
|
||||
empty database 188416 B
|
||||
after the fixture campaign 196608 B
|
||||
after 120 more turns 1056768 B
|
||||
visual profile content 367 B 0.035% of the database
|
||||
|
||||
profile duplication: campaign-scoped against per-position
|
||||
as stored, once per entity 367 B
|
||||
if snapshotted per position 89915 B x245
|
||||
|
||||
packet: persisted or constructed
|
||||
rows written while building one 0
|
||||
packet rows in any table 0 built on read, never stored
|
||||
build time, 2 turns 16.6 ms
|
||||
build time, 122 turns 12.6 ms
|
||||
|
||||
current-scene query behaviour
|
||||
statements, 2 turns 5
|
||||
statements, 122 turns 4
|
||||
does not grow with the campaign
|
||||
```
|
||||
|
||||
The packet's four statements at turn 122 are: the adventure, its visual
|
||||
profiles, the current user, and the head's branch. **Nothing walks the
|
||||
transcript**, which is the property that matters — a scene derivation that
|
||||
scanned actions would have made every future depiction O(turns).
|
||||
|
||||
The `x245` row is the measurement behind §C.3: per-position profiles would have
|
||||
stored the same 367 bytes 245 times in this campaign to say something that never
|
||||
varies.
|
||||
|
||||
**Storage added to an ordinary campaign that uses no profiles: one empty table.**
|
||||
|
||||
---
|
||||
|
||||
## L. Regression counts
|
||||
|
||||
All runs on this branch, after every M10 change, on 2026-09-07.
|
||||
|
||||
| Suite | Result | Time |
|
||||
| --- | --- | --- |
|
||||
| **Backend, full** | **1,191 passed, 14 skipped, 0 failed** — 1,205 collected, exit 0 | 862.9 s |
|
||||
| **M10-specific** | **89 passed** (39 hooks + 8 lineage + 12 authority + 12 no-media + 18 bundle/migration) | 56.4 s |
|
||||
| **Frontend component suite** | **145 passed**, 12 files, 0 failed | 8.17 s |
|
||||
| **Lint** (`oxlint`) | **0 errors**, 15 warnings, exit 0 | — |
|
||||
| **Production build** (`vite build`) | clean — `index-DcHbz7ga.js` 389.66 kB (gzip 119.25 kB), `index-B55Q8MSM.css` 47.50 kB | 0.58 s |
|
||||
| **Docker** (`--no-cache`) | **exit 0**; image runs and imports `app.media` with `registered() == {}` | 25.2 s |
|
||||
| **Browser** | **not run — M10 adds no reader-facing surface.** See §L.2 |
|
||||
|
||||
The 14 skips are M9's and are unchanged — confirmed by running the five files
|
||||
that carry a skip condition on their own (26 passed, 14 skipped): seven need a
|
||||
second machine or an environment the suite cannot create; the rest need a real
|
||||
local model.
|
||||
|
||||
The 15 lint warnings are pre-existing (`no-unused-vars` in two test files, and
|
||||
`react/only-export-components` in components that export a constant beside a
|
||||
component). **M10 modified no frontend file**, so the count is M9's, unchanged.
|
||||
|
||||
### L.1 By group
|
||||
|
||||
Each group was run on its own, so a reviewer can check a claim without running
|
||||
the whole suite. Every group is green.
|
||||
|
||||
| Group | Files | Result |
|
||||
| --- | --- | --- |
|
||||
| **M10** | `test_m10_media_hooks` (39), `_lineage` (8), `_authority` (12), `_no_media` (12), `_bundle` (18) | **89 passed** — 56.4 s |
|
||||
| **M9 recovery** | `test_m9_backup`, `_clean_import`, `_corrupt_bundles`, `_legacy_bundles`, `_portability`, `test_bundle_v2` | **173 passed** — 371.3 s |
|
||||
| **Migration** | `test_tree_migration`, `test_knowledge_migration`, `test_pre_m5_compatibility`, `test_snapshot_compression` | **50 passed** — 27.8 s |
|
||||
| **M5/M6/M7 state, memory, knowledge** | `test_narrative_state`, `test_worldstate_integration`, `test_memory_nodes`, `test_memory_retrieval`, `test_context_memory`, `test_imported_knowledge`, `test_knowledge_retrieval_quality` | **216 passed** — 113.7 s |
|
||||
| **M3/M4 history** | `test_story_tree_baseline`, `test_branch_forking`, `test_head_cursor`, `test_save_points`, `test_take_state`, `test_process_restart` | **137 passed** — 139.6 s |
|
||||
| **Security / offline** | `test_egress`, `test_endpoint_policy`, `test_local_only_surface`, `test_offline_assets`, `test_tls_trust` | **95 passed** — 18.4 s |
|
||||
|
||||
The security group is the one that carries §J's claims: no outbound route, the
|
||||
endpoint policy, the local-only API surface, the offline asset and CSP checks,
|
||||
and the TLS trust union. All were green before M10 and are green now.
|
||||
|
||||
### L.2 On the browser run
|
||||
|
||||
M10 adds **no reader-facing surface**: no page, no control, no copy, no route in
|
||||
the SPA. `git status` shows no file under `frontend/` modified. The five new
|
||||
endpoints are backend-only and nothing in the browser calls them. A real-browser
|
||||
regression pass would therefore be re-verifying M8/M9's surfaces against a build
|
||||
identical to theirs, and its evidence would be M9's evidence. The frontend
|
||||
component suite and the production build were run anyway, and are green.
|
||||
|
||||
This is stated as a decision, not an omission: **if the reviewer wants a browser
|
||||
pass as a matter of process, it has not been done.**
|
||||
|
||||
---
|
||||
|
||||
## M. Findings
|
||||
|
||||
### M.1 A redundant index migration made two databases disagree — **introduced by M10, fixed here**
|
||||
|
||||
**Severity:** low in effect, moderate in kind. **Blocker:** no — fixed.
|
||||
**Owner:** M10 (closed).
|
||||
|
||||
M10 first added migration 93, `CREATE INDEX IF NOT EXISTS ix_visual_profiles_adventure
|
||||
ON visual_profiles (adventure_id)`. But `VisualProfile.adventure_id` declares
|
||||
`index=True`, so `create_all` already builds `ix_visual_profiles_adventure_id` —
|
||||
on a fresh install *and* on an existing database, since `create_all` runs before
|
||||
the migration loop. The result:
|
||||
|
||||
```text
|
||||
upgraded from 92: ix_visual_profiles_adventure, ix_visual_profiles_adventure_id
|
||||
fresh install: ix_visual_profiles_adventure_id
|
||||
```
|
||||
|
||||
Two schemas differing by which path the file took, which is the thing a migration
|
||||
exists to prevent, plus a redundant index on every upgraded database.
|
||||
|
||||
**Found by** `test_a_fresh_database_arrives_at_the_same_place`, which compares a
|
||||
fresh schema against an upgraded one. Neither database examined on its own would
|
||||
have shown it. **Fixed by removing the migration**, not by renaming the index:
|
||||
migration 92's own comment already records the rule for the M7 tables — when the
|
||||
index is declared on the column there is nothing left for a `CREATE INDEX` to do.
|
||||
`LATEST_VERSION` returns to 92.
|
||||
|
||||
### M.2 The state model refuses an over-large scene, and a test asked for one — **not a defect**
|
||||
|
||||
While writing the packet's bounds test I sent 43 entries in `set_scene`'s
|
||||
`present`, exceeding `validate.MAX_LABELS = 40`. The event was correctly refused
|
||||
and the previous scene stayed, so the test measured the wrong scene and failed.
|
||||
Recorded because the diagnosis matters: the product was right and the test was
|
||||
wrong. The test now uses 33 and asserts its own precondition, so it cannot
|
||||
silently measure a scene it did not set.
|
||||
|
||||
### M.3 The Story Engine vocabulary grep hit the file that forbids the vocabulary — **test defect, fixed**
|
||||
|
||||
§9's rule is that no provider vocabulary (ComfyUI, Whisper, `num_inference_steps`,
|
||||
LoRA) appears in the Story Engine. The first version of the test grepped the
|
||||
whole backend and hit `providers.py`, whose docstrings *name* those things
|
||||
precisely in order to exclude them. Fixed by scoping the grep to the story
|
||||
engine, and by adding a complementary test that checks the seam **by behaviour**:
|
||||
no provider registered, and no networking import anywhere under `app/media/`.
|
||||
|
||||
### M.3a A text search for "media" matched "im**media**tely" — **test defect, fixed**
|
||||
|
||||
The first version of the no-media import check read each turn-path module and
|
||||
required the string `media` to be absent. `narrative/store.py` contains the word
|
||||
*immediately*, so the test failed on a module that imports nothing. It now parses
|
||||
the file and inspects its **import statements**, which is what the claim was
|
||||
always about. Recorded because the failure looked briefly like a real coupling
|
||||
and was not, and because the fixed version is the stronger test: a module could
|
||||
have imported the package while never spelling the word in prose.
|
||||
|
||||
### M.4 `ambience` is present and empty — **known gap, deliberate**
|
||||
|
||||
**Severity:** low. **Blocker:** no. **Owner:** whichever milestone builds a
|
||||
coordinator.
|
||||
|
||||
The packet's `ambience` object has the right shape and no content, because
|
||||
filling it would mean extending the `set_scene` event — a change on the prompt
|
||||
path, asking the narrator for something new, which is outside M10's scope. A
|
||||
future provider gets a stable shape today and content when someone decides the
|
||||
narrator should be asked.
|
||||
|
||||
### M.5 Deleting profiles is not recoverable from within the app — **known limit, stated**
|
||||
|
||||
**Severity:** low. **Blocker:** no.
|
||||
|
||||
M9's rebuildable data can be regenerated; a visual profile cannot, because it is
|
||||
something a reader wrote. It travels in the bundle, so a backup or an export
|
||||
recovers it, and deletion is per-entity and explicit. There is no undo for it,
|
||||
and none was invented — that would be a second history model beside the story's.
|
||||
|
||||
**No pre-existing defect was found in M9's or earlier work during this
|
||||
milestone.** The full backend suite was green before M10 began and is green now.
|
||||
|
||||
---
|
||||
|
||||
## N. Planning changes
|
||||
|
||||
| Document | Change | Why |
|
||||
| --- | --- | --- |
|
||||
| `planning/DATA-MODEL.md` | **New §28A** — media extension points as implemented | §20/§27/§28 describe a scene table and job/asset tables. Only one of the three exists, and a reader of the conceptual model needs to know which, and why the scene is derived from §20's own data rather than stored beside it. |
|
||||
| `planning/TECHNICAL-DESIGN.md` | **New §15.1** under Scene and Future Media Boundary | §15 said "persist or derive". The answer is *derive*, and the reason (M5 already persisted it) is the milestone's central fact. Also records the packet's exclusions, the STT asymmetry and the endpoint policy. |
|
||||
| `planning/MEDIA-EXTENSION-CONTRACT.md` | **New §90**, appended | The contract is Phase 0B design and stays readable as such. §90 records what was built, the three places implementation answered an open question (§5 already satisfied, §12 drawn wider, §7-9 collapsed into one table), and what is deliberately unbuilt. |
|
||||
| `planning/BUILD-MILESTONES.md` | **M10 status block** | Milestone status, the shaping finding, what shipped, and the defect its own tests caught. The four post-M8 playtest findings above it are untouched and still M11's. |
|
||||
| `planning/V1-ACCEPTANCE-TESTS.md` | **K01-K04 results** | Acceptance evidence. K01 records that it was already passing; K04 records which of its two clauses it passes on. |
|
||||
| `planning/SECURITY-THREAT-MODEL.md` | **New §42A** | The trust boundary did not widen, but in one place the implementation is deliberately **narrower** than §73 permits. A stricter implementation than the model describes is still a discrepancy, and an undocumented one becomes an accidental relaxation later. |
|
||||
| `planning/VERSION.md` | **v3.6 entry** | Records this milestone's documentation changes and the two decisions (no format bump, no migration). |
|
||||
| `planning/README.md` | Status, milestone map, reading order, report rotation | M10 is implemented; M11 is next. Records **why M9's report stays in `reports/`** against the usual rotation: M9 is not accepted, and M10's baseline is M9's. |
|
||||
| `README.md` | `media/` in the architecture map; `VisualProfile` in the model list; test count 920 → 1,191 | The map is the first thing a new reader reads. |
|
||||
| `DEVELOPMENT.md` | The stricter future media endpoint rule; a note that the suite takes ~15 minutes | Whoever adds the first provider should find the rule before writing the adapter. |
|
||||
|
||||
No planning document was rewritten, and no earlier milestone's evidence was
|
||||
edited.
|
||||
|
||||
---
|
||||
|
||||
## O. M11 handoff
|
||||
|
||||
### O.1 M10's residual risk
|
||||
|
||||
1. **K04 is satisfied structurally, not physically.** No `media_jobs` or
|
||||
`media_assets` table exists. A future coordinator will design them, and the
|
||||
contracts here constrain that design only loosely. *Risk: low — the expensive
|
||||
part (a stable scene identity) is fixed; the cheap part (two tables) is not.*
|
||||
2. **`ambience` is an empty shape** (§M.4). Filling it means extending
|
||||
`set_scene`, which changes what the narrator is asked for. *Risk: low; a
|
||||
provider adapter written today would find the fields and no values.*
|
||||
3. **The seam has no consumer, so it is unexercised by real use.** Every test
|
||||
here uses a dummy provider. The contracts are shaped by the contract document
|
||||
and by what the state model can supply, not by an adapter that had to work
|
||||
against a real generator. *Risk: moderate for the interfaces' ergonomics, nil
|
||||
for the story engine — the first real adapter may want the packet reshaped,
|
||||
and nothing in the story depends on its shape.*
|
||||
4. **A visual profile cannot be recovered from within the app** (§M.5).
|
||||
5. **Profiles are not surfaced to the reader at all.** They are API-only. Whoever
|
||||
builds a media UI owns the browser surface, and no reader-facing vocabulary
|
||||
for them has been invented — deliberately, since M10 was told not to introduce
|
||||
reader-facing branding or surfaces.
|
||||
|
||||
### O.2 M9 carry-forward still relevant
|
||||
|
||||
All six of M9's residual risks are unchanged by M10 — none was addressed and none
|
||||
was made worse:
|
||||
|
||||
| M9 residual | Status after M10 |
|
||||
| --- | --- |
|
||||
| Bundle ceiling ~279 turns | unchanged. M10 adds ~370 bytes per campaign to a file whose ceiling is set by per-position state; it does not move the number. |
|
||||
| `quick_check` rather than `integrity_check` | unchanged; re-exercised on a migrated database (§D) |
|
||||
| No scheduled backup | unchanged |
|
||||
| Stale `chunk_id` in a restored snapshot | unchanged |
|
||||
| Importing machine's context window may differ | unchanged; still M11's |
|
||||
| This machine cannot drive a file into/out of the browser | unchanged; still M11's, and still the cheapest fix is an unconfined Firefox or Xvfb |
|
||||
|
||||
M9's two items "carried to M10" are both **closed**: media tables were owned and
|
||||
the decision is recorded (§F, K04); the `scene` section of the state document was
|
||||
indeed the extension point, and is what the packet is built from.
|
||||
|
||||
### O.3 The post-M8 hands-on findings — still M11's, untouched
|
||||
|
||||
M10 neither implemented nor tested any of them, as its brief required. They are
|
||||
listed here so they cannot be lost when M9's report is eventually archived; the
|
||||
durable copy is in `BUILD-MILESTONES.md`.
|
||||
|
||||
| Finding | Owner |
|
||||
| --- | --- |
|
||||
| **A.** The browser tab still reads `AI D&D` | M11 release polish |
|
||||
| **B.** After Undo, the reader cannot tell where they are | M11 UX/release polish |
|
||||
| **C.** The narration-length setting has no measurable effect | M11 realistic-model behaviour |
|
||||
| **D.** Character identity / coreference confusion — root cause **unknown**, and the playtest database was destroyed | M11 realistic-model / context diagnostic |
|
||||
|
||||
**On D specifically:** M10's fixture is deliberately the same shape — a
|
||||
protagonist, two more characters in the room, and a fourth who is not — but
|
||||
**nothing in M10 asserts anything about coreference**, and the fixture's
|
||||
resemblance is not evidence about the finding. M11 still owes the explicit
|
||||
diagnostic. What M10 does contribute, incidentally, is that a visual profile
|
||||
attaches to the canonical `entity_key` rather than to a display name, so whatever
|
||||
M11 concludes about identity, profiles are keyed to the thing the state model
|
||||
considers one character.
|
||||
|
||||
### O.4 The context-window issue
|
||||
|
||||
Unchanged and still M11's: a deployment's enforced context window may be far
|
||||
below `Settings.context_token_budget` (Ollama defaults to 4,096 when it sees no
|
||||
VRAM). Documented in `DEVELOPMENT.md`; detection, the Settings warning question,
|
||||
and the 100-turn certification all remain open. M10 sends nothing to a model and
|
||||
does not touch it.
|
||||
|
||||
### O.5 Realistic-model / 100-turn work
|
||||
|
||||
Untouched by M10, and M10 adds no new realistic-model obligation: the media seam
|
||||
has no model in it. The 100-turn certification, the contrast/focus measurement,
|
||||
the WCAG audit and the streaming-import question all carry forward unchanged.
|
||||
|
||||
---
|
||||
|
||||
## P. Final verdict
|
||||
|
||||
**1. Is M10's Definition of Done satisfied?**
|
||||
|
||||
> *Future media providers can be added through defined local interfaces without
|
||||
> redesigning core story authority/history.*
|
||||
|
||||
**Yes.** A provider is added by implementing a `Protocol` and registering it;
|
||||
it receives a Scene Packet built from the authoritative state at a position, and
|
||||
returns a `MediaResult`. Nothing in story authority or history changes to
|
||||
accommodate it — proved by the fact that nothing in `app/media/` can even import
|
||||
the code that writes state, and that every media operation leaves the
|
||||
authoritative document byte-identical.
|
||||
|
||||
The honest qualification: **no real adapter has been written against these
|
||||
interfaces**, so their ergonomics are untested (§O.1.3). The story engine's
|
||||
independence, which is what the Definition of Done is actually about, is tested.
|
||||
|
||||
**2. Are K01-K03 all PASS?** **Yes** — all three, with evidence in §F and §G.
|
||||
K01 additionally records that it was already passing before M10 began.
|
||||
|
||||
**3. What is K04's exact status?** **PASS on the acceptance text's deferred
|
||||
branch** ("if media tables are deferred: architecture/types should demonstrate
|
||||
equivalent extension point"), demonstrated by a dummy provider producing an asset
|
||||
against a real packet with the story model unchanged. **Not PASS on the first
|
||||
branch**, since no media table is physically implemented. A reviewer who requires
|
||||
physical tables in v1 should read K04 as **PARTIAL**; the decision and its
|
||||
reasoning are in §F.
|
||||
|
||||
**4. Can all ordinary story operation run with zero media provider?** **Yes** —
|
||||
§I. Turns, state extraction, memory, summaries, knowledge retrieval, Undo, Redo,
|
||||
Retry and Save Point restore, with an empty registry, no media setting in
|
||||
existence, no warning, no connection attempt and no media row written. Restart is
|
||||
covered as a genuine spawned-process restart in the lineage suite (§G); the
|
||||
no-media suite's own restart step is a fresh session read, and §I says so.
|
||||
|
||||
**5. Can a future image provider be added without modifying story authority or
|
||||
history?** **Yes.** It implements `MediaProvider`, is handed a packet, and
|
||||
returns a result. No story table, event type, or history operation changes.
|
||||
|
||||
**6. Can a future video provider consume a multi-turn scene representation
|
||||
without redesigning history?** **Yes.** `packet.build(..., start=, end=)` takes a
|
||||
depth range on the branch and the resulting `scene_id` encodes it
|
||||
(`c7:b3:4-9`). The range is expressed in the coordinates history already uses, so
|
||||
a multi-turn packet is a read of existing structure rather than a new one.
|
||||
|
||||
**7. Does future STT feed editable draft input rather than authoritative state?**
|
||||
**Yes, structurally.** `TranscriptionProvider.transcribe` returns a
|
||||
`DraftTranscription` with `editable=True` and no commit method. A transcriber
|
||||
cannot submit; the ordinary authoritative path is the only way in.
|
||||
|
||||
**8. Is scene/media data branch-safe?** **Yes** — §G. The scene follows the
|
||||
active lineage through Undo, Redo, Retry, divergence, Save Point restore and two
|
||||
genuine process restarts, because it *is* the authoritative state rather than a
|
||||
copy of it. Profiles are campaign-scoped by design and are stable across all of
|
||||
those, which is the correct behaviour for appearance and is argued in §C.3.
|
||||
|
||||
**9. Did M10 create any network dependency?** **No.** No dependency added, no
|
||||
HTTP client, no socket, no subprocess, no provider adapter, no model download, no
|
||||
media endpoint setting. The endpoint *policy* that a future provider will meet is
|
||||
stricter than the one narration uses: loopback only.
|
||||
|
||||
**10. Is there any blocker before M11?** **No blocker.** Two things a reviewer
|
||||
should decide rather than inherit:
|
||||
|
||||
- whether **K04 on the deferred branch** is acceptable for v1, or whether media
|
||||
tables must be physically present (§F);
|
||||
- whether **no browser regression pass** is acceptable given that M10 modified no
|
||||
frontend file (§L.2).
|
||||
|
||||
Neither is a defect; both are decisions that belong to the reviewer.
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user