Files
interactive-story/backend/app/seed.py
JesseMarkowitz 8c65ae99de M2: cut the hosted product away from the local one
94 files, +1,395 -6,578. Three files are new; twenty-four are gone. The
milestone is subtraction, and what is left is the single-user local
storyteller the specification describes.

Removed in full: campaign scripting and its QuickJS sandbox; multi-user
accounts, guest sessions, login, registration and the shared demo key;
the visitor-analytics tables, dashboard and page beacon; the access log
of sign-ins, addresses and devices; per-IP and per-user rate limiting
and quotas; Render deployment config; Postgres and psycopg; cloud
inference providers, the API-key field and the key encryption that
existed to store it; session-cookie signing. None of it was hidden
behind a flag — the routes are gone and answer 404.

Two things were kept that the brief allowed keeping. The `users` table
and its foreign keys stay as an internal ownership detail, because
rewriting them out means a migration across most of the schema to
delete a column that costs nothing; nothing creates a second user and
no request carries an identity. Five inert tables and four inert
columns stay for the same reason, so an M1 campaign database opens
unchanged.

The one addition is app/endpoints.py, which decides where a story may
be sent. Loopback, RFC1918, link-local, unique-local and CGNAT — an
explicit allowlist of networks, not a guess at what `ipaddress` means
by "private", which calls the documentation ranges private and IPv6
loopback reserved. Every address a hostname resolves to must be in it,
so a split answer does not squeak through, and the rule runs both when
the endpoint is saved and before every outbound request, because a name
that resolved to the LAN this morning can resolve elsewhere this
afternoon. Known cloud hosts are named in the refusal so the error says
why rather than looking like broken DNS. TLS is never traded against
it: M1's shared trust context is intact on all four clients and there
is no way to skip verification.

The hardcoded 120-second model timeout is now a setting. That was not
theoretical — on this GPU-less four-core host a cold load of
qwen2.5:3b-instruct took 648.9 seconds to produce the first turn, while
turns 2 to 5 of the same campaign took 3.6 to 13.1. Connect stays short
at 10s so a wrong address still fails fast; the read timeout defaults
to 300s and is bounded at 3600, because "wait longer" must stay a
number.

Two defects found while testing and fixed here. An unknown /api path
fell through the SPA catch-all and came back as HTML with status 200,
so a client asking for JSON parsed a web page instead of learning the
route was gone. And AIDND_CORS_ORIGINS accepted "*", which on an
unauthenticated loopback API would hand every page on the Internet a
write handle on the campaign database; it now refuses to start.

Verified rather than assumed. Offline, on a network with no route out
and no DNS: five turns, retry with both takes retained, restart with an
identical transcript digest, a failed model call leaving the accepted
AI-turn count untouched, and a capture with zero non-loopback unicast
packets. Against a real second machine on the LAN over HTTPS with a
private CA: four turns, restart, and a capture showing 289 packets to
the approved host, 344 loopback, zero anywhere else, zero DNS queries.
Cloud and public endpoints refused with their reasons; no API key
settable; every removed route 404.

604 backend tests pass, down from 648 by the fifteen retired with the
subsystems they tested and up by the twenty-nine added for the endpoint
policy and the removed surface. The scripting tests were not deleted:
eight files used a JavaScript counter as instrumentation for the state
snapshot and rollback machinery, which M2 does not touch, so the
counter moved to the world-state engine and those tests still assert
what they always did. Frontend lint and build are clean; the image
builds, and its wheel-building stage is gone with quickjs.

No M3 work. Undo is still destructive and there is still no Redo.
2026-09-02 11:27:14 -04:00

213 lines
8.0 KiB
Python

"""Seed public demo scenarios on startup.
Every JSON file in ``seed_data/`` describes one demo scenario in the same
model-native shape the export endpoint produces. Seeded scenarios have a NULL
owner and ``is_public=True``, so every visitor (including guests) sees them and
can start an adventure from them, while nobody can edit them. Starting an
adventure copies the scenario's story cards into the adventure.
Seed files are the source of truth for demo content: a scenario is inserted if
missing, reconciled in place when a seed file's content changes, and deleted
when no file claims its title any more, so an edit ships on the next deploy.
Rename a seed by changing its `title` and listing the old one under
`previous_titles`, which moves the rename onto the existing row. When a seed already matches, nothing is written, so
this stays cheap to run on every boot. An adventure already started from a demo
keeps its own copied cards and is unchanged. Only a new adventure
picks up the updated content.
"""
import json
import logging
from pathlib import Path
from sqlalchemy.engine import Engine
from . import models
from .database import SessionLocal
logger = logging.getLogger(__name__)
SEED_DIR = Path(__file__).resolve().parent / "seed_data"
# `title` is in this list so that a rename found through `previous_titles` is
# both detected by `_matches` and written by `_apply_scalars`.
_SCALARS = ("title", "description", "prompt", "memory", "authors_note", "ai_instructions",
"tags", "image", "icon")
_CARD_FIELDS = ("type", "name", "keys", "entry", "notes")
def seed_public_scenarios(engine: Engine) -> None:
if not SEED_DIR.is_dir():
return
files = sorted(SEED_DIR.glob("*.json"))
if not files:
return
db = SessionLocal()
try:
changed = 0
# Every title the files claim, including the ones they used to use. The
# sweep below deletes the seeded rows this set does not name.
claimed: set[str] = set()
complete = True
for path in files:
try:
data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError) as exc:
logger.warning("Skipping seed file %s: %s", path.name, exc)
complete = False
continue
title = (data.get("title") or "").strip()
if not title:
continue
claimed.add(title)
claimed.update(str(old) for old in data.get("previous_titles") or [])
existing = find_seeded(db, title) or _find_renamed(db, data)
if existing is None:
_insert_scenario(db, data)
changed += 1
elif not _matches(existing, data):
_update_scenario(db, existing, data)
changed += 1
changed += _sweep_unclaimed(db, claimed) if complete else 0
if changed:
db.commit()
logger.info("Seeded/updated %d public demo scenario(s).", changed)
except Exception:
db.rollback()
# A seed failure must never take the app down; log and carry on.
logger.exception("Seeding public scenarios failed; continuing without them.")
finally:
db.close()
def _sweep_unclaimed(db, claimed: set[str]) -> int:
"""Deletes seeded scenarios no seed file claims any more, and returns how
many went.
A rename used to strand the row it left behind. `previous_titles` stops new
ones appearing, and this removes the ones already out there, which was
otherwise hand-work on every deployment. Only rows with a NULL owner and
`is_public` are considered, and a player's own scenario is neither, so
nothing anybody created can be reached from here.
An adventure started from a deleted demo survives. `adventures.scenario_id`
is `ON DELETE SET NULL`, so the story and its cards are its
own copies and stay; the adventure loses the cover art it inherited.
The caller skips this when a seed file failed to parse. A file that cannot
be read claims no title, and deleting on that basis would treat a syntax
error as an instruction to remove live content. An empty seed directory
never reaches here at all, for the same reason.
"""
stale = (
db.query(models.Scenario)
.filter(
models.Scenario.user_id.is_(None),
models.Scenario.is_public.is_(True),
models.Scenario.title.notin_(claimed) if claimed else True,
)
.all()
)
for scenario in stale:
logger.info("Removing seeded scenario %r; no seed file claims it.", scenario.title)
db.delete(scenario)
return len(stale)
def _card_tuple(source, get) -> tuple:
return tuple(get(source, f) for f in _CARD_FIELDS)
def find_seeded(db, title: str) -> models.Scenario | None:
"""Returns the seeded scenario with this exact title, if there is one."""
return (
db.query(models.Scenario)
.filter(
models.Scenario.title == title,
models.Scenario.user_id.is_(None),
models.Scenario.is_public.is_(True),
)
.first()
)
def _find_renamed(db, data: dict) -> models.Scenario | None:
"""Returns the row a renamed seed file used to own, so the rename lands on it.
A seed is matched by title, so renaming one inserts a second scenario and
strands the first. The stranded row stays public forever and has to be
deleted by hand on every deployment. List the old title under
`previous_titles` in the seed file and the rename updates the existing row
instead, which also keeps the adventures already started from it pointing at
a scenario that still exists.
Drop a `previous_titles` entry once every deployment has booted past it.
"""
for old in data.get("previous_titles") or []:
found = find_seeded(db, str(old))
if found is not None:
return found
return None
def _matches(scenario: models.Scenario, data: dict) -> bool:
"""True when the DB scenario already equals the seed file, so we can skip
the write and avoid churning rows on every boot."""
if any(getattr(scenario, f) != data.get(f, "") for f in _SCALARS):
return False
if (scenario.stat_schema or None) != (data.get("stat_schema") or None):
return False
have_cards = sorted(_card_tuple(c, lambda o, f: getattr(o, f)) for c in scenario.story_cards)
want_cards = sorted(
_card_tuple(c, lambda o, f: o.get(f, ""))
for c in (data.get("story_cards") or []) if isinstance(c, dict)
)
return have_cards == want_cards
def _insert_scenario(db, data: dict) -> None:
scenario = models.Scenario(user_id=None, is_public=True, title=data.get("title", ""))
_apply_scalars(scenario, data)
db.add(scenario)
db.flush()
_populate_children(db, scenario, data)
def _update_scenario(db, scenario: models.Scenario, data: dict) -> None:
_apply_scalars(scenario, data)
# Replace the child content in full. Demo content is owned by the server and
# cheap to rebuild, and replacing it this way keeps the scenario row, and the
# adventure foreign keys that point at it, intact.
for card in list(scenario.story_cards):
db.delete(card)
db.flush()
_populate_children(db, scenario, data)
def _apply_scalars(scenario: models.Scenario, data: dict) -> None:
for field in _SCALARS:
setattr(scenario, field, data.get(field, ""))
# Phase 12: RPG world-state template (a JSON dict, not a scalar string).
scenario.stat_schema = data.get("stat_schema") or None
def _populate_children(db, scenario: models.Scenario, data: dict) -> None:
for card in data.get("story_cards") or []:
if not isinstance(card, dict):
continue
db.add(
models.StoryCard(
scenario_id=scenario.id,
type=card.get("type", ""),
name=card.get("name", ""),
keys=card.get("keys", ""),
entry=card.get("entry", ""),
notes=card.get("notes", ""),
)
)