Files
interactive-story/backend/app/models.py
T
JesseMarkowitzandClaude Opus 5 1013c94eb1
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
M10: the seam for media, and no media
The media extension contract asks for a scene snapshot a future image or video
provider could be handed: location, who is present, what they hold, what must
stay true, and where in the story it sits. Building one was the milestone's
obvious first task, and it was the wrong one. That snapshot has existed since
M5. `narrative_state["scene"]` holds the summary, the location, the cast and the
coordinate it was written at; a validated `set_scene` event writes it, every
position snapshots it, and every head move restores it. It survives Undo, Redo,
Retry, divergence, Save Point restore and a process restart because it is the
authoritative state rather than a copy of it.

So there is no scenes table here. A second scene store would have been a second
answer to "where is the story now", with its own lineage rules to get wrong —
and the lineage rules are the expensive part, which is the argument for reusing
the ones that already work rather than against it. The Scene Packet is derived
on read, and its identity is computed from the campaign and the position rather
than allocated: the same position yields the same id in another process, after a
restart, and after the packet is thrown away and rebuilt, with no row to keep in
step. That is the part of a future media_assets table that would be expensive to
retrofit, so it is fixed now even though the table is not built.

One table, then: visual_profiles, the only thing the contract's scene list asks
for that nothing already stored. Campaign-scoped and not per-position, because a
character does not change appearance when the story forks — a reader who
diverged would otherwise lose their cast, and the same descriptors would land in
every per-position snapshot, measured at 245 copies of 367 bytes in a 120-turn
campaign to say something that never varies. Keyed by the M5 entity key rather
than a new identity namespace, and one table for characters, locations and items
alike, because a location is an entity with a type and splitting them would
reintroduce the genre shape M5 spent a milestone removing.

What the packet leaves out is the more interesting half. Not the transcript, and
not imported knowledge — none of it, not merely the sources marked hidden. The
rule is what the story established at this position, not everything the narrator
was told, and drawing it by class is what makes it hold for a secret nobody
thought to mark. A hidden Canon source proves it, with a positive control
showing the narrator did receive the sentinel the packet does not carry. Once a
validated event puts the observer in the room, the observer is in the packet:
that is no longer narrator-only knowledge, and a packet that hid it would be
hiding the story from itself.

The providers are contracts and nothing else. Protocols for image, video, audio,
speech and transcription, an empty registry, no adapter, no dependency, no
socket, and no media setting to point anywhere — a setting that exists can be
pointed at a cloud by mistake. A future provider endpoint must be loopback,
stricter than narration's trusted-LAN allowance, because a picture of a scene
carries the scene with it. Transcription returns an editable draft with no
commit method, so STT structurally cannot bypass the authoritative path.

Nothing here can write the story. Not by convention: no module under media/
imports the code that writes state, no media event type exists in the state
vocabulary, and every test in the authority suite compares the authoritative
document byte for byte either side of a media operation — including one where a
provider insists Alice is in a red coat in a corridor, and the campaign goes on
disagreeing.

One defect, found by the milestone's own tests. M10 first added a migration
creating an index that create_all already builds from the column, so an upgraded
database ended up with two indexes and a fresh install with one. Comparing the
two schemas is what caught it; neither database examined alone would have. The
migration is gone rather than renamed, and the right number of migrations for a
new table whose indexes are declared on its columns is zero.

Backend 1,191 passed / 14 skipped / 0 failed, 89 of them M10's. Frontend 145
passed. Lint, production build and Docker build clean. No frontend file changed:
M10 adds no reader-facing surface, and ordinary play — turns, state, memory,
knowledge, Undo, Redo, Retry, Save Point restore, restart — runs with no media
configuration, no warning, no connection attempt and no media row written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qyn3oRd4D6pi72nKBG725B
2026-09-07 03:41:04 -04:00

1255 lines
65 KiB
Python

from datetime import datetime, timezone
from sqlalchemy import (
DDL, JSON, Boolean, DateTime, Float, ForeignKey, Index, Integer, LargeBinary,
String, Text, UniqueConstraint, event,
)
from sqlalchemy.orm import Mapped, Session, mapped_column, relationship
from .compression import CompressedJSON
from .database import Base
from .knowledge import fts as knowledge_fts
def utcnow() -> datetime:
return datetime.now(timezone.utc)
class User(Base):
"""Phase 8: optional accounts.
Three kinds of row share this table:
- The local user has a NULL email and `is_guest` set to False. Single-user
mode creates this row automatically. It owns everything that a database
from before Phase 8 contained.
- Guests have a NULL email and `is_guest` set to True. Multi-user mode
creates one on a visitor's first visit and identifies it only by the
session cookie.
- Registered users have an email. Registration upgrades a guest row in
place, so the guest's data survives without being reassigned.
"""
__tablename__ = "users"
id: Mapped[int] = mapped_column(primary_key=True)
email: Mapped[str | None] = mapped_column(String(320), unique=True, nullable=True)
password_hash: Mapped[str | None] = mapped_column(String(300), nullable=True)
is_guest: Mapped[bool] = mapped_column(Boolean, default=True)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
last_seen_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
# Was the shared demo key's per-day tally. M2 removed the demo key; these
# columns stay so existing databases open unchanged and are never written.
demo_turns_used: Mapped[int] = mapped_column(Integer, default=0)
demo_turns_date: Mapped[str] = mapped_column(String(10), default="")
class Scenario(Base):
__tablename__ = "scenarios"
id: Mapped[int] = mapped_column(primary_key=True)
# NULL owner + is_public = seeded demo content, readable by everyone.
user_id: Mapped[int | None] = mapped_column(
ForeignKey("users.id", ondelete="CASCADE"), nullable=True
)
is_public: Mapped[bool] = mapped_column(Boolean, default=False)
title: Mapped[str] = mapped_column(String(200), default="Untitled Scenario")
description: Mapped[str] = mapped_column(Text, default="")
prompt: Mapped[str] = mapped_column(Text, default="")
# Plot components (AI Dungeon terminology; `memory` == Plot Essentials)
memory: Mapped[str] = mapped_column(Text, default="")
authors_note: Mapped[str] = mapped_column(Text, default="")
ai_instructions: Mapped[str] = mapped_column(Text, default="")
tags: Mapped[str] = mapped_column(String(500), default="")
# Cover art. The value is either an "https://" URL or an inline
# "data:image/...;base64,..." URI. The editor downscales uploads before
# storing them. An empty value tells the UI to fall back to an emoji sigil
# or to generated art. The image is stored in the row rather than on disk,
# because Render's free tier provides no persistent volume. Storing it here
# also keeps export bundles self-contained.
image: Mapped[str] = mapped_column(Text, default="")
# A single emoji or glyph, used when `image` is empty. This is a separate
# column because the value is a character rather than a location, so
# nothing needs to fetch or cache it.
icon: Mapped[str] = mapped_column(String(16), default="")
# Phase 12: the RPG world-state template. It holds stat definitions, which
# include bands and rules, and milestones. A NULL or empty value means the
# scenario has no RPG layer.
stat_schema: Mapped[dict | None] = mapped_column(JSON, nullable=True)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
updated_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow, onupdate=utcnow)
story_cards: Mapped[list["StoryCard"]] = relationship(
back_populates="scenario", cascade="all, delete-orphan"
)
adventures: Mapped[list["Adventure"]] = relationship(back_populates="scenario")
class Adventure(Base):
__tablename__ = "adventures"
id: Mapped[int] = mapped_column(primary_key=True)
user_id: Mapped[int | None] = mapped_column(
ForeignKey("users.id", ondelete="CASCADE"), nullable=True
)
scenario_id: Mapped[int | None] = mapped_column(
ForeignKey("scenarios.id", ondelete="SET NULL"), nullable=True
)
title: Mapped[str] = mapped_column(String(200), default="Untitled Adventure")
memory: Mapped[str] = mapped_column(Text, default="")
authors_note: Mapped[str] = mapped_column(Text, default="")
ai_instructions: Mapped[str] = mapped_column(Text, default="")
# A convenience mirror of whichever summary is eligible at the current
# position, and **never** an input to anything authoritative (M6 corrective,
# review finding M6-F1).
#
# It exists because the Plot panel lets a reader read and edit the summary
# and the export bundle carries it. It is not a store: `summaries` rows are,
# and `summaries.current` decides which one the story is entitled to. This
# column has no lineage of its own, so anything that reads it as truth
# inherits whatever was written last, on whatever line — which is exactly
# how abandoned prose reached an active prompt before the correction.
#
# Kept in step by `summaries.record` when one is written and by
# `attempts.restore_state` when the head moves.
story_summary: Mapped[str] = mapped_column(Text, default="")
# Phase 18: who the player is playing as. The AI never writes these — they
# are user-only, which is what lets them sit in the cached system block
# rather than below the history with the values that change. An empty
# `persona_name` means the adventure has no persona, and every read below
# falls back to the wording used before this existed.
#
# These are adventure columns rather than part of the scenario's
# `stat_schema`, for two reasons. An adventure with no RPG layer still has a
# protagonist, and `worldstate.schema._initials` treats every dict inside a
# stat section as a stat definition, so a persona placed there would be
# instantiated, rendered and given an `initial` value as though it were one.
persona_name: Mapped[str] = mapped_column(String(80), default="")
persona_pronouns: Mapped[str] = mapped_column(String(40), default="")
persona_desc: Mapped[str] = mapped_column(Text, default="")
# Was the campaign scripting engine's shared `state` object. M2 removed
# scripting; the column stays so existing databases open unchanged, and it
# is never written with anything but an empty dict.
script_state: Mapped[dict] = mapped_column(JSON, default=dict)
# Phase 12: live RPG world state (world/player/npc stats + milestones),
# instantiated from the scenario's stat_schema. Empty when there's no RPG layer.
#
# **Legacy as of M5**, and no longer authoritative. M5 replaced the
# relative-delta protocol this column served (ADR 010); the turn engine no
# longer writes it, and nothing reads it to decide anything. It stays so
# that a pre-M5 database opens unchanged and its numbers remain visible to
# whoever wants to look — `narrative_state` below is what the story means
# now. Reinterpreting these values as generic narrative facts would be
# inventing meaning the data does not carry, which the M5 brief forbids.
world_state: Mapped[dict] = mapped_column(JSON, default=dict)
# M5: the authoritative narrative state, as it stands at the active head.
# Genre-neutral (ADR 006), written only by validated typed events (ADR 010),
# and restored from the destination node's snapshot whenever the head moves,
# so it always describes the story being read rather than a story the reader
# has stepped back from.
narrative_state: Mapped[dict] = mapped_column(
CompressedJSON, nullable=True, default=None
)
# Campaign canon: rules the story may not contradict, as configuration
# rather than code (C01, J03). A fantasy campaign forbidding resurrection
# and a science-fiction one forbidding faster-than-light travel use the same
# field and the same validator; neither word appears in the application.
campaign_canon: Mapped[dict | None] = mapped_column(JSON, nullable=True)
@property
def canon_rules(self) -> list[str]:
"""The `rules` list alone, which is the half a person writes.
`campaign_canon` also carries `forbidden_status_changes`, a structured
shape the browser has no editor for and does not need one for — a rule
like "nothing dead becomes alive" is expressible as a sentence. So the
API exposes the sentences and leaves the structured half to whatever
wrote it, rather than round-tripping a shape the UI would flatten.
"""
canon = self.campaign_canon
if not isinstance(canon, dict):
return []
rules = canon.get("rules")
return [r for r in rules if isinstance(r, str)] if isinstance(rules, list) else []
# The ${Placeholder} answers collected when this adventure was started, kept
# so "Update from scenario" can re-fill freshly copied scenario text with the
# same values. NULL for adventures created before this column existed.
placeholders: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# Phase 6: opt-in per adventure (extra AI calls)
auto_summarize: Mapped[bool] = mapped_column(Boolean, default=False)
memory_bank_enabled: Mapped[bool] = mapped_column(Boolean, default=False)
# Phase 14, SP3: how far the memory pass and the summary pass have read,
# each as a coordinate. Each pair
# holds the branch and depth of the last action that pass covered. A
# position moves when an action in front of it is deleted, so the mark
# silently starts covering an action it never read. A depth is a coordinate
# along a path, so deleting an action does not move it. NO_DEPTH, which is
# -1, means that nothing is covered yet, so the first block needs no special
# case. These are plain integers rather than foreign keys, for the reason
# given on `head_branch_id` below. See `context/cursors.py`.
memory_cursor_branch_id: Mapped[int | None] = mapped_column(Integer, nullable=True)
memory_cursor_depth: Mapped[int] = mapped_column(Integer, default=-1)
summary_cursor_branch_id: Mapped[int | None] = mapped_column(Integer, nullable=True)
summary_cursor_depth: Mapped[int] = mapped_column(Integer, default=-1)
# Phase 14: where the story is being played. `head_branch_id` names the
# branch, and `head_depth` gives the depth of its newest node.
#
# `head_branch_id` is deliberately not a ForeignKey. `branches.adventure_id`
# already points from branches to adventures, so a constraint in this
# direction would make the two tables a cycle that `create_all` cannot
# order. The usual fix is `use_alter`, which needs an ALTER statement that
# SQLite does not provide. The column caches a pointer, and
# `tree.head_branch` treats a head that names a missing branch as a bug to
# recover from rather than a state to preserve.
head_branch_id: Mapped[int | None] = mapped_column(Integer, nullable=True)
# The depth of the tip, so the next node is always head_depth + 1.
# NO_DEPTH (-1) for an adventure with no actions yet.
head_depth: Mapped[int] = mapped_column(Integer, default=-1)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
updated_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow, onupdate=utcnow)
scenario: Mapped[Scenario | None] = relationship(back_populates="adventures")
story_cards: Mapped[list["StoryCard"]] = relationship(
back_populates="adventure", cascade="all, delete-orphan"
)
# Every action in the adventure, across all branches. This collection is
# the tree, not the story being played. Ordering it by depth does not make
# it a story, because a path is a selection out of the tree. Code that shows
# a story to a reader goes through `context.history`, which applies the
# branch clause. This relationship exists for ownership and for the
# delete-orphan cascade.
actions: Mapped[list["Action"]] = relationship(
back_populates="adventure",
cascade="all, delete-orphan",
order_by="Action.id",
)
memories: Mapped[list["Memory"]] = relationship(
back_populates="adventure",
cascade="all, delete-orphan",
order_by="Memory.id",
)
# M6: the lineage-anchored generated summaries, newest last.
summaries: Mapped[list["Summary"]] = relationship(
back_populates="adventure",
cascade="all, delete-orphan",
order_by="Summary.id",
)
derived_status: Mapped[list["DerivedStatus"]] = relationship(
back_populates="adventure",
cascade="all, delete-orphan",
order_by="DerivedStatus.id",
)
# M7: the imported knowledge library. Campaign-scoped by construction —
# there is no path from one campaign's sources to another's.
knowledge_sources: Mapped[list["KnowledgeSource"]] = relationship(
back_populates="adventure",
cascade="all, delete-orphan",
order_by="KnowledgeSource.id",
)
# M10: how the campaign's entities look. Derived presentation metadata, not
# story state — see `VisualProfile`.
visual_profiles: Mapped[list["VisualProfile"]] = relationship(
back_populates="adventure",
cascade="all, delete-orphan",
order_by="VisualProfile.id",
)
class Branch(Base):
"""Phase 14: one path through an adventure's story tree.
A branch does not own a copy of the story. It holds the nodes played on it,
and it inherits everything before its fork point from its ancestors. Reading
branch C means reading C's nodes, then B's nodes up to the depth where C
forked, then A's nodes up to the depth where B forked. The `lineage` column
records that list, so a read becomes one OR clause per entry instead of a
walk up parent pointers.
Until forking ships, each adventure has one root branch and every node
belongs to it. This is not a partly migrated state. A linear story is a tree
with one branch, which is why writing these columns changes nothing that a
reader can observe.
This class defines no ORM relationships, by design. `actions.branch_id` and
`memories.branch_id` both use ON DELETE CASCADE, so the database removes a
deleted branch's nodes. A relationship would make SQLAlchemy load those rows
first, and loading every action of a branch is what the windowed reads exist
to avoid.
"""
__tablename__ = "branches"
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
# NULL on a root branch.
parent_branch_id: Mapped[int | None] = mapped_column(
ForeignKey("branches.id", ondelete="CASCADE"), nullable=True
)
# The depth at which this branch left its parent. The fork records this
# value, and no code infers it later. Deriving it from the first depth where
# two branches' nodes differ would produce a wrong answer whenever an
# attempt repeats its parent's text.
fork_depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
# The ancestry, newest first, as [[branch_id, max_depth], ...]. A NULL
# `max_depth` means the entry extends to the tip of that branch. Any other
# value is the fork depth of the branch below it, inclusive. The fork
# computes this list once from the parent's lineage plus one entry, so no
# read has to reconstruct it.
lineage: Mapped[list] = mapped_column(JSON, default=list)
# The name the player gave this line of the story, or NULL if no one named
# it. The column stores NULL rather than a generated name such as
# "branch 4", because it records what the player chose rather than what the
# app derived. A stored default would also become wrong as soon as an
# earlier branch is deleted and the ordinals shift. The UI labels an unnamed
# branch by its fork depth, which deleting a branch does not change.
name: Mapped[str | None] = mapped_column(String(80), nullable=True)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
# M3: the disposition `DATA-MODEL.md` §5 gives a branch, stored as the fact
# that produced it. NULL means active. A value means a divergent write left
# this branch at `superseded_depth`, so its nodes past that depth are
# retained history that no active head is reading.
#
# Nothing reads these to decide behaviour. Redo follows the lineage, so a
# wrong value here cannot make the story wrong; they exist for the cleanup
# and discarded-history features that `STORY-BRANCH-SEMANTICS.md` §28-29
# leave to a later version. See `head.mark_superseded`.
superseded_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
superseded_depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
class Checkpoint(Base):
"""M4: a Save Point — a durable named pointer to a story position.
"Save Point" is what the user reads; `checkpoint` is what the code calls it
(`BROWSER-UX-SPEC.md` §23).
The row holds a name and a coordinate, and no story. `DATA-MODEL.md` §8
describes the pointer as naming a turn; the coordinate here is
`(branch_id, depth)`, which is what M3 made the head and what
`head.node_at` resolves. Restoring one is therefore head movement with a
bounds check rather than a restore system of its own — see ADR 012 and
`head.move_to_node`.
A coordinate rather than an action id, deliberately. One coordinate can
hold several attempts at a turn and exactly one of them is live, so a
retry replaces the row a Save Point would have pinned. "Turn 42 of this
line" survives a retry; "action 918" would point at a take the story no
longer tells.
`branch_id` is the branch the node itself sits on, not the branch that was
being read when the Save Point was made. Those differ whenever the head is
resting in a shared prefix, and the node's own branch is the one that still
names the position after the reader has moved elsewhere.
**Nothing removes a Save Point but the user.** They are not cleaned up for
going stale, for being behind the head, or for pointing into a future the
story has left (`STORY-BRANCH-SEMANTICS.md` §19).
That includes deleting a branch. `branch_id` carries `ON DELETE CASCADE` as
referential integrity — a Save Point must never point at a branch that is
gone — but the branch endpoint refuses to delete a branch any Save Point
names, so the cascade does not fire through the application
(`routers/adventures/branches.py`, `STORY-BRANCH-SEMANTICS.md` §19.1). The
user deletes the Save Point first, which deletes no story, and then the
branch. Deleting the whole campaign does cascade, and should: that is what
the user asked for.
"""
__tablename__ = "checkpoints"
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE")
)
name: Mapped[str] = mapped_column(String(120), default="")
# `DATA-MODEL.md` §8's optional notes, and `BROWSER-UX-SPEC.md` §24's
# optional second field. Empty is the ordinary case.
note: Mapped[str] = mapped_column(Text, default="")
branch_id: Mapped[int] = mapped_column(
ForeignKey("branches.id", ondelete="CASCADE")
)
depth: Mapped[int] = mapped_column(Integer)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
# Bumped by a rename, which is the only edit a Save Point allows. The
# coordinate is never rewritten: `STORY-BRANCH-SEMANTICS.md` §24 keeps a
# Save Point's meaning auditable by making "move it" delete-and-recreate.
updated_at: Mapped[datetime] = mapped_column(
DateTime, default=utcnow, onupdate=utcnow
)
class StateProposal(Base):
"""M5: what the model proposed, and what the application did about it.
`DATA-MODEL.md` §19 requires the model's proposal to be *distinct from*
accepted state, and this table is that separation made physical. The model
writes here; it never writes `state_events`, and it never writes a snapshot.
A row exists whether the proposal was accepted, partly accepted, rejected or
unparseable. A rejected proposal is not authoritative and changes nothing,
but it is the record that explains why the state does not say what the
narration seems to say — without it, a wrong-looking campaign has no trail
to follow. `raw_output` is kept for exactly the case that matters most: the
block that did not parse, which no structured column could hold.
"""
__tablename__ = "state_proposals"
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE")
)
# The node whose narration produced this. NULL only for a manual correction,
# which has a coordinate but no narration behind it.
action_id: Mapped[int | None] = mapped_column(
ForeignKey("actions.id", ondelete="CASCADE"), nullable=True
)
branch_id: Mapped[int | None] = mapped_column(Integer, nullable=True)
depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Which model produced it, so a later comparison of extraction quality has
# something to group by. Empty for a manual correction.
model_name: Mapped[str] = mapped_column(String(200), default="")
# `accepted_story` or `manual_correction` — who is asserting this.
source: Mapped[str] = mapped_column(String(40), default="accepted_story")
# accepted | partially_accepted | rejected | unparseable
status: Mapped[str] = mapped_column(String(30), default="accepted")
# The block as written, including when it did not parse.
raw_output: Mapped[str] = mapped_column(Text, default="")
# The parsed payload, the events accepted, and every rejection with its
# reason. Compressed for the same reason the prompt is: a busy turn's
# rejections are the largest thing here and nothing reads them in bulk.
detail: Mapped[dict | None] = mapped_column(
CompressedJSON, nullable=True, deferred=True
)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
class StateEvent(Base):
"""M5: one accepted change to the authoritative narrative state.
The audit half of `DATA-MODEL.md` §17's hybrid. Append-only, ordered, and
**never read to reconstruct state** — that is the snapshot's job, and mixing
the two would make restore proportional to campaign length, which ADR 012
and M4 both forbid.
What this table answers is §8's list: what changed, why, which turn caused
it, whether a model or the user asserted it, and what the value was before.
`before` is stored per event rather than derived, because deriving it would
mean replaying — the thing the hybrid exists to avoid.
Events carry the story coordinate as well as the action id. The coordinate
survives a retry replacing the live take at that position, exactly as a Save
Point's does; the action id says which attempt actually proposed it.
"""
__tablename__ = "state_events"
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE")
)
proposal_id: Mapped[int | None] = mapped_column(
ForeignKey("state_proposals.id", ondelete="SET NULL"), nullable=True
)
action_id: Mapped[int | None] = mapped_column(
ForeignKey("actions.id", ondelete="CASCADE"), nullable=True
)
branch_id: Mapped[int | None] = mapped_column(Integer, nullable=True)
depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Order within one proposal, so a turn's events replay for a reader in the
# order they were applied.
sequence: Mapped[int] = mapped_column(Integer, default=0)
event_type: Mapped[str] = mapped_column(String(60), default="")
payload: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# What the affected value was immediately before this event, so the audit
# can answer "what did it used to be" without reconstruction. NULL when the
# event established something that did not exist.
before: Mapped[dict | None] = mapped_column(JSON, nullable=True)
source: Mapped[str] = mapped_column(String(40), default="accepted_story")
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
class Memory(Base):
"""Phase 6: an auto-summarized (or hand-written) fact about the adventure.
The vector lives in `embedding_blob` as packed float32 (see vectors.py).
NULL until embedded, which also marks it for backfill when an embedding
model becomes available.
Cosine ranking runs in Python, so the vectors travel over the wire. Measure
that cost in bytes rather than in rows. A few hundred vectors stored as JSON
come to about 10 MB, fetched again on every turn. Size any new column by the
bytes it adds, not by the number of rows.
"""
__tablename__ = "memories"
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
text: Mapped[str] = mapped_column(Text, default="")
# The vector, stored as little-endian float32. This column is deferred
# because it is wider than the rest of the row combined and only one code
# path reads it. Code that loads memories in bulk, such as the Memories
# drawer, eviction, and the embed queue, must select the columns it needs
# instead of loading whole entities.
embedding_blob: Mapped[bytes | None] = mapped_column(
LargeBinary, nullable=True, deferred=True
)
# The stretch of story this memory summarizes, given as depths on
# `branch_id`. Both are NULL for a hand-written memory, which summarizes no
# actions. `source_end` is the depth of the node the memory
# attaches to, and `depth` below mirrors it. `source_start` is where the
# stretch begins, which is where the summarizer resumes if the memory is
# withdrawn.
source_start: Mapped[int | None] = mapped_column(Integer, nullable=True)
source_end: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Phase 14: the node that produced this memory, meaning the last action the
# memory summarizes. Derived data attaches to the node it came from, which
# is what makes forking cheap. Memories on a shared ancestor are shared
# automatically, and a memory that covers part of branch B is not visible
# from any path that does not go through B.
#
# Every memory has a coordinate, including a hand-written one, which takes
# the head as of the moment it was written (SP7). A NULL depth used to mean
# that the memory belonged to the adventure rather than to a path. A fork
# cannot cap a NULL, so such a memory followed the reader onto branches
# whose story it did not describe.
branch_id: Mapped[int | None] = mapped_column(
ForeignKey("branches.id", ondelete="CASCADE"), nullable=True
)
depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Whether `embedding_blob` is set. `memorybank.set_vector` keeps this column
# current. Readers need only the yes-or-no answer, and fetching six
# kilobytes of vector to get it is too expensive.
embedded: Mapped[bool] = mapped_column(Boolean, default=False)
# M6: how much weight the narrator should give this memory
# (`CONTEXT-AND-MEMORY.md` §14). `accepted_story` is something the story
# actually established; `heuristic` is an interpretation of it. The
# application owns this classification — the extractor may hint, but
# `memorybank.classify_authority` decides — so a guess can never become
# canon merely by being written down. Authoritative state changes still go
# only through the M5 event path (ADR 013).
authority: Mapped[str] = mapped_column(String(20), default="accepted_story")
pinned: Mapped[bool] = mapped_column(Boolean, default=False)
forgotten: Mapped[bool] = mapped_column(Boolean, default=False) # evicted, kept for UI
use_count: Mapped[int] = mapped_column(Integer, default=0)
last_used_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
adventure: Mapped[Adventure] = relationship(back_populates="memories")
class Summary(Base):
"""M6: one generated rolling summary, anchored to the story it summarizes.
The inherited design kept the summary in a single `adventures.story_summary`
column with a lineage cursor recording how far it had read. The cursor was
lineage-aware; the prose was not. After an Undo and a divergence the column
still held sentences describing the abandoned line, and the context builder
injected it unconditionally — the leak `STORY-BRANCH-SEMANTICS.md` §32 and
acceptance test E03 forbid.
A summary is therefore a row on a path, exactly as a `Memory` is, and it is
filtered through the same `lineage.Path.clause` chokepoint. `branch_id` and
`depth` are the coordinate it was written at; `source_start`/`source_end`
are the stretch of story it covers. A summary whose coordinate is not on the
active capped lineage is not eligible, and is never deleted for it — the
abandoned line keeps its own derived data (§11).
"""
__tablename__ = "summaries"
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
text: Mapped[str] = mapped_column(Text, default="")
# The coordinate this summary was written at: the last node it covers.
branch_id: Mapped[int | None] = mapped_column(
ForeignKey("branches.id", ondelete="CASCADE"), nullable=True
)
depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
# The stretch of story it summarizes, as depths on `branch_id`.
source_start: Mapped[int | None] = mapped_column(Integer, nullable=True)
source_end: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Why it was generated: "interval" for the automatic pass, "manual" when the
# reader wrote or edited it themselves.
trigger: Mapped[str] = mapped_column(String(20), default="interval")
model_name: Mapped[str] = mapped_column(String(200), default="")
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
adventure: Mapped[Adventure] = relationship(back_populates="summaries")
class DerivedStatus(Base):
"""M6: the outcome of one kind of background derived work, per campaign.
M2 shipped with the whole memory bank dead and the suite green: the
summariser and the embedder raised inside a fire-and-forget task, and
nothing recorded it (`BUILD-MILESTONES.md`, note from M2). Derived work is
allowed to fail — the accepted turn, the state and the head must all
survive it — but it is not allowed to fail *invisibly*.
One row per (adventure, kind), rewritten in place. This is deliberately not
a job queue: it records what happened last, so a reader can see that
memories stopped being written and why, and so a maintainer can retry.
"""
__tablename__ = "derived_status"
__table_args__ = (UniqueConstraint("adventure_id", "kind", name="uq_derived_kind"),)
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
# "memory", "summary" or "embedding".
kind: Mapped[str] = mapped_column(String(20))
# "ok" (did work), "idle" (ran, nothing pending) or "failed".
status: Mapped[str] = mapped_column(String(20), default="ok")
detail: Mapped[str] = mapped_column(Text, default="")
failures: Mapped[int] = mapped_column(Integer, default=0)
last_attempt_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
last_success_at: Mapped[datetime | None] = mapped_column(DateTime, nullable=True)
adventure: Mapped[Adventure] = relationship(back_populates="derived_status")
class KnowledgeSource(Base):
"""M7: one local file the reader imported as campaign knowledge.
A first-class record rather than a Story Card. Phase 0B found Story Cards
could not carry what an imported-knowledge system needs — classification,
provenance, a content identity, a lifecycle, chunking, or an index — and
`IMPORTED-KNOWLEDGE-DESIGN.md` §73 settles that they are not the production
store. Nothing here writes a Story Card and nothing reads one.
Two things about a source are **not** derivable and must survive anything:
the accepted content and its classification. Everything else here is either
metadata about where it came from or a description of derived work that can
be rebuilt (`chunks`, the FTS rows, `KnowledgeEmbedding`).
## Why the content is in the column
`IMPORTED-KNOWLEDGE-DESIGN.md` §11 requires the campaign to stop depending
on the original file the moment the import succeeds. Two designs satisfy
that: copy the bytes into an application-owned directory with the database
as metadata authority, or store the text here. This build stores the text.
It is the simpler of the two by some distance — one transaction covers the
source, its chunks and its index, so a failed import cannot leave a file
behind with no row or a row with no file; export carries the content with no
second archive format; and there is no directory whose contents can drift
away from the rows describing them. Sources are capped at
`knowledge.MAX_SOURCE_BYTES`, so the column stays small enough for that to
be the right trade.
`original_filename` is metadata and nothing else. **It is never used as a
path.** The import surface is an HTTP upload, so no backend pathname is ever
accepted in the first place (H08); see `knowledge/importer.py`.
"""
__tablename__ = "knowledge_sources"
id: Mapped[int] = mapped_column(primary_key=True)
# Campaign-scoped, and only campaign-scoped: `IMPORTED-KNOWLEDGE-DESIGN.md`
# §65-66 make cross-campaign retrieval a defect, not a missing feature.
# There is deliberately no branch coordinate. An imported file is campaign
# source material; it does not become a different file because the story
# forked (`CONTEXT-AND-MEMORY.md` §39). Nothing in M7 derives a knowledge
# record from story history, which is the only case that would need one.
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
)
title: Mapped[str] = mapped_column(String(200), default="")
original_filename: Mapped[str] = mapped_column(String(255), default="")
# "canon", "reference" or "inspiration". Exactly one, always set, editable
# without reimport. This is semantic, not cosmetic: it decides the framing
# the chunk is given in the prompt, the weight it carries in ranking, and
# which budget it competes in.
classification: Mapped[str] = mapped_column(String(20), default="reference")
enabled: Mapped[bool] = mapped_column(Boolean, default=True)
# "normal" or "hidden". Hidden is narrator-only knowledge — the secret a
# mystery turns on. It is not a permission system: the person who imported
# the file can always read it here. It means the protagonist does not know
# it, and the prompt says so (`IMPORTED-KNOWLEDGE-DESIGN.md` §67-69).
visibility: Mapped[str] = mapped_column(String(20), default="normal")
# Canon that must be considered whether or not it resembles the query —
# "resurrection is impossible" does not stop applying because nobody said
# the word (`CONTEXT-AND-MEMORY.md` §41-42). Canon only, and it still costs
# measured budget and still appears in provenance.
always_include: Mapped[bool] = mapped_column(Boolean, default=False)
# SHA-256 of the normalized text. Identity, and the duplicate test.
content_hash: Mapped[str] = mapped_column(String(64), default="", index=True)
# The accepted source text, exactly as it was decoded. Not the normalized
# form: the reader inspects what they imported.
content: Mapped[str] = mapped_column(Text, default="")
byte_size: Mapped[int] = mapped_column(Integer, default=0)
media_type: Mapped[str] = mapped_column(String(80), default="text/plain")
# What produced the chunks now on disk, so a later parser change can be
# detected rather than guessed at.
parser_version: Mapped[int] = mapped_column(Integer, default=1)
chunking_version: Mapped[int] = mapped_column(Integer, default=1)
# The lexical half: "ready" once chunks and FTS rows are committed,
# "failed" if building them raised. A source is retrievable only when this
# is "ready", which is what makes a half-built import unreachable rather
# than ambiguous (`IMPORTED-KNOWLEDGE-DESIGN.md` §57).
index_state: Mapped[str] = mapped_column(String(20), default="pending")
index_detail: Mapped[str] = mapped_column(Text, default="")
# The semantic half, kept separate on purpose. Lexical retrieval is a
# supported production path, not a fallback, so a source whose embeddings
# failed still says "lexical available, semantic failed" rather than
# reporting one health for both.
embed_state: Mapped[str] = mapped_column(String(20), default="idle")
embed_detail: Mapped[str] = mapped_column(Text, default="")
notes: Mapped[str] = mapped_column(Text, default="")
imported_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
updated_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow, onupdate=utcnow)
adventure: Mapped[Adventure] = relationship(back_populates="knowledge_sources")
chunks: Mapped[list["KnowledgeChunk"]] = relationship(
back_populates="source",
cascade="all, delete-orphan",
order_by="KnowledgeChunk.chunk_index",
)
class KnowledgeChunk(Base):
"""M7: one retrievable passage of an imported source.
Derived data. Deleting every chunk of a source and rebuilding it from
`KnowledgeSource.content` must produce the same chunks in the same order —
the chunker is deterministic — which is what makes reindexing safe and what
lets an export carry the source alone.
`adventure_id` is denormalized from the source. Retrieval filters by
campaign on every query, and carrying the column here means the FTS join
reaches the campaign scope without a third table in the hot path.
"""
__tablename__ = "knowledge_chunks"
id: Mapped[int] = mapped_column(primary_key=True)
source_id: Mapped[int] = mapped_column(
ForeignKey("knowledge_sources.id", ondelete="CASCADE"), index=True
)
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
)
chunk_index: Mapped[int] = mapped_column(Integer, default=0)
# The Markdown heading trail above this passage, joined with " > ". Empty
# for plain text and for a passage above the first heading. It is carried
# into the prompt, because "Old Abbey > The Crypt" is most of what tells the
# narrator what the passage is about.
heading_path: Mapped[str] = mapped_column(Text, default="")
text: Mapped[str] = mapped_column(Text, default="")
token_count: Mapped[int] = mapped_column(Integer, default=0)
content_hash: Mapped[str] = mapped_column(String(64), default="")
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
source: Mapped[KnowledgeSource] = relationship(back_populates="chunks")
embedding: Mapped["KnowledgeEmbedding | None"] = relationship(
back_populates="chunk", cascade="all, delete-orphan", uselist=False
)
# M7: the FTS5 lexical index travels with the table it indexes.
#
# An FTS5 table is a virtual table, and SQLAlchemy's metadata has no way to
# describe one — so left to itself, `create_all` would build every knowledge
# table and no index, and `drop_all` would leave the index behind holding
# rowids for chunks that no longer exist. Hanging the DDL off
# `knowledge_chunks` fixes both ends at once: the index is created with the
# table it points at, and dropped before it, on every path that builds or tears
# down a schema — a fresh install, an existing database gaining the M7 tables,
# and a test's setup and teardown.
#
# `execute_if(dialect="sqlite")` because FTS5 is SQLite's. This build stores
# campaigns in SQLite and nothing else; the Postgres branches elsewhere in the
# tree are inherited from upstream and unused (`DEVELOPMENT.md`).
event.listen(
KnowledgeChunk.__table__,
"after_create",
DDL(knowledge_fts.DDL).execute_if(dialect="sqlite"),
)
event.listen(
KnowledgeChunk.__table__,
"before_drop",
DDL(f"DROP TABLE IF EXISTS {knowledge_fts.TABLE}").execute_if(dialect="sqlite"),
)
class KnowledgeEmbedding(Base):
"""M7: the vector for one chunk, with enough metadata to distrust it.
A separate table rather than a column on the chunk, for one reason: it makes
the rebuildable boundary a table boundary. "Rebuild the semantic index" is
`DELETE FROM knowledge_embeddings`, and nothing about the source, its
classification or its chunks is in the blast radius.
`model` and `dimensions` are what make a stale vector detectable rather than
silently wrong. `vectors.cosine` already refuses to score two vectors of
different lengths, but a same-width vector from a different model would
score plausible nonsense, so retrieval checks the model name too.
"""
__tablename__ = "knowledge_embeddings"
id: Mapped[int] = mapped_column(primary_key=True)
chunk_id: Mapped[int] = mapped_column(
ForeignKey("knowledge_chunks.id", ondelete="CASCADE"), unique=True, index=True
)
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
)
# Little-endian float32, the same packing the memory bank uses (vectors.py).
vector: Mapped[bytes] = mapped_column(LargeBinary)
model: Mapped[str] = mapped_column(String(200), default="")
dimensions: Mapped[int] = mapped_column(Integer, default=0)
# What the vector was computed against. A parser or chunker change moves the
# text under the vector, and these say so without re-reading the chunk.
parser_version: Mapped[int] = mapped_column(Integer, default=1)
chunking_version: Mapped[int] = mapped_column(Integer, default=1)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
chunk: Mapped[KnowledgeChunk] = relationship(back_populates="embedding")
class VisualProfile(Base):
"""M10: how one entity looks, so a future depiction can be consistent.
The only thing M10 persists, and the reason is that it was the only thing
the media contract asks for that nothing already stored. The scene snapshot
§5 asks for already exists as `narrative_state["scene"]` and has since M5;
building a second one beside it would have been a duplicate representation
with its own lineage rules to get wrong.
## Not story state, and structurally so
A visual profile is **presentation metadata**. Nothing here is a fact the
story established: `MEDIA-EXTENSION-CONTRACT.md` §35 and §37 are explicit
that a depiction — and therefore a description written to guide one — must
never become canon on its own, and that promoting a visual detail into canon
would have to be a deliberate act by the reader.
So these rows are deliberately **outside** the M5 pipeline. They are not
events, they are not validated by `narrative/validate.py`, they are not in
the state document, and they are not snapshotted per position. Writing one
cannot change `narrative_state`, because nothing in `media/` imports the
code that may. That is the guarantee, and it is a structural one rather than
a rule somebody has to remember.
## Campaign-scoped, not per-position — which is the interesting decision
Every other derived record in this schema carries a `(branch_id, depth)`
coordinate, because it describes a *moment*: a memory summarises a stretch,
a summary covers a range, a snapshot records an outcome. A visual profile
describes none of those. It says what someone looks like, and a character
does not change appearance because the story forked.
Making it per-position would have been actively wrong twice over. It would
have meant a profile written on one branch was invisible on another, so a
reader who diverged would lose their cast's appearance — the opposite of the
continuity the profile exists for. And it would have put a descriptor
document into every per-position state snapshot, which M9 measured as
already 74% of a campaign bundle; the profiles would have been duplicated
once per turn to say something that never varies.
So the key is `(adventure_id, entity_key)` and there is exactly one profile
per entity per campaign. It is stable across Undo, Redo, Save Point restore
and divergence for the same reason it is simple: there is nothing there to
move.
## `entity_key` is the M5 key, and no second identity namespace
The key is the entity key the narrative state already uses — `"mara"`,
`"the_office"`, `"silver_key"` — not a new id, not a name, and not a media
identifier. `MEDIA-EXTENSION-CONTRACT.md` §7-9 describe character, location
and item profiles separately; this is one table for all three, because M5's
entity model is genre-neutral by design (`DATA-MODEL.md` §9) and a
character, a location, an item, a vehicle and a spaceship are all entities
with a `type`. Splitting them here would have reintroduced the genre shape
M5 spent a milestone removing.
There is no `kind` column for the same reason: the entity already has a
`type`, and storing it again would be a second source of truth for one fact.
## The columns, and why they are shaped this way
The contract's examples are fantasy-shaped — hair, eyes, build; architecture,
hearths, oil lamps — and the brief is explicit that they are examples rather
than a schema. A fixed column per fantasy attribute would not hold an
orbital station, a corporate office or a car.
So: `descriptors` is an open map of trait to value, `features` is a list of
distinctive visible things, and `style_notes` is free text about how it
should be rendered. `{"hair": "dark auburn"}` and
`{"hull": "pitted white composite"}` are the same shape, and neither needed
a migration to become possible.
"""
__tablename__ = "visual_profiles"
__table_args__ = (
# One profile per entity per campaign. The uniqueness is the model: a
# second profile for the same entity would be a second answer to "what
# does this look like", with nothing to decide between them.
UniqueConstraint("adventure_id", "entity_key", name="uq_visual_entity"),
)
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE"), index=True
)
#: The narrative-state entity key. Not a display name: two characters may
#: share a name, and M9's report recorded that the state model permits it.
entity_key: Mapped[str] = mapped_column(String(200))
#: Trait -> value. Open by construction; see the class docstring.
descriptors: Mapped[dict] = mapped_column(JSON, default=dict)
#: Distinctive visible things, as short phrases.
features: Mapped[list] = mapped_column(JSON, default=list)
#: How it should be rendered, rather than what it is.
style_notes: Mapped[str] = mapped_column(Text, default="")
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
updated_at: Mapped[datetime] = mapped_column(
DateTime, default=utcnow, onupdate=utcnow
)
adventure: Mapped[Adventure] = relationship(back_populates="visual_profiles")
class StoryCard(Base):
"""Owned by either a scenario or an adventure (exactly one set)."""
__tablename__ = "story_cards"
id: Mapped[int] = mapped_column(primary_key=True)
scenario_id: Mapped[int | None] = mapped_column(
ForeignKey("scenarios.id", ondelete="CASCADE"), nullable=True
)
adventure_id: Mapped[int | None] = mapped_column(
ForeignKey("adventures.id", ondelete="CASCADE"), nullable=True
)
type: Mapped[str] = mapped_column(String(100), default="")
name: Mapped[str] = mapped_column(String(200), default="")
keys: Mapped[str] = mapped_column(Text, default="") # comma-separated triggers
entry: Mapped[str] = mapped_column(Text, default="")
notes: Mapped[str] = mapped_column(Text, default="")
# Set on adventure copies only. It records which piece of the scenario the
# card came from, as either "card:<scenario_card_id>" or "npc:<npc_key>".
# The "Update from scenario" action refreshes or removes exactly these
# cards. A NULL value means the player wrote the card, so the update leaves
# it alone, or that the copy predates this column, in which case the update
# matches it by name and then sets this value.
source_ref: Mapped[str | None] = mapped_column(String(64), nullable=True)
scenario: Mapped[Scenario | None] = relationship(back_populates="story_cards")
adventure: Mapped[Adventure | None] = relationship(back_populates="story_cards")
def _change_label(parts: list[str]) -> str:
"""Names a world-state path for the inline turn summary.
`npc.gwen.trust` becomes "gwen trust". Every other shape uses its last
segment, so `player.hp` becomes "hp".
"""
if len(parts) == 3 and parts[0] == "npc":
return f"{parts[1]} {parts[2]}"
return parts[-1] if parts else ""
class Action(Base):
__tablename__ = "actions"
# Phase 14: every story read selects one branch up to one depth, then
# another branch up to another depth, and so on. The pair (branch_id, depth)
# is the index those clauses need.
__table_args__ = (Index("ix_actions_branch_depth", "branch_id", "depth"),)
id: Mapped[int] = mapped_column(primary_key=True)
adventure_id: Mapped[int] = mapped_column(ForeignKey("adventures.id", ondelete="CASCADE"))
# Phase 14: the node's place in the tree. `depth` is a position along one
# path rather than a global turn number. Node A4 and node B4 are
# alternatives, not duplicates.
#
# Both columns are nullable because ALTER TABLE cannot add a NOT NULL column
# without a default, and no default makes sense for a branch. The migration
# fills these columns for existing rows, and `tree.place_action` fills them
# for new rows. From SP2 onward, a NULL `branch_id` marks a row that no read
# can see.
branch_id: Mapped[int | None] = mapped_column(
ForeignKey("branches.id", ondelete="CASCADE"), nullable=True
)
depth: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Phase 14, SP9: the node that this node was played after. It names the take
# that was live when this row was written, not whatever sits at depth - 1
# now.
#
# This column answers one question: which takes belong to the same turn. A
# coordinate cannot answer it. A take that is forked onto its own branch
# leaves the coordinate that its siblings still occupy, so the pager would
# show it as 1/1 next to their 1/3. Forking a branch does not change a
# node's parent.
#
# Code reads this column only to group takes, using one indexed lookup
# rather than a walk. Paths still resolve through `lineage`, which is why
# adding this column required no change to any read of the story.
#
# The value is NULL on a root node, and on pre-SP9 rows that the migration
# could not place. For those rows, `attempts.group` falls back to the
# coordinate, which is how they were written.
parent_id: Mapped[int | None] = mapped_column(
ForeignKey("actions.id", ondelete="SET NULL"), nullable=True, index=True
)
# Phase 14, SP4: whether this node is the one the story uses at its
# coordinate. Retry no longer rewrites a row. It writes a sibling at the
# same branch and depth, so one coordinate can hold several attempts while
# exactly one of them is on the path. `lineage.Path.clause` is the only
# place that reads this column, for the same reason it is the only place
# that knows about branches. If a discarded attempt reaches a read, the page
# renders the same turn twice.
#
# A node with no siblings is live, so the default is True and every pre-SP4
# row is already correct. The migration does not need to visit them.
live: Mapped[bool] = mapped_column(Boolean, default=True, nullable=False)
type: Mapped[str] = mapped_column(String(20)) # start|do|say|story|continue|ai
text: Mapped[str] = mapped_column(Text, default="")
# Reasoning-model "thinking" that preceded the text (AI actions only).
reasoning: Mapped[str | None] = mapped_column(Text, nullable=True)
# The full assembled prompt for this turn, used by the Insights viewer.
# This is the largest column in the database. It averages 163 KB per row in
# production and 232 KB on the longest adventure, and it accounts for 89% of
# everything stored. Only one endpoint reads it, one action at a time.
#
# The column is expensive in two ways, so it has two protections.
# `deferred=True` protects reads, because SQLAlchemy loads the column only
# when code touches the attribute. A page load therefore costs nothing. Code
# that reads actions in bulk must not touch this attribute, which is why
# `world_delta` below exists. `CompressedJSON` protects storage, because
# this column determines when the free tier's 512 MB limit is reached. The
# attribute behaves like a plain dict in both cases. See compression.py.
context_snapshot: Mapped[dict | None] = mapped_column(
CompressedJSON, nullable=True, deferred=True
)
# The small slice of the snapshot that IS needed in bulk: this turn's RPG
# state changes, for the inline chips under an AI message (world_changes)
# and for re-attaching the emit block when replaying history to the model.
# Mirrors the active variant, same as text/reasoning/context_snapshot.
world_delta: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# Phase 14, SP4: the RPG world state as it stood after this node was
# played. These columns record the node's outcome rather than its starting
# position. (`state_after` held the scripting engine's state, which M2
# removed; it is now always written empty.)
#
# Two operations need this outcome, and neither can use a snapshot taken
# before the turn. Switching between siblings must restore the state that
# the chosen attempt produced, and the attempts differ in exactly that.
# Rolling back to before a turn means restoring the state that the preceding
# node left behind, which is one lookup along the path.
#
# The value is NULL on pre-SP4 rows for which the migration could not derive
# one. Every caller tolerates that. A missing snapshot means that the caller
# leaves the live state unchanged. It never means reset the state.
#
# These columns are deferred, because code reads them only for the single
# node being switched to, undone, or retried past.
state_after: Mapped[dict | None] = mapped_column(
JSON, nullable=True, deferred=True
)
world_state_after: Mapped[dict | None] = mapped_column(
JSON, nullable=True, deferred=True
)
# M5: the authoritative narrative state as it stood after this node played.
# The genre-neutral successor to `world_state_after`, and the reason Undo,
# Redo and Save Point restore stay bounded: a position's state is one row
# read, not a replay of every event since the campaign began
# (`TECHNICAL-DESIGN.md` §10.4, and the M4 note that made it load-bearing
# for Save Points too).
#
# `DATA-MODEL.md` §17 selects the hybrid — validated events for audit, a
# snapshot for reads and restore. `state_events` is the audit half; this
# column is the restore half, and nothing reconstructs a document from
# events.
#
# Deferred and compressed for the reasons `context_snapshot` is: only the
# single node being moved to reads it, and a document carrying a campaign's
# entities and facts is larger than the RPG dict it replaces. `world_delta`
# has an M5 counterpart in `state_changes` for the bulk read.
narrative_state_after: Mapped[dict | None] = mapped_column(
CompressedJSON, nullable=True, deferred=True
)
# The small slice needed in bulk: the events accepted here, the ones
# refused, and short lines for the chip under an AI message. Same role
# `world_delta` played, and a separate column for the same reason — the
# context builder reads it for every action in the replayed history, and
# the snapshot beside it is deferred so a turn never loads the prompt
# archive. Shape: {"accepted": [...], "rejected": [...], "summary": [...]}.
state_changes: Mapped[dict | None] = mapped_column(JSON, nullable=True)
created_at: Mapped[datetime] = mapped_column(DateTime, default=utcnow)
adventure: Mapped[Adventure] = relationship(back_populates="actions")
@property
def state_events_replay(self) -> list[dict]:
"""M5: the accepted events this turn produced, for replay into the prompt.
Read from `state_changes`' companion slice in the bulk-read column
rather than from the deferred snapshot, because the context builder
calls this for every action in the replayed history and loading the
prompt archive per action is the egress mistake this project keeps a
regression test about.
"""
changes = self.state_changes
if not isinstance(changes, dict):
return []
events = changes.get("accepted")
return events if isinstance(events, list) else []
@property
def state_rejections(self) -> list[dict]:
"""M5: what this turn proposed that the application refused.
Fed back to the model as a correction for one turn only. A refusal it
has already had a chance to fix is stale, and repeating it forever would
price one bad turn into the rest of the campaign.
"""
changes = self.state_changes
if not isinstance(changes, dict):
return []
rejected = changes.get("rejected")
return rejected if isinstance(rejected, list) else []
@property
def state_summary(self) -> list[str]:
"""M5: the short lines shown under an AI message: what changed here."""
changes = self.state_changes
if not isinstance(changes, dict):
return []
lines = changes.get("summary")
return [str(line) for line in lines] if isinstance(lines, list) else []
@property
def world_changes(self) -> list[dict]:
"""Compact per-turn RPG state changes (Phase 12), for the inline summary
under an AI message. Labels are path-based (no schema needed):
`npc.gwen.trust` -> "gwen trust".
The summary reports refused changes as well as accepted ones. A stat the
engine clamped carries `clamped`, and a stat it refused outright becomes
a `rejected` entry carrying the reason. Reporting only the accepted
changes made a clamp indistinguishable from a change that never
happened: a value the model pushed past its ceiling came back as a
delta of 0 and rendered as an ordinary chip, so a refused update read on
screen as an applied one.
Reads `world_delta`, never `context_snapshot`. This runs for every
action in a list response, and touching the deferred snapshot here would
drag the entire prompt archive out of the database."""
wd = self.world_delta if isinstance(self.world_delta, dict) else None
if wd is None:
return []
clamped_paths = {
str(e.get("path", "")) for e in (wd.get("clamped") or []) if isinstance(e, dict)
}
out: list[dict] = []
for entry in wd.get("applied") or []:
path = str(entry.get("path", ""))
parts = path.split(".")
section, name = parts[0], parts[-1]
if section == "flags":
out.append({"kind": "flag", "label": name, "on": bool(entry.get("new"))})
elif section == "milestones":
out.append({"kind": "milestone", "label": name})
else:
old, new = entry.get("old"), entry.get("new")
delta = new - old if isinstance(old, (int, float)) and isinstance(new, (int, float)) else None
chip = {
"kind": "stat",
"label": _change_label(parts),
"delta": delta,
"value": new,
"clamped": path in clamped_paths,
}
# Carried only when the engine wrote one. It is empty for every
# accepted change, and a key per chip per action is paid on
# every page load.
if entry.get("fix"):
chip["fix"] = str(entry["fix"])
out.append(chip)
for entry in wd.get("rejected") or []:
if not isinstance(entry, dict):
continue
parts = str(entry.get("path", "")).split(".")
chip = {
"kind": "rejected",
"label": _change_label(parts),
"reason": str(entry.get("reason", "")),
}
if entry.get("fix"):
chip["fix"] = str(entry["fix"])
out.append(chip)
return out
class Settings(Base):
__tablename__ = "settings"
id: Mapped[int] = mapped_column(primary_key=True)
# Phase 8: one row per user (pre-Phase-8 DBs had a single id=1 row, which
# the migration assigns to the local user).
user_id: Mapped[int | None] = mapped_column(
ForeignKey("users.id", ondelete="CASCADE"), nullable=True, unique=True
)
endpoint_url: Mapped[str] = mapped_column(String(500), default="http://localhost:11434/v1")
# Was a cloud provider's API key, encrypted at rest. Ollama does not use
# one and M2 removed cloud providers, so nothing reads or writes this now.
# The column stays so existing databases open unchanged; an old value is
# left where it is rather than migrated or decrypted.
api_key: Mapped[str] = mapped_column(String(500), default="")
model: Mapped[str] = mapped_column(String(200), default="")
api_mode: Mapped[str] = mapped_column(String(20), default="chat") # chat|completion
temperature: Mapped[float] = mapped_column(Float, default=0.8)
# 800 leaves room for a full scene; 400 tended to truncate mid-paragraph
# and left reasoning models with nothing after their thinking.
max_output_tokens: Mapped[int] = mapped_column(Integer, default=800)
# Was an OpenRouter-style thinking budget. Ollama's OpenAI-compatible
# endpoint ignores the field, so M2 stopped sending it and removed it from
# the Settings API and UI. The column stays so existing databases open
# unchanged and is never read.
reasoning_max_tokens: Mapped[int] = mapped_column(Integer, default=0)
context_token_budget: Mapped[int] = mapped_column(Integer, default=16384)
# How long to wait for the model, in seconds, before giving up on a turn.
# A cold load of a mid-sized model on a CPU-only machine can take minutes,
# while the same turn takes seconds once the model is resident. See
# `providers.openai_compatible.DEFAULT_READ_TIMEOUT`.
model_timeout_seconds: Mapped[int] = mapped_column(Integer, default=300)
narrator_prompt: Mapped[str] = mapped_column(
Text,
default=(
"You are a masterful storyteller continuing an interactive adventure. "
"Continue the story naturally in second person, staying consistent with "
"everything established so far. Write vivid prose. Never speak for the "
"player or break character. Do not conclude the story; always leave room "
"for the player's next action."
),
)
# Phase 6: auto-summarization + memory bank
summary_model: Mapped[str] = mapped_column(String(200), default="") # "" = main model
embedding_model: Mapped[str] = mapped_column(String(200), default="") # "" = bank disabled
# This was 200. It was lowered mainly to improve retrieval quality. Ranking
# 200 memories to choose 5 selects from a large amount of noise, and the
# oldest memories describe a part of the story that the player has left
# behind. Cheaper reads are a secondary benefit rather than the reason.
memory_bank_capacity: Mapped[int] = mapped_column(Integer, default=80)
memory_top_k: Mapped[int] = mapped_column(Integer, default=5)
# The hosted visitor dashboard's two counter tables and the access log that
# recorded sign-ins, addresses and devices used to be mapped here. M2 removed
# the hosted deployment they served.
#
# The tables are left in the database rather than dropped: they are inert,
# nothing reads or writes them, and a destructive migration would risk an
# existing campaign database for tidiness alone. They are not product
# functionality.
@event.listens_for(Session, "before_flush")
def _place_new_nodes_on_the_tree(session, flush_context, instances):
from . import tree
tree.place_new_nodes(session)