Files
interactive-story/backend/app/schemas.py
JesseMarkowitzandClaude Opus 5 480414efe0 M7: a first-class imported knowledge library
A campaign can import local .txt and .md files as Canon, Reference or
Inspiration, and the class is load-bearing rather than a label: it decides the
words a passage is framed with in the prompt, the weight it carries when
passages are ranked, and which budget it competes in when the context is tight.

This is a separate subsystem, which is the Phase 0B decision
(IMPORTED-KNOWLEDGE-DESIGN.md §73). Story Cards do not carry classification,
provenance, content identity, chunking, an index or a lifecycle, and they were
not promoted into something that does. Nothing here reads or writes one.

The subsystem, in backend/app/knowledge/:

  classes      the three classes, their weights, and the prompt framing
  chunking     deterministic, heading-aware, 60-800 tokens, no overlap
  fts          SQLite FTS5 with porter stemming; scoped and bounded in SQL
  importer     validate, hash, store, chunk, index — in one transaction
  embeddings   local Ollama vectors through the shared provider
  retrieval    query construction, hybrid merge, rerank
  inject       the budgeted cut and the rendered prompt sections

Relevance admission is a separate stage from ranking, and that separation is
the milestone's most expensive lesson. An independent review found the first
implementation deciding relevance with a floor expressed as a share of the best
candidate — which the best clears by construction — so a passage was admitted on
every turn regardless of the scene. A query about tide tables and container
tonnage retrieved all five sources of a fantasy campaign, narrator-only hidden
Canon among them.

So the pipeline is now:

  candidate generation -> admission -> ranking -> class weighting -> budget

Admission reads raw, candidate-set-independent signals: the cosine the model
returned, and how many distinct meaningful query terms a passage contains.
Ranking reads normalized ones, because bm25 has no fixed range and cosine's zero
is not zero. Normalization decides order among things that matched; it can never
decide whether anything matched. Authority is applied after admission, so a
class orders what matched and never rescues what did not.

Retrieval may therefore return nothing, and on a scene unrelated to the library
it does.

The other decisions that each replaced an obvious wrong one:

- The class multiplies relevance rather than adding to it. An additive bonus
  satisfies "Canon outranks Reference" and makes "do not include irrelevant
  Canon" impossible, because a large enough constant wins on its own.
- The semantic floor is measured, not guessed: 113 production-path pairs against
  nomic-embed-text put targeted matches at 0.55-0.85 and off-topic pairs at
  0.36-0.56, and 0.58 sits between them. Because it is a property of that model
  and not of cosine similarity, it is keyed to the model rather than applied to
  whatever is configured: an embedding model with no measured calibration in
  this build does not borrow the number. Semantic admission is skipped, the
  campaign retrieves lexically, and the reason is stated in the knowledge status
  and in the turn's provenance. Degrading to lexical keeps the library usable;
  lending the threshold to an unmeasured model is how the admitted-everything
  defect would return.
- One lexical term is not evidence. Two distinct meaningful terms, or one that
  is neither a standing campaign entity nor a negligible share of the query.
  The stop list grew from 42 words to 261, all function words — no subject
  matter, because a stop list that removes subject matter stops finding "The
  Silver Key".
- Lexical retrieval is a production path, not a fallback. It finds the proper
  nouns and invented terms a setting bible is made of, and the library is fully
  usable with no embedding model configured.

Safety is structural rather than filtered. Imported text reaches the prompt
whole, inside a section that says what it is, under a rule stating the authority
order in words and refusing every instruction inside it. No endpoint accepts a
filesystem path, so H08 has no mechanism to escape from. Nothing renders
imported content as HTML, so a script tag is five visible characters and a
remote image is never fetched. Import, chunking, indexing, retrieval and a turn
open no socket at all; only embeddings do, through the endpoint allowlist the
memory bank already uses.

Provenance is the rendered text, not a foreign key: deleting a source cannot
turn a historical turn's evidence into dangling ids.

Schema: knowledge_sources, knowledge_chunks, knowledge_embeddings, and an FTS5
virtual table attached to knowledge_chunks as a DDL hook so it is created and
dropped with the table it indexes. Migration 92. A pre-M7 database opens
unchanged and needs no sources to play.

Bundle: the source content and the reader's judgements about it travel; the
passages, index rows and vectors are rebuilt on import, so a restored campaign
is searchable immediately without a reindex step.

One runtime dependency: python-multipart, Starlette's multipart parser. It is
what makes the upload surface possible, and the upload surface is why no
pathname is ever accepted.

The test doubles were the reason the defect shipped, so they were corrected too.
The retrieval stub scored unrelated text at 0.06-0.20 where the real model
scores it at 0.43-0.44, and its docstring said it had deliberately removed the
constant component that "would put a similarity floor under every pair" — which
is exactly the property real models have. The stub now has that floor, one test
fails if it is ever removed, and another reproduces the superseded rule and
asserts it is still fooled by the same fixture. Run against the pre-corrective
implementation, the new suite fails 13 of 18.

Tests: 939 passed, 14 skipped (836/7 at M6). 110 new across seven files, one of
which mocks nothing between itself and Ollama and re-measures the similarity
separation on every run. 43/43 checks in a real Firefox, reproduced.
Docker build clean.

Four other defects found by review or by the browser run were fixed here rather
than carried: an unreachable relevance constant that appeared to enforce
something and did not; acceptance tests using the wrong fixture files, so G07's
trap was never exercised; a bidirectional override surviving into displayed
filenames; and, from the implementation pass, the Insights panel showing M5's
two state sections as raw keys and the source inspector refetching on every
keystroke.

M7 was independently reviewed, which returned PASS WITH CORRECTIVE WORK
REQUIRED. Both blocking findings are closed, and closeout resolved the
embedding-model calibration boundary the corrective pass had left as debt.
planning/reports/M7-IMPLEMENTATION-REPORT.md carries the review, the corrective
closeout and the closeout verification in sequence, none overwriting another.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HdaXiFbscatQaLS7dJk6b
2026-09-06 15:40:13 -04:00

673 lines
22 KiB
Python

from datetime import datetime
from typing import Annotated, Literal
from pydantic import BaseModel, ConfigDict, Field, computed_field
from . import images
# Length caps (Phase 9). The VARCHAR caps are a correctness requirement rather
# than only an abuse limit. Postgres enforces column lengths and SQLite never
# did, so a longer value has to be a 422 here rather than a 500 at INSERT. The
# text-column caps are generous abuse ceilings that a legitimate player does not
# reach.
NAME_MAX = 200 # Titles and names. VARCHAR(200).
TAGS_MAX = 500 # VARCHAR(500).
CARD_TYPE_MAX = 100 # VARCHAR(100).
PROSE_MAX = 50_000 # Memory, author's note, prompts, entries, and notes.
ACTION_MAX = 20_000 # One player action.
MEMORY_TEXT_MAX = 5_000
# A scenario cover image, stored inline as a base64 data URI. A 400x300 WebP at
# the quality the editor encodes runs about 20 to 40 kB. A cap of 400 kB leaves
# room for a client that downscales less aggressively, and it stops anyone from
# storing a multi-megabyte PNG in a row that every list request reads.
IMAGE_MAX = 400_000
ICON_MAX = 16 # One emoji or glyph. VARCHAR(16).
BRANCH_NAME_MAX = 80 # What a player called one line of the story. VARCHAR(80).
PERSONA_NAME_MAX = 80 # The protagonist's name. VARCHAR(80).
PERSONA_PRONOUNS_MAX = 40 # "they/them" and the like. VARCHAR(40).
# M4: what a player called a Save Point. VARCHAR(120). Wider than a branch name
# because these are sentences rather than labels — "Before entering the abbey"
# is the example the specification uses throughout.
CHECKPOINT_NAME_MAX = 120
Name = Annotated[str, Field(max_length=NAME_MAX)]
Tags = Annotated[str, Field(max_length=TAGS_MAX)]
CardType = Annotated[str, Field(max_length=CARD_TYPE_MAX)]
Prose = Annotated[str, Field(max_length=PROSE_MAX)]
ActionText = Annotated[str, Field(max_length=ACTION_MAX)]
Image = Annotated[str, Field(max_length=IMAGE_MAX)]
Icon = Annotated[str, Field(max_length=ICON_MAX)]
PersonaName = Annotated[str, Field(max_length=PERSONA_NAME_MAX)]
PersonaPronouns = Annotated[str, Field(max_length=PERSONA_PRONOUNS_MAX)]
CheckpointName = Annotated[str, Field(max_length=CHECKPOINT_NAME_MAX)]
class ORMModel(BaseModel):
model_config = ConfigDict(from_attributes=True)
# ---------- Story cards ----------
class StoryCardBase(BaseModel):
type: CardType = ""
name: Name = ""
keys: Prose = ""
entry: Prose = ""
notes: Prose = ""
class StoryCardCreate(StoryCardBase):
scenario_id: int | None = None
adventure_id: int | None = None
class StoryCardUpdate(BaseModel):
type: CardType | None = None
name: Name | None = None
keys: Prose | None = None
entry: Prose | None = None
notes: Prose | None = None
class StoryCardOut(ORMModel, StoryCardBase):
id: int
scenario_id: int | None
adventure_id: int | None
# ---------- Scenarios ----------
class ScenarioBase(BaseModel):
title: Name = "Untitled Scenario"
description: Prose = ""
prompt: Prose = ""
memory: Prose = ""
authors_note: Prose = ""
ai_instructions: Prose = ""
tags: Tags = ""
# Cover art, either an https URL or a base64 data URI. See `app/images.py`.
image: Image = ""
# The emoji or glyph shown when `image` is empty.
icon: Icon = ""
# Phase 12: the RPG world-state template, holding stat definitions, bands,
# rules, and milestones. `None` means the scenario has no RPG layer.
stat_schema: dict | None = None
class ScenarioCreate(ScenarioBase):
pass
class ScenarioUpdate(BaseModel):
title: Name | None = None
description: Prose | None = None
prompt: Prose | None = None
memory: Prose | None = None
authors_note: Prose | None = None
ai_instructions: Prose | None = None
tags: Tags | None = None
image: Image | None = None
icon: Icon | None = None
stat_schema: dict | None = None
class ScenarioOut(ORMModel, ScenarioBase):
id: int
is_public: bool = False # Shared demo content, read-only for everyone.
created_at: datetime
updated_at: datetime
story_cards: list[StoryCardOut] = []
class ScenarioListItem(ORMModel):
id: int
title: str
description: str
tags: str
is_public: bool = False
updated_at: datetime
# Read from the row so that `image_url` can be derived, and excluded from
# the response, because a list of base64 data URIs would be megabytes of
# JSON.
image: str = Field("", exclude=True)
icon: str = ""
@computed_field
@property
def image_url(self) -> str:
return images.public_url(self.id, self.image, self.updated_at)
# ---------- Adventures ----------
class AdventureCreate(BaseModel):
scenario_id: int | None = None
title: Name | None = None
# The `${Placeholder}` values collected from the player at the start, which
# is the AI Dungeon behavior.
placeholders: dict[str, str] = {}
# Phase 18: who the player is playing as, collected by the same modal. These
# are independent of `placeholders`: a scenario that asks for `${Name}` is
# asking its own question, and nothing here fills it in.
persona_name: PersonaName = ""
persona_pronouns: PersonaPronouns = ""
persona_desc: Prose = ""
class AdventureUpdate(BaseModel):
title: Name | None = None
memory: Prose | None = None
authors_note: Prose | None = None
ai_instructions: Prose | None = None
story_summary: Prose | None = None
auto_summarize: bool | None = None
memory_bank_enabled: bool | None = None
persona_name: PersonaName | None = None
persona_pronouns: PersonaPronouns | None = None
persona_desc: Prose | None = None
class AdventureRefresh(BaseModel):
"""The body for "Update from scenario".
`placeholders` supplies answers the adventure has no stored value for. See
`AdventureCreate.placeholders`. The answers are merged over the stored ones
and saved.
"""
placeholders: dict[str, str] = {}
class RefreshPlan(BaseModel):
"""What a refresh would change. The confirm dialog is built from this."""
scenario_id: int
scenario_title: str
has_changes: bool
# Maps a field name to `{"old": ..., "new": ...}`, for differing fields
# only.
fields: dict[str, dict] = {}
# Maps "added", "updated", or "removed" to a list of card names.
cards: dict[str, list[str]] = {}
# Maps "added" or "removed" to a list of stat paths. Live values are
# otherwise kept.
world_state: dict[str, list[str]] = {}
# The `${Placeholder}` names the scenario asks for that the adventure has
# no stored answer to. The client collects these and sends them back.
placeholders_needed: list[str] = []
class ActionOut(ORMModel):
id: int
adventure_id: int
type: str
text: str
reasoning: str | None = None
# Phase 12: the compact RPG state changes for this turn, read from the
# model property. Legacy as of M5 and empty on new turns; kept so a pre-M5
# campaign's chips still render.
world_changes: list[dict] = []
# M5: what this turn changed, as short lines for the chip under an AI
# message. Read from `Action.state_summary`, which reads the small
# bulk-loaded column rather than the deferred snapshot.
state_summary: list[str] = []
# SP9: the pager, such as `2/4`. It reports how many attempts this turn has
# and which one is on screen. It is keyed on the parent, so it counts the
# attempts of this turn rather than every node that shares a depth, and it
# keeps counting them after one has been forked onto its own branch.
#
# A turn nobody has retaken reads 1/1, which is most turns, and the client
# draws no pager for a count of one. The attempts themselves come from
# `GET /actions/{id}/variants`, so this payload stays small.
take_count: int = 1
take_index: int = 0
# Which line this node is on, so the pager can distinguish the two kinds of
# step without asking the server. An attempt on this branch is a leaf with
# nothing below it, so showing it is a local change. An attempt on another
# branch has a story of its own, so moving to it is a branch switch.
branch_id: int | None = None
created_at: datetime
class VariantOut(BaseModel):
# Since SP4 every attempt is its own node, so each one has an id, and the
# client needs that id. A fork is addressed by the attempt being promoted,
# not by its position in a group that renumbers whenever an attempt is
# added.
id: int
index: int
text: str
reasoning: str | None = None
# See `ActionOut.branch_id`. It decides whether choosing this attempt is a
# local step or a branch switch.
branch_id: int | None = None
created_at: str | None = None
active: bool = False
class VariantSelect(BaseModel):
index: int = Field(ge=0)
class BranchOut(ORMModel):
"""One line through the story tree (Phase 14, SP5).
This carries enough to draw the tree and nothing more. `fork_depth` is where
this line leaves its parent, and `depth` is where it currently ends, so a
fork is two numbers rather than a walk. `own_actions` counts the turns played
on this branch itself. The rest of its story is borrowed from its ancestors,
which is why the number is smaller than a reader expects.
"""
id: int
parent_branch_id: int | None = None
fork_depth: int | None = None
depth: int
own_actions: int = 0
# M4: how many Save Points name a position on this line. Deleting the branch
# deletes them with its story, so the panel warns with a number rather than
# a vague caution. Zero for a line nobody has bookmarked, which is most.
save_points: int = 0
is_head: bool = False
# NULL for a branch nobody has named. The client labels those from the fork
# depth rather than the server inventing a name. See the column comment.
name: str | None = None
created_at: datetime
class BranchRename(BaseModel):
"""A name a player chose, or `null` to make the branch unnamed again."""
name: Annotated[str, Field(max_length=BRANCH_NAME_MAX)] | None = None
# ---------- Narrative state (M5) ----------
class StateGroup(BaseModel):
"""One labelled section of the state inspector.
Rows carry the key as well as the label, because a manual correction has to
name an entity and the user should not have to guess the identifier.
"""
title: str
rows: list[dict] = []
class NarrativeStateOut(BaseModel):
"""The authoritative state at the active head.
`groups` is the display form and `document` is the state itself. Both are
returned because they answer different questions: the panel renders the
first, and a correction form — or a test — needs the second to name a key.
"""
groups: list[StateGroup] = []
empty: bool = True
document: dict = {}
class StateEventIn(BaseModel):
"""One typed event, as a client proposes it.
Deliberately loose about which fields are present: the event vocabulary is
defined in `narrative/events.py` and enforced by `narrative/validate.py`,
and duplicating those rules here would create a second, drifting copy of the
allowlist. What this model does is bound the shapes — a type that is a
string, values that are scalars, labels that are short strings — so a
payload cannot smuggle a structure past Pydantic and reach the validator as
something other than an event.
"""
model_config = ConfigDict(extra="allow")
type: Annotated[str, Field(max_length=60)]
class StateCorrection(BaseModel):
"""A manual correction: the user overruling what the story established.
`note` records why, in the user's words, and is kept on the proposal record
so the audit says more than "the user changed this".
"""
events: Annotated[list[StateEventIn], Field(min_length=1, max_length=20)]
note: Prose = ""
class StateEventOut(ORMModel):
"""One accepted change, for the audit view."""
id: int
action_id: int | None = None
branch_id: int | None = None
depth: int | None = None
# The reader-facing position, matching the Save Point panel's vocabulary.
turn: int | None = None
sequence: int = 0
event_type: str
payload: dict = {}
before: dict | None = None
source: str = "accepted_story"
created_at: datetime
# ---------- Save Points (M4) ----------
#
# "Save Point" is the user-facing term and `checkpoint` is the internal one
# (`BROWSER-UX-SPEC.md` §23). The wire format uses the internal name, as the
# rest of this module does.
class CheckpointOut(ORMModel):
"""One Save Point: a name and the position it names.
The position is reported three ways because the panel needs three different
things from it. `turn` is what a reader counts — the same `depth + 1` the
branch list shows. `depth` and `branch_id` are the coordinate itself.
`on_path` says whether the position lies on the story being read, which is
how the panel can tell a Save Point on this line from one naming a line the
story has left; restoring either works, but they are not the same offer.
`resolved` is false when the coordinate no longer names a live turn, which
an action deleted out of the middle of a story can do. Restore refuses such
a Save Point rather than moving the head somewhere approximate, so the list
says so before the button is pressed.
"""
id: int
adventure_id: int
name: str
note: str = ""
branch_id: int
depth: int
turn: int = 0
on_path: bool = True
resolved: bool = True
created_at: datetime
updated_at: datetime
class CheckpointCreate(BaseModel):
"""A Save Point at wherever the story is being read.
The position is not a field. A Save Point is made at the campaign's active
head, which the server already knows, and accepting a coordinate from the
client would be the second way to name a position — the thing this milestone
exists not to build.
"""
name: CheckpointName
note: Prose = ""
class CheckpointRename(BaseModel):
"""A new label, and nothing else.
There is deliberately no coordinate here. `STORY-BRANCH-SEMANTICS.md` §24
keeps a Save Point's meaning auditable by refusing to move one: rename it,
or delete it and make another where you are.
"""
name: CheckpointName | None = None
note: Prose | None = None
class ActionUpdate(BaseModel):
text: ActionText
class ActionCreate(BaseModel):
type: Literal["do", "say", "story", "continue"]
text: ActionText = ""
# The node this action is played after (SP9). Omitting it means the tip,
# which is what every ordinary turn uses.
#
# Naming an attempt the story moved past is what creates a branch. Stepping
# between attempts costs nothing and creates nothing, and the fork happens
# on the first text written below one. That is the first moment the player
# states which line they mean. Before it, they were reading.
after_id: int | None = None
class TakeCreate(BaseModel):
"""Another attempt at a turn (SP9).
`text` is what the player says instead, and it applies only when the turn was
the player's. An AI turn's other attempt is generated, so the field is
ignored there rather than rejected. The client makes the same request for
both, and the node type decides what happens.
"""
text: ActionText = ""
class AdventureOut(ORMModel):
id: int
scenario_id: int | None
title: str
memory: str
authors_note: str
ai_instructions: str
story_summary: str
auto_summarize: bool
memory_bank_enabled: bool
persona_name: str
persona_pronouns: str
persona_desc: str
created_at: datetime
updated_at: datetime
story_cards: list[StoryCardOut] = []
# The newest window of the story, not all of it. Older pages arrive from
# `GET /{id}/actions` as the reader scrolls up. `action_count` is the whole
# story's length, which is how the client knows more actions exist above.
actions: list[ActionOut] = []
action_count: int = 0
# M3. Whether the history controls have anywhere to go from where the story
# is. The client cannot work either out for itself: `can_undo` needs the
# campaign opening, which may be off the top of the loaded window, and
# `can_redo` needs the retained future, which the client is never sent.
can_undo: bool = False
can_redo: bool = False
class ActionPage(BaseModel):
"""A slice of the story, counted back from the newest action."""
actions: list[ActionOut] = []
total: int = 0
# Whether anything older than this slice exists. The server computes it, so
# the client never has to do arithmetic on positions to find the end.
has_more: bool = False
# The same two flags `AdventureOut` carries, so that the response to Undo,
# Redo or a turn updates the controls without a second request.
can_undo: bool = False
can_redo: bool = False
# ---------- Memory bank (Phase 6) ----------
class MemoryOut(ORMModel):
id: int
adventure_id: int
text: str
pinned: bool
forgotten: bool
embedded: bool # model property: embedding vector present
use_count: int
last_used_at: datetime | None
source_start: int | None
source_end: int | None
created_at: datetime
class MemoryCreate(BaseModel):
text: Annotated[str, Field(max_length=MEMORY_TEXT_MAX)]
class MemoryUpdate(BaseModel):
text: Annotated[str, Field(max_length=MEMORY_TEXT_MAX)] | None = None
pinned: bool | None = None
forgotten: bool | None = None
# ---------------------------------------------------------------- M7: knowledge
class KnowledgeSourceOut(BaseModel):
"""One imported source, as a list row.
Deliberately without `content`. A library of twenty files would otherwise
put every byte of every one of them on a screen that shows none of it;
`KnowledgeSourceDetail` is what serves the text when it is asked for.
"""
id: int
title: str
original_filename: str
classification: str
enabled: bool
visibility: str
always_include: bool
content_hash: str
byte_size: int
media_type: str
chunk_count: int
embedded_count: int
# The two halves of derived state, kept apart on purpose. Lexical retrieval
# is a supported production path, so "the vectors failed" and "the index
# failed" are different sentences with different consequences.
index_state: str
index_detail: str
embed_state: str
embed_detail: str
parser_version: int
chunking_version: int
imported_at: str | None = None
updated_at: str | None = None
class KnowledgeSourceDetail(KnowledgeSourceOut):
"""A source with its text, for the inspector.
`content` is the file as it was decoded, not the normalized form used for
hashing and search: the reader inspects what they imported
(`IMPORTED-KNOWLEDGE-DESIGN.md` §61).
"""
content: str
notes: str = ""
class KnowledgeChunkOut(BaseModel):
id: int
chunk_index: int
heading_path: str
text: str
token_count: int
content_hash: str
embedded: bool
embedding_model: str = ""
class KnowledgeSourceUpdate(BaseModel):
"""What a reader may change about a source without reimporting it.
Everything here is metadata or state. Nothing rewrites content, and nothing
is destructive: changing a classification re-frames and re-weights the same
passages, and disabling a source removes it from retrieval while leaving the
rows exactly where they are.
"""
title: str | None = None
classification: str | None = None
enabled: bool | None = None
visibility: str | None = None
always_include: bool | None = None
notes: str | None = None
class AdventureListItem(ORMModel):
id: int
scenario_id: int | None
scenario_title: str | None = None
title: str
updated_at: datetime
action_count: int = 0
# The end of the most recent narration, so a Continue card can show the
# story rather than only a turn count.
snippet: str = ""
# Cover art inherited from the parent scenario. See `app/images.py`.
image_url: str = ""
icon: str = ""
# ---------- Auth (Phase 8) ----------
class AuthCredentials(BaseModel):
email: Annotated[str, Field(max_length=320)] # VARCHAR(320).
# The upper bound keeps the scrypt cost constant. Without it, hashing a
# megabyte password would give an attacker free CPU time.
password: Annotated[str, Field(max_length=128)]
# ---------- Settings ----------
class SettingsOut(ORMModel):
endpoint_url: str
model: str
api_mode: str
temperature: float
max_output_tokens: int
context_token_budget: int
model_timeout_seconds: int
narrator_prompt: str
summary_model: str
embedding_model: str
memory_bank_capacity: int
memory_top_k: int
ScenarioOut.model_rebuild()
# ---------- AI Chat (power users) ----------
# A scratchpad for talking to a model directly, with no story framing. The
# server persists nothing, so these caps are per-request abuse limits only.
CHAT_MESSAGE_MAX = 100_000 # One message.
CHAT_TOTAL_MAX = 400_000 # The whole conversation sent per request.
CHAT_MESSAGES_MAX = 200 # Turns per request.
class ChatMessage(BaseModel):
role: Literal["system", "user", "assistant"]
content: Annotated[str, Field(max_length=CHAT_MESSAGE_MAX)]
class ChatRequest(BaseModel):
messages: Annotated[list[ChatMessage], Field(min_length=1, max_length=CHAT_MESSAGES_MAX)]
# If this field is empty or omitted, the user's configured model is used.
model: Name | None = None
temperature: Annotated[float, Field(ge=0, le=5)] | None = None
max_tokens: Annotated[int, Field(ge=1, le=100_000)] | None = None
class SettingsUpdate(BaseModel):
endpoint_url: Annotated[str, Field(max_length=500)] | None = None # VARCHAR(500).
model: Name | None = None
api_mode: Annotated[str, Field(max_length=20)] | None = None
temperature: Annotated[float, Field(ge=0, le=5)] | None = None
max_output_tokens: Annotated[int, Field(ge=1, le=100_000)] | None = None
context_token_budget: Annotated[int, Field(ge=256, le=200_000)] | None = None
# Seconds to wait for the model. The floor is high enough that a normal
# turn cannot trip it; the ceiling exists so that "wait longer" stays a
# number rather than becoming "wait forever".
model_timeout_seconds: Annotated[int, Field(ge=30, le=3600)] | None = None
narrator_prompt: Prose | None = None
summary_model: Name | None = None
embedding_model: Name | None = None
memory_bank_capacity: Annotated[int, Field(ge=1, le=1000)] | None = None
memory_top_k: Annotated[int, Field(ge=1, le=50)] | None = None