Files
JesseMarkowitzandClaude Opus 5 ef25b0a876 Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
2026-09-10 06:13:55 -04:00

744 lines
26 KiB
Python

from datetime import datetime
from typing import Annotated, Literal
from pydantic import BaseModel, ConfigDict, Field, computed_field
from . import images
# Length caps (Phase 9). The VARCHAR caps are a correctness requirement rather
# than only an abuse limit. Postgres enforces column lengths and SQLite never
# did, so a longer value has to be a 422 here rather than a 500 at INSERT. The
# text-column caps are generous abuse ceilings that a legitimate player does not
# reach.
NAME_MAX = 200 # Titles and names. VARCHAR(200).
TAGS_MAX = 500 # VARCHAR(500).
CARD_TYPE_MAX = 100 # VARCHAR(100).
PROSE_MAX = 50_000 # Memory, author's note, prompts, entries, and notes.
ACTION_MAX = 20_000 # One player action.
MEMORY_TEXT_MAX = 5_000
# A scenario cover image, stored inline as a base64 data URI. A 400x300 WebP at
# the quality the editor encodes runs about 20 to 40 kB. A cap of 400 kB leaves
# room for a client that downscales less aggressively, and it stops anyone from
# storing a multi-megabyte PNG in a row that every list request reads.
IMAGE_MAX = 400_000
ICON_MAX = 16 # One emoji or glyph. VARCHAR(16).
BRANCH_NAME_MAX = 80 # What a player called one line of the story. VARCHAR(80).
PERSONA_NAME_MAX = 80 # The protagonist's name. VARCHAR(80).
PERSONA_PRONOUNS_MAX = 40 # "they/them" and the like. VARCHAR(40).
# M4: what a player called a Save Point. VARCHAR(120). Wider than a branch name
# because these are sentences rather than labels — "Before entering the abbey"
# is the example the specification uses throughout.
CHECKPOINT_NAME_MAX = 120
Name = Annotated[str, Field(max_length=NAME_MAX)]
Tags = Annotated[str, Field(max_length=TAGS_MAX)]
CardType = Annotated[str, Field(max_length=CARD_TYPE_MAX)]
Prose = Annotated[str, Field(max_length=PROSE_MAX)]
ActionText = Annotated[str, Field(max_length=ACTION_MAX)]
Image = Annotated[str, Field(max_length=IMAGE_MAX)]
Icon = Annotated[str, Field(max_length=ICON_MAX)]
PersonaName = Annotated[str, Field(max_length=PERSONA_NAME_MAX)]
PersonaPronouns = Annotated[str, Field(max_length=PERSONA_PRONOUNS_MAX)]
CheckpointName = Annotated[str, Field(max_length=CHECKPOINT_NAME_MAX)]
class ORMModel(BaseModel):
model_config = ConfigDict(from_attributes=True)
# ---------- Story cards ----------
class StoryCardBase(BaseModel):
type: CardType = ""
name: Name = ""
keys: Prose = ""
entry: Prose = ""
notes: Prose = ""
class StoryCardCreate(StoryCardBase):
scenario_id: int | None = None
adventure_id: int | None = None
class StoryCardUpdate(BaseModel):
type: CardType | None = None
name: Name | None = None
keys: Prose | None = None
entry: Prose | None = None
notes: Prose | None = None
class StoryCardOut(ORMModel, StoryCardBase):
id: int
scenario_id: int | None
adventure_id: int | None
# ---------- Scenarios ----------
class ScenarioBase(BaseModel):
title: Name = "Untitled Scenario"
description: Prose = ""
prompt: Prose = ""
memory: Prose = ""
authors_note: Prose = ""
ai_instructions: Prose = ""
tags: Tags = ""
# Cover art, either an https URL or a base64 data URI. See `app/images.py`.
image: Image = ""
# The emoji or glyph shown when `image` is empty.
icon: Icon = ""
# Phase 12: the RPG world-state template, holding stat definitions, bands,
# rules, and milestones. `None` means the scenario has no RPG layer.
stat_schema: dict | None = None
class ScenarioCreate(ScenarioBase):
pass
class ScenarioUpdate(BaseModel):
title: Name | None = None
description: Prose | None = None
prompt: Prose | None = None
memory: Prose | None = None
authors_note: Prose | None = None
ai_instructions: Prose | None = None
tags: Tags | None = None
image: Image | None = None
icon: Icon | None = None
stat_schema: dict | None = None
class ScenarioOut(ORMModel, ScenarioBase):
id: int
is_public: bool = False # Shared demo content, read-only for everyone.
created_at: datetime
updated_at: datetime
story_cards: list[StoryCardOut] = []
class ScenarioListItem(ORMModel):
id: int
title: str
description: str
tags: str
is_public: bool = False
updated_at: datetime
# Read from the row so that `image_url` can be derived, and excluded from
# the response, because a list of base64 data URIs would be megabytes of
# JSON.
image: str = Field("", exclude=True)
icon: str = ""
@computed_field
@property
def image_url(self) -> str:
return images.public_url(self.id, self.image, self.updated_at)
# ---------- Adventures ----------
class AdventureCreate(BaseModel):
scenario_id: int | None = None
title: Name | None = None
# M8: the opening scene, for a campaign started without a scenario.
#
# A scenario's `prompt` already becomes the campaign's `start` action, and
# this is the same thing said directly. It exists because M8's setup flow
# creates a campaign from a form rather than from a template
# (`BROWSER-UX-SPEC.md` §41), and without it every new campaign opens on a
# blank page — the reader has to invent the situation *and* the first move
# in one box. Ignored when `scenario_id` is given, which already supplies one.
opening: Prose = ""
# M8: the campaign's own rules, as a list of sentences.
#
# The column has existed since migration 82 and both the prompt
# (`context/builder._canon_section`) and the state validator
# (`narrative/apply`) already read it — it simply had no way in from the
# browser, so a fixture had to write it with SQL. This is the highest
# authority in the campaign, which is exactly why a person setting one up
# needs to be able to state it.
canon_rules: list[Name] = []
# The `${Placeholder}` values collected from the player at the start, which
# is the AI Dungeon behavior.
placeholders: dict[str, str] = {}
# Phase 18: who the player is playing as, collected by the same modal. These
# are independent of `placeholders`: a scenario that asks for `${Name}` is
# asking its own question, and nothing here fills it in.
persona_name: PersonaName = ""
persona_pronouns: PersonaPronouns = ""
persona_desc: Prose = ""
# M11: how long the reader wants turns to be. The setup screen also puts a
# sentence about it into `ai_instructions`; this is the half the prompt
# builder can do arithmetic with.
narration_length: Literal["", "brief", "medium", "long"] = ""
class AdventureUpdate(BaseModel):
title: Name | None = None
memory: Prose | None = None
authors_note: Prose | None = None
ai_instructions: Prose | None = None
narration_length: Literal["", "brief", "medium", "long"] | None = None
story_summary: Prose | None = None
auto_summarize: bool | None = None
memory_bank_enabled: bool | None = None
persona_name: PersonaName | None = None
persona_pronouns: PersonaPronouns | None = None
persona_desc: Prose | None = None
# M8. See `AdventureCreate.canon_rules`. Editable after setup because canon
# is the thing a reader most often gets wrong first and needs to correct —
# "resurrection is impossible" is easier to write once the story has tried it.
canon_rules: list[Name] | None = None
class AdventureRefresh(BaseModel):
"""The body for "Update from scenario".
`placeholders` supplies answers the adventure has no stored value for. See
`AdventureCreate.placeholders`. The answers are merged over the stored ones
and saved.
"""
placeholders: dict[str, str] = {}
class RefreshPlan(BaseModel):
"""What a refresh would change. The confirm dialog is built from this."""
scenario_id: int
scenario_title: str
has_changes: bool
# Maps a field name to `{"old": ..., "new": ...}`, for differing fields
# only.
fields: dict[str, dict] = {}
# Maps "added", "updated", or "removed" to a list of card names.
cards: dict[str, list[str]] = {}
# Maps "added" or "removed" to a list of stat paths. Live values are
# otherwise kept.
world_state: dict[str, list[str]] = {}
# The `${Placeholder}` names the scenario asks for that the adventure has
# no stored answer to. The client collects these and sends them back.
placeholders_needed: list[str] = []
class ActionOut(ORMModel):
id: int
adventure_id: int
type: str
text: str
reasoning: str | None = None
# Phase 12: the compact RPG state changes for this turn, read from the
# model property. Legacy as of M5 and empty on new turns; kept so a pre-M5
# campaign's chips still render.
world_changes: list[dict] = []
# M5: what this turn changed, as short lines for the chip under an AI
# message. Read from `Action.state_summary`, which reads the small
# bulk-loaded column rather than the deferred snapshot.
state_summary: list[str] = []
# SP9: the pager, such as `2/4`. It reports how many attempts this turn has
# and which one is on screen. It is keyed on the parent, so it counts the
# attempts of this turn rather than every node that shares a depth, and it
# keeps counting them after one has been forked onto its own branch.
#
# A turn nobody has retaken reads 1/1, which is most turns, and the client
# draws no pager for a count of one. The attempts themselves come from
# `GET /actions/{id}/variants`, so this payload stays small.
take_count: int = 1
take_index: int = 0
# Which line this node is on, so the pager can distinguish the two kinds of
# step without asking the server. An attempt on this branch is a leaf with
# nothing below it, so showing it is a local change. An attempt on another
# branch has a story of its own, so moving to it is a branch switch.
branch_id: int | None = None
created_at: datetime
class VariantOut(BaseModel):
# Since SP4 every attempt is its own node, so each one has an id, and the
# client needs that id. A fork is addressed by the attempt being promoted,
# not by its position in a group that renumbers whenever an attempt is
# added.
id: int
index: int
text: str
reasoning: str | None = None
# See `ActionOut.branch_id`. It decides whether choosing this attempt is a
# local step or a branch switch.
branch_id: int | None = None
created_at: str | None = None
active: bool = False
class VariantSelect(BaseModel):
index: int = Field(ge=0)
class BranchOut(ORMModel):
"""One line through the story tree (Phase 14, SP5).
This carries enough to draw the tree and nothing more. `fork_depth` is where
this line leaves its parent, and `depth` is where it currently ends, so a
fork is two numbers rather than a walk. `own_actions` counts the turns played
on this branch itself. The rest of its story is borrowed from its ancestors,
which is why the number is smaller than a reader expects.
"""
id: int
parent_branch_id: int | None = None
fork_depth: int | None = None
depth: int
own_actions: int = 0
# M4: how many Save Points name a position on this line. Deleting the branch
# deletes them with its story, so the panel warns with a number rather than
# a vague caution. Zero for a line nobody has bookmarked, which is most.
save_points: int = 0
is_head: bool = False
# NULL for a branch nobody has named. The client labels those from the fork
# depth rather than the server inventing a name. See the column comment.
name: str | None = None
created_at: datetime
class BranchRename(BaseModel):
"""A name a player chose, or `null` to make the branch unnamed again."""
name: Annotated[str, Field(max_length=BRANCH_NAME_MAX)] | None = None
# ---------- Narrative state (M5) ----------
class StateGroup(BaseModel):
"""One labelled section of the state inspector.
Rows carry the key as well as the label, because a manual correction has to
name an entity and the user should not have to guess the identifier.
"""
title: str
rows: list[dict] = []
class NarrativeStateOut(BaseModel):
"""The authoritative state at the active head.
`groups` is the display form and `document` is the state itself. Both are
returned because they answer different questions: the panel renders the
first, and a correction form — or a test — needs the second to name a key.
"""
groups: list[StateGroup] = []
empty: bool = True
document: dict = {}
#: M11 (post-M8 finding D): entities that share a display name, keyed by the
#: name. Reported rather than refused — two people called Alice is ordinary
#: fiction — but reported, because until M11 it happened silently and one of
#: the finding's candidate failure modes is exactly this.
duplicate_names: dict[str, list[str]] = {}
#: M11: the changes in *this* correction that were refused, and why.
#:
#: `narrative/validate.py` states the rule — "what is never allowed is a
#: rejected event mutating anything, or a rejection being silent" — and until
#: M11 the human-facing half of it was missing. A correction where one event
#: of four was refused returned 201 with the other three applied and said
#: nothing, so the reader believed they had made a change they had not. The
#: refusals were recorded on the proposal for the audit trail; they were
#: simply never shown to the person who wrote them.
refused: list[dict] = []
class StateEventIn(BaseModel):
"""One typed event, as a client proposes it.
Deliberately loose about which fields are present: the event vocabulary is
defined in `narrative/events.py` and enforced by `narrative/validate.py`,
and duplicating those rules here would create a second, drifting copy of the
allowlist. What this model does is bound the shapes — a type that is a
string, values that are scalars, labels that are short strings — so a
payload cannot smuggle a structure past Pydantic and reach the validator as
something other than an event.
"""
model_config = ConfigDict(extra="allow")
type: Annotated[str, Field(max_length=60)]
class StateCorrection(BaseModel):
"""A manual correction: the user overruling what the story established.
`note` records why, in the user's words, and is kept on the proposal record
so the audit says more than "the user changed this".
"""
events: Annotated[list[StateEventIn], Field(min_length=1, max_length=20)]
note: Prose = ""
class StateEventOut(ORMModel):
"""One accepted change, for the audit view."""
id: int
action_id: int | None = None
branch_id: int | None = None
depth: int | None = None
# The reader-facing position, matching the Save Point panel's vocabulary.
turn: int | None = None
sequence: int = 0
event_type: str
payload: dict = {}
before: dict | None = None
source: str = "accepted_story"
created_at: datetime
# ---------- Save Points (M4) ----------
#
# "Save Point" is the user-facing term and `checkpoint` is the internal one
# (`BROWSER-UX-SPEC.md` §23). The wire format uses the internal name, as the
# rest of this module does.
class CheckpointOut(ORMModel):
"""One Save Point: a name and the position it names.
The position is reported three ways because the panel needs three different
things from it. `turn` is what a reader counts — the same `depth + 1` the
branch list shows. `depth` and `branch_id` are the coordinate itself.
`on_path` says whether the position lies on the story being read, which is
how the panel can tell a Save Point on this line from one naming a line the
story has left; restoring either works, but they are not the same offer.
`resolved` is false when the coordinate no longer names a live turn, which
an action deleted out of the middle of a story can do. Restore refuses such
a Save Point rather than moving the head somewhere approximate, so the list
says so before the button is pressed.
"""
id: int
adventure_id: int
name: str
note: str = ""
branch_id: int
depth: int
turn: int = 0
on_path: bool = True
resolved: bool = True
created_at: datetime
updated_at: datetime
class CheckpointCreate(BaseModel):
"""A Save Point at wherever the story is being read.
The position is not a field. A Save Point is made at the campaign's active
head, which the server already knows, and accepting a coordinate from the
client would be the second way to name a position — the thing this milestone
exists not to build.
"""
name: CheckpointName
note: Prose = ""
class CheckpointRename(BaseModel):
"""A new label, and nothing else.
There is deliberately no coordinate here. `STORY-BRANCH-SEMANTICS.md` §24
keeps a Save Point's meaning auditable by refusing to move one: rename it,
or delete it and make another where you are.
"""
name: CheckpointName | None = None
note: Prose | None = None
class ActionUpdate(BaseModel):
text: ActionText
class ActionCreate(BaseModel):
type: Literal["do", "say", "story", "continue"]
text: ActionText = ""
# The node this action is played after (SP9). Omitting it means the tip,
# which is what every ordinary turn uses.
#
# Naming an attempt the story moved past is what creates a branch. Stepping
# between attempts costs nothing and creates nothing, and the fork happens
# on the first text written below one. That is the first moment the player
# states which line they mean. Before it, they were reading.
after_id: int | None = None
class TakeCreate(BaseModel):
"""Another attempt at a turn (SP9).
`text` is what the player says instead, and it applies only when the turn was
the player's. An AI turn's other attempt is generated, so the field is
ignored there rather than rejected. The client makes the same request for
both, and the node type decides what happens.
"""
text: ActionText = ""
class AdventureOut(ORMModel):
id: int
scenario_id: int | None
title: str
memory: str
authors_note: str
ai_instructions: str
narration_length: str
story_summary: str
auto_summarize: bool
memory_bank_enabled: bool
persona_name: str
persona_pronouns: str
persona_desc: str
# M8. Read from the `canon_rules` property on the model, which pulls the
# sentence list out of the stored `campaign_canon` document.
canon_rules: list[str] = []
created_at: datetime
updated_at: datetime
story_cards: list[StoryCardOut] = []
# The newest window of the story, not all of it. Older pages arrive from
# `GET /{id}/actions` as the reader scrolls up. `action_count` is the whole
# story's length, which is how the client knows more actions exist above.
actions: list[ActionOut] = []
action_count: int = 0
# M3. Whether the history controls have anywhere to go from where the story
# is. The client cannot work either out for itself: `can_undo` needs the
# campaign opening, which may be off the top of the loaded window, and
# `can_redo` needs the retained future, which the client is never sent.
can_undo: bool = False
can_redo: bool = False
class ImportedAdventureOut(AdventureOut):
"""A campaign that has just been restored from a bundle (M9).
Exactly `AdventureOut` plus what could not be rebuilt. The extra field is on
a subclass rather than on the base, because "which of your search indexes
failed to rebuild" is a fact about one import and not a property of a
campaign — putting it on `AdventureOut` would attach it to every read of
every campaign forever.
An empty list is the ordinary answer and means the whole campaign, its
evidence and its derived indexes all landed. A non-empty one means the
authoritative import succeeded and a rebuildable index did not, which is a
distinction M9 requires a caller to be able to draw: the campaign is intact,
and Reindex is the repair.
"""
import_warnings: list[str] = []
class ActionPage(BaseModel):
"""A slice of the story, counted back from the newest action."""
actions: list[ActionOut] = []
total: int = 0
# Whether anything older than this slice exists. The server computes it, so
# the client never has to do arithmetic on positions to find the end.
has_more: bool = False
# The same two flags `AdventureOut` carries, so that the response to Undo,
# Redo or a turn updates the controls without a second request.
can_undo: bool = False
can_redo: bool = False
# ---------- Memory bank (Phase 6) ----------
class MemoryOut(ORMModel):
id: int
adventure_id: int
text: str
pinned: bool
forgotten: bool
embedded: bool # model property: embedding vector present
use_count: int
last_used_at: datetime | None
source_start: int | None
source_end: int | None
created_at: datetime
class MemoryCreate(BaseModel):
text: Annotated[str, Field(max_length=MEMORY_TEXT_MAX)]
class MemoryUpdate(BaseModel):
text: Annotated[str, Field(max_length=MEMORY_TEXT_MAX)] | None = None
pinned: bool | None = None
forgotten: bool | None = None
# ---------------------------------------------------------------- M7: knowledge
class KnowledgeSourceOut(BaseModel):
"""One imported source, as a list row.
Deliberately without `content`. A library of twenty files would otherwise
put every byte of every one of them on a screen that shows none of it;
`KnowledgeSourceDetail` is what serves the text when it is asked for.
"""
id: int
title: str
original_filename: str
classification: str
enabled: bool
visibility: str
always_include: bool
content_hash: str
byte_size: int
media_type: str
chunk_count: int
embedded_count: int
# The two halves of derived state, kept apart on purpose. Lexical retrieval
# is a supported production path, so "the vectors failed" and "the index
# failed" are different sentences with different consequences.
index_state: str
index_detail: str
embed_state: str
embed_detail: str
parser_version: int
chunking_version: int
imported_at: str | None = None
updated_at: str | None = None
class KnowledgeSourceDetail(KnowledgeSourceOut):
"""A source with its text, for the inspector.
`content` is the file as it was decoded, not the normalized form used for
hashing and search: the reader inspects what they imported
(`IMPORTED-KNOWLEDGE-DESIGN.md` §61).
"""
content: str
notes: str = ""
class KnowledgeChunkOut(BaseModel):
id: int
chunk_index: int
heading_path: str
text: str
token_count: int
content_hash: str
embedded: bool
embedding_model: str = ""
class KnowledgeSourceUpdate(BaseModel):
"""What a reader may change about a source without reimporting it.
Everything here is metadata or state. Nothing rewrites content, and nothing
is destructive: changing a classification re-frames and re-weights the same
passages, and disabling a source removes it from retrieval while leaving the
rows exactly where they are.
"""
title: str | None = None
classification: str | None = None
enabled: bool | None = None
visibility: str | None = None
always_include: bool | None = None
notes: str | None = None
class AdventureListItem(ORMModel):
id: int
scenario_id: int | None
scenario_title: str | None = None
title: str
updated_at: datetime
action_count: int = 0
# The end of the most recent narration, so a Continue card can show the
# story rather than only a turn count.
snippet: str = ""
# Cover art inherited from the parent scenario. See `app/images.py`.
image_url: str = ""
icon: str = ""
# ---------- Auth (Phase 8) ----------
class AuthCredentials(BaseModel):
email: Annotated[str, Field(max_length=320)] # VARCHAR(320).
# The upper bound keeps the scrypt cost constant. Without it, hashing a
# megabyte password would give an attacker free CPU time.
password: Annotated[str, Field(max_length=128)]
# ---------- Settings ----------
class SettingsOut(ORMModel):
endpoint_url: str
model: str
api_mode: str
temperature: float
max_output_tokens: int
context_token_budget: int
model_timeout_seconds: int
context_window_override: int | None
narrator_prompt: str
summary_model: str
embedding_model: str
memory_bank_capacity: int
memory_top_k: int
ScenarioOut.model_rebuild()
# ---------- AI Chat (power users) ----------
# A scratchpad for talking to a model directly, with no story framing. The
# server persists nothing, so these caps are per-request abuse limits only.
CHAT_MESSAGE_MAX = 100_000 # One message.
CHAT_TOTAL_MAX = 400_000 # The whole conversation sent per request.
CHAT_MESSAGES_MAX = 200 # Turns per request.
class ChatMessage(BaseModel):
role: Literal["system", "user", "assistant"]
content: Annotated[str, Field(max_length=CHAT_MESSAGE_MAX)]
class ChatRequest(BaseModel):
messages: Annotated[list[ChatMessage], Field(min_length=1, max_length=CHAT_MESSAGES_MAX)]
# If this field is empty or omitted, the user's configured model is used.
model: Name | None = None
temperature: Annotated[float, Field(ge=0, le=5)] | None = None
max_tokens: Annotated[int, Field(ge=1, le=100_000)] | None = None
class SettingsUpdate(BaseModel):
endpoint_url: Annotated[str, Field(max_length=500)] | None = None # VARCHAR(500).
model: Name | None = None
api_mode: Annotated[str, Field(max_length=20)] | None = None
temperature: Annotated[float, Field(ge=0, le=5)] | None = None
max_output_tokens: Annotated[int, Field(ge=1, le=100_000)] | None = None
context_token_budget: Annotated[int, Field(ge=256, le=200_000)] | None = None
# Seconds to wait for the model. The floor is high enough that a normal
# turn cannot trip it; the ceiling exists so that "wait longer" stays a
# number rather than becoming "wait forever".
model_timeout_seconds: Annotated[int, Field(ge=30, le=3600)] | None = None
# The window an inference server enforces, for servers that cannot be asked.
# Bounded like the budget it caps. It is never a way to *raise* the prompt
# past a window the server did report — `contextwindow._declared_or` — so
# the ceiling here only bounds what an operator can usefully claim.
context_window_override: Annotated[int, Field(ge=256, le=200_000)] | None = None
narrator_prompt: Prose | None = None
summary_model: Name | None = None
embedding_model: Name | None = None
memory_bank_capacity: Annotated[int, Field(ge=1, le=1000)] | None = None
memory_top_k: Annotated[int, Field(ge=1, le=50)] | None = None