Add a reasoning-off setting (-1 reasoning budget)
Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.
Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.
Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
This commit is contained in:
co-authored by
Claude Opus 5
parent
aed28d9295
commit
d66fd1d6a2
@@ -307,7 +307,8 @@ class Settings(Base):
|
||||
# and left reasoning models with nothing after their thinking.
|
||||
max_output_tokens: Mapped[int] = mapped_column(Integer, default=800)
|
||||
# Separate thinking budget for reasoning models (OpenRouter-style
|
||||
# `reasoning: {max_tokens}`); 0 = param not sent. Added on top of
|
||||
# `reasoning: {max_tokens}`); 0 = param not sent, -1 = reasoning explicitly
|
||||
# off (`reasoning: {effort: none}`). Added on top of
|
||||
# max_output_tokens so story output keeps its full budget.
|
||||
reasoning_max_tokens: Mapped[int] = mapped_column(Integer, default=0)
|
||||
context_token_budget: Mapped[int] = mapped_column(Integer, default=16384)
|
||||
|
||||
Reference in New Issue
Block a user