Add a reasoning-off setting (-1 reasoning budget)

Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.

Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.

Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
This commit is contained in:
parththakkar106
2026-08-02 11:10:35 +05:30
co-authored by Claude Opus 5
parent aed28d9295
commit d66fd1d6a2
5 changed files with 70 additions and 9 deletions
+2 -1
View File
@@ -307,7 +307,8 @@ class Settings(Base):
# and left reasoning models with nothing after their thinking.
max_output_tokens: Mapped[int] = mapped_column(Integer, default=800)
# Separate thinking budget for reasoning models (OpenRouter-style
# `reasoning: {max_tokens}`); 0 = param not sent. Added on top of
# `reasoning: {max_tokens}`); 0 = param not sent, -1 = reasoning explicitly
# off (`reasoning: {effort: none}`). Added on top of
# max_output_tokens so story output keeps its full budget.
reasoning_max_tokens: Mapped[int] = mapped_column(Integer, default=0)
context_token_budget: Mapped[int] = mapped_column(Integer, default=16384)