Add a reasoning-off setting (-1 reasoning budget)

Models like DeepSeek V4 Flash reason by default, and the reasoning budget
setting could only ever add thinking tokens - there was no value that turned
thinking off. A negative budget now sends `reasoning: {effort: "none"}`.

Uses effort:none rather than exclude:true deliberately - exclude still thinks
and still bills, it only hides the trace.

Zero keeps its old meaning (send no `reasoning` field at all) so endpoints that
reject unknown fields, like the default Ollama one, are unaffected. Reusing the
existing int column this way avoids a migration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
This commit is contained in:
parththakkar106
2026-08-02 11:10:35 +05:30
co-authored by Claude Opus 5
parent aed28d9295
commit d66fd1d6a2
5 changed files with 70 additions and 9 deletions
+14 -3
View File
@@ -39,7 +39,8 @@ class OpenAICompatibleProvider(Provider):
self.api_mode = api_mode # "chat" | "completion"
# Thinking budget for reasoning models, on top of max_tokens. 0 = the
# `reasoning` param is not sent (endpoints that don't know it may
# reject unknown fields).
# reject unknown fields); negative = explicitly ask the endpoint to
# turn reasoning off.
self.reasoning_max_tokens = reasoning_max_tokens
def _headers(self) -> dict:
@@ -50,8 +51,18 @@ class OpenAICompatibleProvider(Provider):
def _apply_reasoning_budget(self, body: dict) -> None:
"""Give reasoning models their own thinking budget (OpenRouter-style),
raising max_tokens so the actual output keeps its full budget."""
if self.reasoning_max_tokens > 0 and self.api_mode == "chat":
raising max_tokens so the actual output keeps its full budget.
A negative budget means the opposite: send `effort: "none"` to switch
reasoning off on models that do it by default (DeepSeek V4 Flash, say).
That's distinct from `exclude: true`, which still thinks — and bills —
but hides the trace. Zero stays "send nothing at all" so endpoints that
reject unknown fields (Ollama) keep working."""
if self.api_mode != "chat":
return
if self.reasoning_max_tokens < 0:
body["reasoning"] = {"effort": "none"}
elif self.reasoning_max_tokens > 0:
body["reasoning"] = {"max_tokens": self.reasoning_max_tokens}
body["max_tokens"] += self.reasoning_max_tokens