Reasoning-model UX: better empty-response error, higher default budget, live link
- When a model streams reasoning but no story text, return a specific, actionable error (raise Max output tokens / set a reasoning cap / use a non-reasoning model) instead of the opaque "empty response". - Bump default max_output_tokens 400 -> 800: 400 truncated scenes and left reasoning models with no room after thinking. Affects new settings rows; existing users keep their value. - README: add a "Try it live" link to the Render deployment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
8188ed3999
commit
82fb6c8535
@@ -5,6 +5,10 @@ with your own AI model. Create scenarios, play open-ended adventures where an LL
|
||||
world, and extend the engine with **JavaScript scripts compatible with real AI Dungeon
|
||||
scripting**.
|
||||
|
||||
> ### ▶️ Try it live: **[ai-dnd-1gmp.onrender.com](https://ai-dnd-1gmp.onrender.com)**
|
||||
> Play a demo scenario as a guest — no sign-up, no API key needed. (Hosted on Render's free
|
||||
> tier, so the first load after it's been idle takes ~30–60s to wake up.)
|
||||
|
||||
Built with FastAPI + SQLite on the backend and React (Vite) on the frontend. Works with **any
|
||||
OpenAI-compatible endpoint**: Ollama and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM
|
||||
in the cloud — endpoint, key, and model are all runtime settings, and OpenRouter's free-tier
|
||||
|
||||
@@ -239,7 +239,9 @@ class Settings(Base):
|
||||
model: Mapped[str] = mapped_column(String(200), default="")
|
||||
api_mode: Mapped[str] = mapped_column(String(20), default="chat") # chat|completion
|
||||
temperature: Mapped[float] = mapped_column(Float, default=0.8)
|
||||
max_output_tokens: Mapped[int] = mapped_column(Integer, default=400)
|
||||
# 800 leaves room for a full scene; 400 tended to truncate mid-paragraph
|
||||
# and left reasoning models with nothing after their thinking.
|
||||
max_output_tokens: Mapped[int] = mapped_column(Integer, default=800)
|
||||
# Separate thinking budget for reasoning models (OpenRouter-style
|
||||
# `reasoning: {max_tokens}`); 0 = param not sent. Added on top of
|
||||
# max_output_tokens so story output keeps its full budget.
|
||||
|
||||
@@ -288,7 +288,17 @@ async def generate_turn(
|
||||
|
||||
text = "".join(chunks).strip()
|
||||
if not text:
|
||||
yield sse({"type": "error", "detail": "The AI returned an empty response."})
|
||||
# If the model streamed reasoning but no story text, it spent its whole
|
||||
# budget thinking — say so instead of a mysterious "empty response".
|
||||
if reasoning_chunks:
|
||||
detail = (
|
||||
"The model used its entire token budget on reasoning and returned no "
|
||||
'story text. Raise "Max output tokens" in Settings, set a "Reasoning '
|
||||
'max tokens" cap, or switch to a non-reasoning model.'
|
||||
)
|
||||
else:
|
||||
detail = "The AI returned an empty response."
|
||||
yield sse({"type": "error", "detail": detail})
|
||||
return
|
||||
|
||||
# onOutput
|
||||
|
||||
Reference in New Issue
Block a user