Reasoning-model UX: better empty-response error, higher default budget, live link

- When a model streams reasoning but no story text, return a specific,
  actionable error (raise Max output tokens / set a reasoning cap / use a
  non-reasoning model) instead of the opaque "empty response".
- Bump default max_output_tokens 400 -> 800: 400 truncated scenes and
  left reasoning models with no room after thinking. Affects new settings
  rows; existing users keep their value.
- README: add a "Try it live" link to the Render deployment.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TEpWMjnfPqzZ13nPMoGhs5
This commit is contained in:
parththakkar106
2026-07-20 17:05:29 +05:30
co-authored by Claude Opus 4.8
parent 8188ed3999
commit 82fb6c8535
3 changed files with 18 additions and 2 deletions
+4
View File
@@ -5,6 +5,10 @@ with your own AI model. Create scenarios, play open-ended adventures where an LL
world, and extend the engine with **JavaScript scripts compatible with real AI Dungeon
scripting**.
> ### ▶️ Try it live: **[ai-dnd-1gmp.onrender.com](https://ai-dnd-1gmp.onrender.com)**
> Play a demo scenario as a guest — no sign-up, no API key needed. (Hosted on Render's free
> tier, so the first load after it's been idle takes ~30–60s to wake up.)
Built with FastAPI + SQLite on the backend and React (Vite) on the frontend. Works with **any
OpenAI-compatible endpoint**: Ollama and LM Studio locally, or OpenRouter / OpenAI / Groq / vLLM
in the cloud — endpoint, key, and model are all runtime settings, and OpenRouter's free-tier
+3 -1
View File
@@ -239,7 +239,9 @@ class Settings(Base):
model: Mapped[str] = mapped_column(String(200), default="")
api_mode: Mapped[str] = mapped_column(String(20), default="chat") # chat|completion
temperature: Mapped[float] = mapped_column(Float, default=0.8)
max_output_tokens: Mapped[int] = mapped_column(Integer, default=400)
# 800 leaves room for a full scene; 400 tended to truncate mid-paragraph
# and left reasoning models with nothing after their thinking.
max_output_tokens: Mapped[int] = mapped_column(Integer, default=800)
# Separate thinking budget for reasoning models (OpenRouter-style
# `reasoning: {max_tokens}`); 0 = param not sent. Added on top of
# max_output_tokens so story output keeps its full budget.
+11 -1
View File
@@ -288,7 +288,17 @@ async def generate_turn(
text = "".join(chunks).strip()
if not text:
yield sse({"type": "error", "detail": "The AI returned an empty response."})
# If the model streamed reasoning but no story text, it spent its whole
# budget thinking — say so instead of a mysterious "empty response".
if reasoning_chunks:
detail = (
"The model used its entire token budget on reasoning and returned no "
'story text. Raise "Max output tokens" in Settings, set a "Reasoning '
'max tokens" cap, or switch to a non-reasoning model.'
)
else:
detail = "The AI returned an empty response."
yield sse({"type": "error", "detail": detail})
return
# onOutput