Commit Graph
2 Commits
Author SHA1 Message Date
parththakkar106andClaude Opus 5 3b9e6b3d50 Give the length hint a floor, not just a wall
A ceiling alone is a one-sided instruction, and models read it in opposite
directions. A verbose one is held back by it; a terse one has nothing to act
on except "write only as much as the moment needs -- a typical turn is much
shorter" and collapses to two paragraphs. Same prompt, wildly different turn
lengths depending on which model is behind it.

State a floor as well, so the guidance is a band. The two bounds are
deliberately asymmetric -- "must not exceed" for the wall the endpoint
enforces, "should not stop short of" for the floor -- so neither reads as a
number to hit, which is the property the earlier A/B says decides whether
this hint helps or hurts. "Prefer the lower end" inherits the anti-overshoot
job the deleted "much shorter" line was doing, but now with a number under
it, so a terse model lands on the floor instead of at forty words.

Below MIN_LENGTH_FLOOR_WORDS the floor is dropped and the tight-cap wording
is left byte-identical: at a tight cap a short turn is the correct turn, and
that phrasing is the one measured to keep the state block alive (0/6
truncations at cap 250 against 2/6 unhinted). So this only moves loose caps.
MAX_LENGTH_FLOOR_WORDS keeps the share from demanding 555 words minimum at
cap 2400 -- a big cap means long turns are allowed, not compulsory.

Shipped without an A/B run, deliberately. Two things to watch live: whether a
stated range invites landing mid-range on verbose models (drop the share to
~0.25 if so), and whether the state block still survives -- nothing reads
finish_reason yet, so truncation is silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-21 00:14:56 +05:30
parththakkar106andClaude Opus 5 f295893204 Ask the model for a turn that fits inside the output cap
max_output_tokens is a hard wall the endpoint enforces mid-sentence. The
```state block is emitted after the narration, so a long turn hits the wall
partway through the block and the deltas are lost — silently, since nothing
reads finish_reason.

builder.length_hint() derives a word limit from the cap ((cap - 50 headroom)
* 0.75 words/token * 0.90 buffer) and injects it just above EMIT_REMINDER,
which keeps the last slot it needs. Reserved in build_context like the
reminder is.

Phrased as a ceiling, not a budget. Measured against gemma-4-26b at cap 800,
n=5 per arm: no hint 174 words, "keep this turn under about N words" 246,
"hard limit ... a typical turn is much shorter" 170. A budget reads as a
target to fill — every budget run was longer than every unhinted one, pushing
turns toward the wall the hint exists to avoid. Ceiling phrasing still works
at tight caps: at 250, unhinted hit finish_reason=length 2/6, hinted 0/6.

tests/test_length_hint.py, 11 tests; each mechanism verified by sabotage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UeQVy5bEjLhfgWNc27Efet
2026-08-08 12:20:26 +05:30