Lay the prompt out so the endpoint can cache most of it

Prompt caching bills on a shared prefix: the endpoint reuses the request up
to the first byte that differs from last time and no further. The live
world-state block sat third from the top of the system message, so every turn
re-priced the instructions, the plot essentials and the whole story history
underneath it. The retrieved memories and the rewritten summary did it again.

Everything fixed is emitted first now, and everything that moves goes after
the history, ordered least-volatile first — which is also where recency serves
it best, the reasoning that already put the emit reminder last. The three tail
sections that are last for their own reasons stay last. The moved sections are
still charged to the token budget; only their position changed.

Two smaller halves of the same problem. OpenRouter serves a model from
whichever upstream is free and each upstream holds its own cache, so a
deepseek model now names deepseek as its preferred upstream — a preference,
not a restriction, so a turn still runs if that upstream is down. And the
endpoint's usage block is read back off the response and kept per attempt, so
the hit rate shows up in Insights and the debug log instead of being assumed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
This commit is contained in:
parththakkar106
2026-08-23 06:31:30 +05:30
co-authored by Claude Opus 5
parent 28322b4b82
commit a408c7b6f7
20 changed files with 425 additions and 25 deletions
+21
View File
@@ -1298,6 +1298,26 @@ function ScriptReport({ script }) {
)
}
// What the endpoint charged for the turn, and how much of the prompt it read
// back out of its cache instead of billing in full. Only shown on a past turn:
// the "next turn" view has not been sent anywhere yet, so it has no usage. A
// cached read costs a tenth of a fresh one, which is the whole reason the
// prompt is laid out static-first — so this is the number that says whether
// that layout is working.
function CacheReport({ usage }) {
if (!usage) return null
const prompt = usage.prompt_tokens || 0
const cached = usage.prompt_tokens_details?.cached_tokens || 0
if (!prompt) return null
const pct = Math.round((cached / prompt) * 100)
return (
<div className="insights-history">
Prompt cache: {cached} of {prompt} prompt tokens read from cache ({pct}%)
{usage.cost != null && ` · cost $${Number(usage.cost).toFixed(5)}`}
</div>
)
}
const LEGEND_VISIBLE = 6 // enough to cover what actually moves the budget
// Where the prompt's tokens actually went: one stacked bar scaled to the
@@ -1435,6 +1455,7 @@ function InsightsPanel({ advId, inspectActionId, onClearInspect, refreshKey }) {
{history.total > history.included && ' (older history trimmed)'}
{history.oldest_truncated && ' — oldest entry cut mid-text'}
</div>
<CacheReport usage={report.usage} />
{cards.length > 0 && (
<div className="insights-cards">
{cards.map((c, i) => (