Prompt caching bills on a shared prefix: the endpoint reuses the request up to the first byte that differs from last time and no further. The live world-state block sat third from the top of the system message, so every turn re-priced the instructions, the plot essentials and the whole story history underneath it. The retrieved memories and the rewritten summary did it again. Everything fixed is emitted first now, and everything that moves goes after the history, ordered least-volatile first — which is also where recency serves it best, the reasoning that already put the emit reminder last. The three tail sections that are last for their own reasons stay last. The moved sections are still charged to the token budget; only their position changed. Two smaller halves of the same problem. OpenRouter serves a model from whichever upstream is free and each upstream holds its own cache, so a deepseek model now names deepseek as its preferred upstream — a preference, not a restriction, so a turn still runs if that upstream is down. And the endpoint's usage block is read back off the response and kept per attempt, so the hit rate shows up in Insights and the debug log instead of being assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
35 lines
1.1 KiB
Python
35 lines
1.1 KiB
Python
from abc import ABC, abstractmethod
|
|
from dataclasses import dataclass
|
|
from typing import AsyncIterator
|
|
|
|
|
|
@dataclass
|
|
class PromptParts:
|
|
"""Assembled context, provider-agnostic. Providers map this to their wire format."""
|
|
|
|
system: str # narrator prompt + AI instructions + memory
|
|
story: str # the story text so far (already token-budgeted)
|
|
|
|
|
|
class ProviderError(Exception):
|
|
"""User-presentable provider failure (connection refused, bad key, model not found…)."""
|
|
|
|
|
|
class Provider(ABC):
|
|
# The endpoint's own token accounting for the most recent call, when it
|
|
# reported any — notably `prompt_tokens_details.cached_tokens`, which is
|
|
# the only direct read on whether the prompt prefix is being cached.
|
|
# Callers read it after the call they made; one provider is built per
|
|
# request, so there is nothing to race.
|
|
last_usage: dict | None = None
|
|
|
|
@abstractmethod
|
|
def generate(
|
|
self,
|
|
parts: PromptParts,
|
|
*,
|
|
temperature: float,
|
|
max_tokens: int,
|
|
) -> AsyncIterator[tuple[str, str]]:
|
|
"""Yield ("text" | "reasoning", chunk) pairs. Raises ProviderError on failure."""
|