Files
interactive-story/backend/app/debuglog.py
T
parththakkar106andClaude Opus 5 a408c7b6f7 Lay the prompt out so the endpoint can cache most of it
Prompt caching bills on a shared prefix: the endpoint reuses the request up
to the first byte that differs from last time and no further. The live
world-state block sat third from the top of the system message, so every turn
re-priced the instructions, the plot essentials and the whole story history
underneath it. The retrieved memories and the rewritten summary did it again.

Everything fixed is emitted first now, and everything that moves goes after
the history, ordered least-volatile first — which is also where recency serves
it best, the reasoning that already put the emit reminder last. The three tail
sections that are last for their own reasons stay last. The moved sections are
still charged to the token budget; only their position changed.

Two smaller halves of the same problem. OpenRouter serves a model from
whichever upstream is free and each upstream holds its own cache, so a
deepseek model now names deepseek as its preferred upstream — a preference,
not a restriction, so a turn still runs if that upstream is down. And the
endpoint's usage block is read back off the response and kept per attempt, so
the hit rate shows up in Insights and the debug log instead of being assumed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DfMCsN1KBLsTqMkj5hSgrY
2026-08-23 06:31:30 +05:30

69 lines
1.8 KiB
Python

"""In-memory ring buffer of recent provider requests/responses for the debug page.
API keys never enter the log: only the request body (which carries no
credentials) and response text are recorded, both truncated.
"""
import itertools
from collections import deque
from datetime import datetime, timezone
MAX_ENTRIES = 30
MAX_TEXT = 6000
_entries: deque[dict] = deque(maxlen=MAX_ENTRIES)
_ids = itertools.count(1)
def _clip(text: str) -> str:
if len(text) <= MAX_TEXT:
return text
return text[:MAX_TEXT] + f"\n… [{len(text) - MAX_TEXT} more chars truncated]"
def _clip_obj(obj):
if isinstance(obj, str):
return _clip(obj)
if isinstance(obj, dict):
return {k: _clip_obj(v) for k, v in obj.items()}
if isinstance(obj, list):
return [_clip_obj(v) for v in obj]
return obj
def start_entry(url: str, model: str, body: dict) -> dict:
entry = {
"id": next(_ids),
"time": datetime.now(timezone.utc).isoformat(),
"url": url,
"model": model,
"request": _clip_obj(body),
"status": "pending",
"response": "",
"usage": None,
"error": None,
}
_entries.appendleft(entry)
return entry
def finish_entry(
entry: dict,
*,
response: str = "",
error: str | None = None,
usage: dict | None = None,
) -> None:
"""`usage` is the endpoint's own token accounting when it reported any.
On OpenRouter it carries `prompt_tokens_details.cached_tokens`, which is
the only direct read on whether the prompt prefix is actually being
cached — a number worth seeing beside the request that produced it."""
entry["response"] = _clip(response)
entry["usage"] = usage
entry["error"] = error
entry["status"] = "error" if error else "ok"
def recent() -> list[dict]:
return list(_entries)