Files
interactive-story/backend/app/routers/adventures/insights.py
T
JesseMarkowitzandClaude Opus 5 ef25b0a876 Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
2026-09-10 06:13:55 -04:00

102 lines
4.3 KiB
Python

"""Read-only views of the context an adventure would send, or did send.
The dry run assembles a prompt without calling the model. The per-action endpoint
returns the prompt a turn was actually generated from. Neither writes anything.
"""
from fastapi import Depends, HTTPException
from sqlalchemy.orm import Session
from ... import derived, memorybank, models, summaries
from ... import contextwindow
from ...context import ContextOverflow, build_context
from ...database import get_db
from ...knowledge import retrieval as knowledge_retrieval
from ..settings import get_settings
from .deps import CurrentUser, current_adventure, router
@router.get("/{adventure_id}/context")
async def dry_run_context(
db: Session = Depends(get_db),
user: models.User = CurrentUser,
adventure: models.Adventure = Depends(current_adventure),
):
"""Returns what the app would send to the AI if the player continued now."""
settings = get_settings(db, user)
memories = await memorybank.retrieve_memories(adventure, settings, update_stats=False)
# M7: retrieved here too, and by the same call the turn makes. A dry run
# that skipped the library would show a prompt the next turn will not send,
# which is the one thing this panel must never do.
knowledge = await knowledge_retrieval.retrieve(adventure, settings)
# M11: and by the same probe the turn makes, for the same reason — a panel
# that showed a 16,384-token budget while the next turn will be capped to
# 4,096 would be showing a prompt that is not the one about to be sent.
window = await contextwindow.probe(settings.endpoint_url, settings.model,
declared=settings.context_window_override)
try:
_, _, report = build_context(
adventure, settings, memories, knowledge=knowledge, window=window
)
except ContextOverflow as exc:
# M6: a dry run of a prompt that cannot be built is still an answer, and
# a more useful one than a 500. The reader opened this panel to find out
# what would be sent; "nothing, because the protected context does not
# fit, and here is by how much" is exactly that.
raise HTTPException(422, str(exc)) from exc
return report
@router.get("/{adventure_id}/derived")
def derived_status(
db: Session = Depends(get_db),
adventure: models.Adventure = Depends(current_adventure),
):
"""M6: whether background memory, summary and embedding work is healthy.
The surface that makes a dead memory bank findable. M2 shipped with the
whole bank failing inside a fire-and-forget task and nothing anywhere said
so — not the UI, not a log a player would read, not a failing test
(`BUILD-MILESTONES.md`, note from M2). This endpoint is where that now
shows.
"""
# Resolved once, not once per row: which summary the current head is
# entitled to. Asking inside the comprehension would be one query per
# summary, which is the shape M5 spent a finding removing.
eligible = summaries.current(db, adventure)
eligible_id = eligible.id if eligible is not None else None
status = derived.report(db, adventure.id)
return {
"status": status,
"failing": [row["kind"] for row in status if row["status"] == "failed"],
"summaries": [
{
"id": row.id,
"branch_id": row.branch_id,
"depth": row.depth,
"trigger": row.trigger,
"model": row.model_name,
"eligible": row.id == eligible_id,
"created_at": row.created_at.isoformat() if row.created_at else None,
"preview": row.text[:200],
}
for row in summaries.all_for(db, adventure)
],
}
@router.get("/{adventure_id}/actions/{action_id}/context")
def action_context(
adventure_id: int,
action_id: int,
db: Session = Depends(get_db),
adventure: models.Adventure = Depends(current_adventure),
):
action = db.get(models.Action, action_id)
if action is None or action.adventure_id != adventure_id:
raise HTTPException(404, "Action not found")
if action.context_snapshot is None:
raise HTTPException(404, "No context snapshot for this action")
return action.context_snapshot