Files
interactive-story/backend/app/context/builder.py
JesseMarkowitzandClaude Opus 5 ef25b0a876 Stop re-reading the whole prompt every turn, and let a lost run carry on
M01, the hundred-turn campaign, is the one REQUIRED test still
outstanding. Everything here is about it finishing, and being worth
believing when it does. No requirement changed, no acceptance test was
retired or relaxed, and M11 §P.1's "no performance requirement" still
stands: what changed is the cost of a turn, not what a turn contains.

An inference server caches a prompt by its prefix. The history window
gave up its oldest action every turn, which changed the prompt near the
front and threw that cache away, so nearly the whole prompt was
reprocessed every turn however little had actually changed. The window
now snaps the oldest depth to a block and holds it, stepping every few
turns. Measured on real builder output at an 8,192-token budget: 124.0s
per turn against 362.4s. The cost is history depth, bounded by
TRIM_FRACTION at a quarter of the window, which is the dial between
recent history and speed.

A run that dies no longer starts again from turn one. m11_long_run
checkpoints resume.json after the prologue, after every scheduled step
and after every turn, and --resume reattaches to the same campaign. A
finished run deletes it, so the file's presence means an unfinished run
and starting fresh over one is refused. The model timeout is an option
rather than a hard-coded 600s, a turn that overruns is a failed turn
instead of an unhandled exception that ends the run with no summary,
and a run that has stopped producing turns writes its evidence and
stops.

Two checks could not fail. M04's planted clue went into an add_fact
"detail" key that the event does not define, so it was dropped and
fact_still_in_state could never be true; it is now in "value" and
proved at turn one, which stops a run measuring nothing for hours.
m11_browser degraded silently without a narrator into two failures that
read exactly like a product regression, and now requires one, with
--no-narrator as an explicit opt-out that marks the run partial.

Window discovery speaks Ollama's native API, so against vLLM or
llama.cpp's own server the window goes unverified and the budget
uncapped -- M11's own failure mode reached by another route.
context_window_override lets the operator state what they launched the
server with, and is used only where discovery left a hole: a verified
window always wins, so a declaration can lower an unknown ceiling into
existence and never raise a known one. "verified" still means the
server answered, so window_verified in a turn's provenance keeps the
meaning M11's report counts on.

planning/README.md said the M11 tree was staged rather than committed,
in two places; it was committed and signed. Planning package v3.8.

Backend 1,376 passed, 17 skipped, 0 failed; frontend 161; lint and
build clean. Every M11 harness re-run on this tree: browser 38/0/0,
offline 23/0, identity clean, contrast unchanged, recovery 14/0 on a
small bundle. M01 itself has not been run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9LiyxBxnMTXV2wRjdyDGB
2026-09-10 06:13:55 -04:00

913 lines
46 KiB
Python

"""Context assembly per AI Dungeon's memory system
(help.aidungeon.com/faq/the-memory-system):
[AI Instructions] always included
[Player Character] always included when the adventure has a persona
[Plot Essentials] always included (classic "Memory")
[Story Summary] always included (manual in Phase 3, auto in Phase 6)
[Used Memories] top-K memory-bank retrievals (Phase 6, when enabled)
[Triggered Story Cards] "World Lore: <entry>", conditional; first dropped when over budget
[Story history] newest actions that fit the remaining token budget
[Author's Note] injected AUTHORS_NOTE_DEPTH actions before the end of history
[Latest player action] (+ script frontMemory right after it, Phase 4)
The list above comes from AI Dungeon's design. The order does not. This module
emits every fixed section first and every changing section after the history,
because prompt caching bills on a shared prefix. A section that changes near the
top of the prompt re-prices everything below it. See the comments on the static
block and the live sections in `build_context`.
"""
from dataclasses import dataclass
import tiktoken
from sqlalchemy.orm import object_session
from .. import contextwindow, derived, models, narrative, summaries, worldstate
from ..knowledge import inject as knowledge_inject
from ..knowledge import records as knowledge_records
from . import encoding, history
AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
# `CARD_BUDGET_SHARE = 0.4` was here, and is gone with the injection it bounded
# (M9). It is named rather than deleted silently because two other places
# reasoned about their own share against it.
NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene")
SEPARATOR = "\n\n"
#: How much of the history window one trim gives up, as one-over-this. A
#: quarter: large enough that the window then holds still for several turns,
#: small enough that the narrator never loses most of its recent history at once.
#:
#: **This is the dial.** Lower it for bigger blocks — fewer prompt re-reads and
#: faster long campaigns, at the cost of retaining less recent history. Raise it
#: for the reverse. Nothing else has to change: `trim_block` is the only reader,
#: and `test_trim_fraction_is_the_dial_between_history_and_speed` pins that.
#: Measured at 4, on an 8,192-token budget: 124.0s per turn against 362.4s with
#: trimming off.
TRIM_FRACTION = 4
#: Never trim less than this, or the window slides by one action again and the
#: whole point is lost.
MIN_TRIM_BLOCK = 2
# Output-length guidance. The endpoint enforces `max_output_tokens` as a hard
# limit, and it truncates the reply mid-sentence when the model reaches it. The
# state block is emitted last, so truncation removes it. Asking the model to
# finish inside the limit prevents the truncation.
LENGTH_HEADROOM = 50 # Tokens reserved from the cap for the state block.
# Models cannot count their own tokens, but they do follow a word budget, so the
# hint states a number of words. English prose averages 0.75 words per token.
WORDS_PER_TOKEN = 0.75
# Models regularly exceed a word budget, and the cap it protects is a hard
# limit. Aiming 10% below the real ceiling leaves room for that overshoot, so it
# does not consume the state block.
LENGTH_BUFFER = 0.90
MIN_LENGTH_HINT_WORDS = 40 # Below this, the hint adds nothing useful.
# A ceiling on its own gives one-sided guidance, and models respond to it
# differently. A verbose model treats it as a limit. A terse model has only the
# instruction to write as much as the moment needs, and it produces two
# paragraphs. Adding a floor turns the guidance into a range, so the same prompt
# produces a similar length from either model. The floor is a share of the
# ceiling so that it can never approach the ceiling.
LENGTH_FLOOR_SHARE = 0.35
# Below this word count, a floor means nothing, because a short turn is the
# correct turn at a tight cap. The wording used at a tight cap is also the
# wording that was measured to preserve the state block, so it is unchanged.
MIN_LENGTH_FLOOR_WORDS = 60
# The floor prevents a collapse to two paragraphs. It does not ask for an essay.
# At a 2400-token cap, the share alone would request a minimum of 555 words. A
# reader who wants longer turns can ask for them in the author's note.
MAX_LENGTH_FLOOR_WORDS = 300
#: M11, post-M8 finding C: what the campaign's own narration-length choice means
#: in words. Until M11 the choice became one English sentence in the campaign's
#: instructions and moved no number at all, while the numeric hint below was
#: derived from the *global* `max_output_tokens` and therefore read identically
#: for brief, medium and long — at the default cap, "must not exceed 506 words,
#: and it should not stop short of about 177" whichever the reader picked. A
#: setting with a visible control and no measurable effect is worse than no
#: setting, because the reader spends trust on it.
#:
#: These bands are (floor, ceiling) in words. They are a design decision made
#: here rather than a ratified requirement — `BUILD-MILESTONES.md` records
#: "Brief ~100-200 words" as a candidate — and they are deliberately wide enough
#: that a scene can breathe inside one.
LENGTH_BANDS = {
"brief": (70, 180),
"medium": (150, 380),
"long": (320, 700),
}
#: Where the floor lands when a band's ceiling has to be cut down to fit the
#: token cap: keep it proportional rather than letting it collide with the
#: ceiling.
BAND_FLOOR_SHARE = 0.5
# Built from the table vendored in `encoding.py`, not fetched: the upstream
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
# is called on every turn.
# M6: added to the configured reply budget when reserving output space. It
# absorbs the section separators added after budgeting and the drift between
# this tokenizer and the serving model's. Fixed rather than proportional: what
# it covers does not grow with the size of the budget.
OUTPUT_SAFETY_MARGIN = 64
class ContextOverflow(RuntimeError):
"""Raised when protected context alone cannot fit in the token budget.
Protected means the narrator rules, the campaign canon, the authoritative
narrative state, the reader's own input, and the reserve for the reply
(`CONTEXT-AND-MEMORY.md` §30). None of those may be dropped to make room for
old prose, so when they do not fit there is no prompt to build and saying so
is the only honest answer.
"""
def _encoding() -> tiktoken.Encoding:
return encoding.get_encoding()
def count_tokens(text: str) -> int:
return len(_encoding().encode(text))
def truncate_to_last_tokens(text: str, budget: int) -> str:
tokens = _encoding().encode(text)
if len(tokens) <= budget:
return text
return _encoding().decode(tokens[-budget:])
@dataclass
class Section:
label: str
text: str
@property
def tokens(self) -> int:
return count_tokens(self.text)
def length_hint(max_output_tokens: int, narration_length: str = "") -> str:
"""Ask for a turn that fits inside the output cap, stated as a word budget.
Returns an empty string when the cap is too small to state usefully. The
model can exceed the hint, so the hint earns its tokens only when there is
enough room for that overshoot to stay inside the cap.
M11: `narration_length` is the campaign's own choice — `brief`, `medium` or
`long`, or empty for a campaign that never made one. It narrows the range
*within* what the token cap allows; it can never widen it, because the cap
is what the endpoint will actually emit and a hint that asked for more than
that would be asking for a truncated turn.
**The generation budget is deliberately not touched.** Capping
`max_output_tokens` per length would make a brief turn likelier to hit the
endpoint's limit mid-sentence, and the state block is emitted *last* — so
the first thing a truncated reply loses is the turn's state. That is the
trade `BUILD-MILESTONES.md` names when it says "do not hard-truncate prose".
"""
words = int((max_output_tokens - LENGTH_HEADROOM) * WORDS_PER_TOKEN * LENGTH_BUFFER)
if words < MIN_LENGTH_HINT_WORDS:
return ""
band = LENGTH_BANDS.get((narration_length or "").strip().lower())
if band is not None:
band_floor, band_ceiling = band
# The cap still wins. A `long` campaign on a 300-token reply cap gets
# the cap's number, not 700, and the floor moves down with it.
words = min(words, band_ceiling)
floor = min(band_floor, int(words * BAND_FLOOR_SHARE))
tail = (
" Finish the narration and append the state block well inside the limit."
)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
tail = " Finish the narration and append the state block well inside the limit."
# State the number as a ceiling, never as a budget. In measurements, the
# wording "keep this turn under about N words" read to the model as a target
# to fill. It raised the average from 174 words to 246 across five runs, and
# every hinted run was longer than every unhinted run. The hint therefore
# pushed turns toward the limit it exists to avoid. Naming the number as a
# limit, and adding that a typical turn is much shorter, held the average at
# 170 while still preserving the state block at tight caps.
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
if floor < MIN_LENGTH_FLOOR_WORDS:
return (
f"[Hard limit: this turn must not exceed {words} words. Write only as "
f"much as the moment needs — a typical turn is much shorter.{tail}]"
)
# Both numbers are bounds, and the wording is deliberately asymmetric. The
# ceiling uses "must not exceed", because the endpoint enforces it. The floor
# uses "should not stop short of". Neither reads as a target, which the
# measurement above shows is what matters. The clause that asks the model to
# prefer the lower end does the job the earlier wording did, which was to
# keep a verbose model away from the ceiling. It now has a number beneath it,
# so a terse model reading the same clause stops at the floor rather than at
# forty words.
return (
f"[Hard limit: this turn must not exceed {words} words, and it should not "
f"stop short of about {floor}. Prefer the lower end of that range unless "
f"the scene genuinely needs more.{tail}]"
)
def render_persona(adventure: models.Adventure) -> str:
"""Returns the Player Character section, or "" when there is no persona.
The three fields are independent. A name alone is enough, a description
alone is enough, and the wording holds together for either. Pronouns are
stated because the summarizer in `memorybank` writes about the protagonist
in the third person, and a model that has to infer a pronoun from a name
will sometimes infer wrongly and then repeat that error in every memory it
writes.
"""
name = adventure.persona_name.strip()
pronouns = adventure.persona_pronouns.strip()
desc = adventure.persona_desc.strip()
if not (name or desc):
return ""
head = f"You are {name}" if name else ""
if head and pronouns:
head += f" ({pronouns})"
# Joined with a space, not `SEPARATOR`: this is one short paragraph about
# one character, and a blank line inside it reads as two unrelated notes.
body = " ".join(part for part in (f"{head}." if head else "", desc) if part)
return f"Player character:\n{body}"
def _script_memory(adventure: models.Adventure) -> dict:
"""Script-provided memory overrides (populated by Phase 4 scripting)."""
state = adventure.script_state if isinstance(adventure.script_state, dict) else {}
memory = state.get("memory")
return memory if isinstance(memory, dict) else {}
def _history_text(action: models.Action) -> str:
"""Returns an AI turn as the model should see it in replayed history.
Replayed history is **prose only**. The protocol block is not reconstructed
into it, and the M5 corrective pass is why (review Finding 4).
Replaying the block was meant to teach the model the output format by
example. What it actually did was put a second, older account of the world
into the same prompt as the authoritative one, with nothing marking which
governed. A fact the reader had explicitly withdrawn through a manual
correction was dropped from the state section and then handed straight back
in the history section, as an accepted event, phrased exactly as the model
had first asserted it. C04 requires a correction to reach the narrator's
context; a correction the next prompt contradicts has not reached it.
Two other things were wrong with it. The blocks are implementation
metadata, not story, and every other consumer of stored text — memory,
summaries, export, the transcript — treats an action's text as prose. And a
turn's accepted events are a record of what was true *then*, which is
precisely what a later correction, retcon or invalidation revises.
The format instruction survives without the examples: `EMIT_RULE` carries a
worked example in the system block and `EMIT_REMINDER` repeats the demand
last, where recency is strongest.
"""
return action.text
def _memory_line(memory: dict) -> str:
"""One retrieved memory, marked with its authority (M6)."""
mark = " [inferred]" if memory.get("authority") == "heuristic" else ""
return f"-{mark} {memory['text']}"
def _canon_section(adventure: models.Adventure) -> str:
"""The campaign's own rules, rendered for the system block.
Canon is configuration (C01, J03): the campaign writes what is true and what
is forbidden, and both the prompt and the validator read the same field.
Putting it in the system block is what makes C01 a narration-time constraint
as well as a validation-time one — the model is told the rule rather than
only refused after breaking it.
"""
canon = adventure.campaign_canon
if not isinstance(canon, dict):
return ""
lines: list[str] = []
rules = canon.get("rules")
if isinstance(rules, list):
lines += [f"- {rule}" for rule in rules if isinstance(rule, str) and rule.strip()]
forbidden = canon.get("forbidden_status_changes")
if isinstance(forbidden, list):
for rule in forbidden:
if isinstance(rule, dict) and rule.get("from") and rule.get("to"):
lines.append(
f"- Nothing that is {rule['from']} can become {rule['to']}."
)
if not lines:
return ""
body = "\n".join(lines)
return f"Campaign canon (these are true and may not be contradicted):\n{body}"
def trim_block(history_budget: int, max_output_tokens: int) -> int:
"""How many `depth` steps of history one trim gives up.
Derived from **configuration**, never from the story, because the answer has
to be the same on two consecutive turns. A block size that moved with the
measured size of recent actions would move the boundary it defines, and a
boundary that moves is precisely what this exists to stop.
An AI action is bounded by `max_output_tokens` and a player action is small
beside it, so `max_output_tokens` is the scale of one row of history — a
setting, rather than a guess about the data.
"""
per_action = max(1, max_output_tokens)
fits = max(1, history_budget // per_action)
return max(MIN_TRIM_BLOCK, fits // TRIM_FRACTION)
def history_floor(depths: list[int | None], costs: list[int], budget: int,
block: int) -> int | None:
"""The depth of the oldest action to include, snapped to a block boundary.
## Why this is not just "whatever fits"
Taking whatever fits is what the builder did, and it is correct. It is also
the reason a long campaign costs a full prompt re-read every turn.
Inference servers cache the prompt they have already processed, keyed on the
**prefix**. While the story only grows at the end, each turn re-uses that
cache and pays for its own new tokens alone. As soon as the budget is full,
"whatever fits" drops the *oldest* action every turn — a change near the
front of the prompt — and everything after it has to be processed again.
So the floor is snapped forward to a multiple of `block` and then held. It
moves in steps: several cheap turns that re-use the cache, then one turn that
pays to re-read, rather than every turn paying. The cost is history depth —
right after a step the window holds up to `block` actions fewer than the
budget would allow, which is what `TRIM_FRACTION` bounds.
Measured against the reference deployment, on prompts this builder produced,
at an 8,192 budget where `block` is 3:
floor held, story grew by one action 14-20 s
floor stepped, prompt re-read 333-338 s
mean over two whole cycles 124.0 s
floor disabled, every turn re-read 362.4 s (361, 361, 365, 361)
2.9x, and the shape is the point rather than the ratio: the saving grows with
`block`, which grows with the budget, so the configuration that hurt most
before benefits most now.
Returns None when nothing needs trimming, which covers two cases that must
both stay as they were: a story short enough to fit whole (the window is a
growing prefix already, and snapping would drop its opening for no reason),
and an action so large that not even the newest one fits, which the caller
truncates.
"""
if not depths or any(depth is None for depth in depths):
# Legacy rows, or a path this cannot place on the tree. Trimming needs a
# stable coordinate; without one, behave exactly as before.
return None
spent = 0
oldest_fitting: int | None = None
for depth, cost in zip(reversed(depths), reversed(costs)):
if spent + cost > budget:
break
spent += cost
oldest_fitting = depth
if oldest_fitting is None:
return None
if oldest_fitting == depths[0]:
# Everything offered fits. There is nothing to drop, and snapping here
# would throw away the start of a short story to no purpose.
return None
block = max(1, block)
return -(-oldest_fitting // block) * block
def _visible_npcs(actions: list[models.Action], stat_schema: dict) -> dict[str, str]:
"""Returns the NPCs whose trigger words appear in the recent story.
These are the NPCs in scene, and the prompt includes stats for them only.
The result maps an NPC id to its display name.
`actions` holds only the most recent actions. See `NPC_WINDOW`.
"""
recent = SEPARATOR.join(a.text for a in actions).lower()
visible: dict[str, str] = {}
for npc_key, ndef in (stat_schema.get("npcs") or {}).items():
if not isinstance(ndef, dict):
continue
if any(trigger in recent for trigger in worldstate.npc_triggers(ndef, npc_key)):
visible[npc_key] = worldstate.npc_name(ndef, npc_key)
return visible
def match_cards(cards: list[models.StoryCard], window_text: str) -> list[dict]:
"""Returns one record per matched story card, naming the keyword that matched.
Matching follows AI Dungeon's rules. It ignores case, respects spaces, and
matches partial words, so "boat" matches "boats".
Public since Phase 18b: `memorybank.cast_brief` runs the same rule over the
block it is about to summarize, so that the summarizer is told who the
characters in that stretch of story are. One rule, one implementation.
"""
haystack = window_text.lower()
matched = []
for card in cards:
for key in (k.strip().lower() for k in card.keys.split(",")):
if key and key in haystack:
matched.append(
{"id": card.id, "name": card.name, "keyword": key, "entry": card.entry}
)
break
return matched
def build_context(
adventure: models.Adventure,
settings: models.Settings,
memory_bank: dict | None = None,
exclude_action_id: int | None = None,
knowledge: knowledge_records.Result | None = None,
window: contextwindow.Window | None = None,
) -> tuple[str, str, dict]:
"""Returns (system_text, story_text, context_report). `memory_bank` is the
result of memorybank.retrieve_memories (None when the bank is off);
`exclude_action_id` omits one action from the story (see history.py).
M7: `knowledge` is the result of `knowledge.retrieval.retrieve` — the ranked
imported passages, before any budget has been applied. It arrives already
retrieved for the same reason `memory_bank` does: retrieval may need an
embedding call, this function is synchronous, and a prompt builder that can
make network requests is a prompt builder that can fail halfway through a
prompt. None means the campaign has no library, or the caller did not ask.
M11: `window` is what the inference server was found to actually accept
(`contextwindow.probe`), and it arrives the same way and for the same
reason — asking the server is a network call and this function does not make
those. A **verified** window is a ceiling on the configured budget, which is
the whole of M11's no-silent-overflow invariant: the prompt this returns
cannot be longer than what the runtime will read, so `llama.cpp` never gets
the chance to drop the system block off the front. `None` means nobody
checked, and then the configured budget stands and the report says it was
not verified.
"""
# M11: the budget every section below is priced against. Capped by what the
# server was verified to accept; the configured value when nothing was
# verified, or when the reader has asked for something smaller.
budget = contextwindow.effective_budget(settings.context_token_budget, window)
script_mem = _script_memory(adventure)
# M7: priced before anything else, because the answer changes what is left.
# `plan` prices only the protected half — the untrusted-data rule and any
# always-in-force Canon — and both are counted with the system block below.
knowledge_plan = knowledge_inject.plan(
knowledge if knowledge is not None else knowledge_records.Result(),
count_tokens,
budget,
)
# ----- The static block, which is identical on every turn -----
# This ordering exists to reduce cost. Prompt caching matches a prefix. The
# endpoint reuses the prompt up to the first byte that differs from the
# previous request, and no further. A section that changes near the top
# therefore re-prices everything below it, and what sits below it is the
# story history, which is most of the prompt. Sections that change from turn
# to turn go after the history, among the live sections. Placing them there
# also gives them the most recency, which is why `EMIT_REMINDER` goes last.
system_sections: list[Section] = [Section("narrator", settings.narrator_prompt.strip())]
# RPG world state (Phase 12): the instructions for reporting changes. The
# live values go into a live section below. The guide derived from the
# schema and the emit rule do not change while the scenario is unchanged.
stat_schema = adventure.scenario.stat_schema if adventure.scenario else None
has_ws = worldstate.has_schema(stat_schema)
persona_name = adventure.persona_name.strip()
# M5: the typed-event protocol replaces the delta rule for every campaign,
# with or without an inherited stat schema. State is no longer an opt-in
# RPG layer — a story has entities, places and possessions whatever genre it
# is, so the rule is unconditional.
system_sections.append(Section("state_rule", narrative.extract.EMIT_RULE))
canon_text = _canon_section(adventure)
if canon_text:
system_sections.append(Section("campaign_canon", canon_text))
# M7: the imported-knowledge framing rule, and any Canon the campaign has
# marked as always in force. Both go here, directly *below* the campaign's
# own canon, which is the authority order stated in words in
# `knowledge.classes.KNOWLEDGE_RULE` and reinforced by the position.
#
# In the system block rather than among the live sections, for two reasons.
# They change only when the reader edits their library, so they belong in
# the cached prefix; and being counted with the protected sections is what
# makes an over-large always-include a `ContextOverflow` with an explanation
# rather than a prompt that silently loses its history.
for protected_section in knowledge_plan.protected:
system_sections.append(
Section(protected_section.label, protected_section.text)
)
if isinstance(script_mem.get("context"), str) and script_mem["context"].strip():
system_sections.append(Section("script_context", script_mem["context"].strip()))
if adventure.ai_instructions.strip():
system_sections.append(Section("ai_instructions", adventure.ai_instructions.strip()))
# Phase 18. This sits in the static block because only the user can edit it,
# so it never changes mid-story and stays inside the cached prefix. It is
# emitted whether or not the adventure has an RPG layer: an adventure with
# no stats still has a protagonist, and that is the case the persona was
# added for.
persona_text = render_persona(adventure)
if persona_text:
system_sections.append(Section("persona", persona_text))
if adventure.memory.strip():
system_sections.append(
Section("plot_essentials", f"Plot essentials:\n{adventure.memory.strip()}")
)
# ----- Live sections, which hold everything that changes -----
# This code builds them here and places them after the history further down.
# They are ordered from least to most volatile, so a turn that changes only
# the fastest-moving section leaves the others cached. The summary is
# rewritten every few turns. Lore changes with the scene. The retrieved
# memories change on most turns, and the stat values change on nearly every
# turn. `world_lore` is added below, because the history window determines
# which cards trigger and that window is not known yet.
# M6: the summary the *current lineage* is entitled to, not whatever was
# written last. A summary is derived data anchored to the story it covers,
# so an Undo or a divergence makes an old one ineligible rather than
# leaking it into a story it does not describe (E03, `app/summaries.py`).
db = object_session(adventure)
summary_row = summaries.current(db, adventure) if db is not None else None
summary_text = summary_row.text.strip() if summary_row is not None else ""
summary_section = (
Section("story_summary", f"Story summary:\n{summary_text}")
if summary_text
else None
)
memories_section = None
if memory_bank and memory_bank.get("used"):
# M6: an inference must not read as a record. A heuristic memory is
# marked in the prompt itself, because the narrator decides what to
# treat as established from what it is shown, and an unlabelled guess
# sitting beside accepted history is how a guess becomes canon
# (`CONTEXT-AND-MEMORY.md` §14). Authoritative state changes still come
# only from the M5 event path, whatever a memory says.
lines_text = "\n".join(_memory_line(m) for m in memory_bank["used"])
memories_section = Section(
"used_memories",
"Memories from earlier in the story. Lines marked [inferred] are "
"interpretation, not established fact — do not treat them as "
f"settled truth:\n{lines_text}",
)
world_state_section = None
refusal_note = ""
# M5: the authoritative narrative state, as the model is shown it. Read from
# the campaign's live document, which head movement keeps pointed at the
# position being read — so an undone story is described by the state it had
# then, not by the state it reached later.
state_block = narrative.render.for_prompt(adventure.narrative_state)
if state_block:
world_state_section = Section("narrative_state", state_block)
# Corrections for the previous AI turn only. A refusal the model has
# already had one chance to fix is stale, and repeating it every turn
# would price a correction into the whole rest of the adventure.
recent = history.tail(adventure, NPC_WINDOW, exclude_action_id)
last_ai = next((a for a in reversed(recent) if a.type == "ai"), None)
if last_ai is not None:
refusal_note = narrative.extract.render_rejections(last_ai.state_rejections)
authors_note_text = adventure.authors_note.strip()
if isinstance(script_mem.get("authorsNote"), str) and script_mem["authorsNote"].strip():
authors_note_text = script_mem["authorsNote"].strip()
authors_note = f"[Author's note: {authors_note_text}]" if authors_note_text else ""
front_memory = ""
if isinstance(script_mem.get("frontMemory"), str):
front_memory = script_mem["frontMemory"].strip()
length_note = length_hint(settings.max_output_tokens, adventure.narration_length)
# The live sections sit below the history, but they are still part of the
# prompt, so they still count against the budget. `world_lore` is the
# exception, because the code below budgets it out of `available`.
reserved = (
sum(s.tokens for s in system_sections)
+ sum(
s.tokens
for s in (summary_section, memories_section, world_state_section)
if s is not None
)
+ count_tokens(authors_note)
+ count_tokens(front_memory)
+ count_tokens(length_note)
+ count_tokens(narrative.extract.EMIT_REMINDER)
+ count_tokens(refusal_note)
)
# ----- M6: the output reserve, and what happens when it does not fit -----
#
# `context_token_budget` is the whole window the model is given, so the
# narrator's reply has to be subtracted from it before any history is
# chosen. Until M6 it was not: the builder spent the entire budget on input
# and left the reply to fit in whatever the endpoint had left, which is a
# truncated turn on a model whose window is the budget
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
#
# The margin covers what is added after this arithmetic — the separators
# between sections, and the difference between our tokenizer's count and the
# serving model's. It is small and fixed rather than proportional, because
# what it absorbs does not scale with the budget.
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
protected = reserved + output_reserve
if protected >= budget:
# Failing here is the point. The alternative — carrying on with a token
# or two of history — builds a prompt that is known to overflow, and
# the reader gets a truncated reply with no explanation. §32: "fail
# gracefully if protected context alone is too large."
raise ContextOverflow(
f"The protected context needs {protected} tokens "
f"({reserved} of prompt plus {output_reserve} reserved for the "
f"reply) but the context budget is {budget}. "
+ (
"That budget is what this server was found to accept, so raising "
"the setting alone will not help — load the model with a larger "
"window. Or lower the maximum reply length, or shorten the "
"campaign's canon, instructions and persona."
if budget < settings.context_token_budget else
"Raise the context budget, lower the maximum reply length, or "
"shorten the campaign's canon, instructions and persona."
)
)
available = budget - protected
# ----- M7: retrieved imported knowledge, out of a share of `available` -----
#
# Chosen here, before the history window is sized, because what knowledge
# spends is what the history does not get: a window fetched against the
# whole of `available` would read turns there was never room for.
#
# Bounded rather than trimmed afterwards. The passages that fit are selected
# against a share of the budget and the rest is recorded as dropped, so the
# section stops growing when the budget is exhausted however large the
# library becomes. Always-included Canon is not spent from this — it was
# priced into `reserved` above — so Reference and Inspiration cannot crowd
# out a standing campaign rule, and none of them can reach the current
# state, the reader's input or the reply reserve, which are all above.
knowledge_sections = [
Section(section.label, section.text)
for section in knowledge_inject.select(knowledge_plan, available)
]
knowledge_spent = sum(
section.tokens + count_tokens(SEPARATOR) for section in knowledge_sections
)
available_after_knowledge = max(0, available - knowledge_spent)
# Only the newest actions can reach the prompt, because the code below
# either truncates the text to `available` tokens or stops at the budget.
# Fetch a window that is provably larger than that and no larger. Otherwise
# a long adventure reads its whole history on every turn and uses only the
# end of it.
actions = history.window_covering(
adventure, available_after_knowledge, count_tokens, exclude_action_id
)
# ----- Story cards: legacy, and no longer part of the narrator's prompt (M9)
#
# Until M9 a keyword-triggered story card was injected here as
# `World Lore: <entry>`, taking up to 40% of what was left after the
# imported knowledge had been placed.
#
# `IMPORTED-KNOWLEDGE-DESIGN.md` §73 settles that Story Cards are not the
# production imported-knowledge store and, in as many words, that they "must
# not become an alternate untracked path around the new knowledge
# authority/provenance rules". That is exactly what this was. A card entry
# arrived in front of the narrator as a world fact with:
#
# * no class — nothing said whether it was Canon, Reference or Inspiration,
# so nothing framed how far the narrator could rely on it;
# * no visibility — no narrator-only distinction at all;
# * no source, no hash, no lifecycle, nothing to disable it with;
# * no browser surface, since M8 removed the editor — so a reader could
# neither see it nor switch it off;
# * and no row in the context inspector, which renders `knowledge` and
# never rendered `cards`.
#
# It also competed with imported Canon for one budget, which is the
# arrangement M7 spent a milestone separating.
#
# M9's decision, recorded in the milestone report: story cards are
# **compatibility-only legacy data**. Nothing is deleted. The rows stay, the
# `/api/story-cards` endpoints stay, the bundle carries them out and back so
# a round trip destroys nothing, and `memorybank.cast_brief` still reads them
# as the summariser's character roster — a roster names who is on stage so a
# memory says "Aldric" rather than "he", it never reaches the narrator, and
# every memory written from it is authority-classified by the application
# afterwards. What stops is the one path that asserted campaign facts to the
# narrator without any of the controls §73 requires.
#
# `cards` stays in the report and is now always empty for a new turn.
# Removing the key would break the historical snapshots that have one, which
# M9 has just made portable: an old turn's evidence says story cards were
# included, and it must go on saying so.
card_records: list[dict] = []
lore_section = None
used = 0
# ----- Story history: newest first until the remaining budget is spent -----
history_budget = available_after_knowledge - used
# Where the window starts, snapped to a block so it holds still for several
# turns instead of sliding by one action every turn. `history_floor` says
# why that matters and what it costs. None means trim nothing, and then
# everything below is exactly what it was before.
costs = [count_tokens(_history_text(a)) + count_tokens(SEPARATOR)
for a in actions]
block = trim_block(history_budget, settings.max_output_tokens)
floor_depth = history_floor([a.depth for a in actions], costs,
history_budget, block)
windowed = actions
if floor_depth is not None:
kept = [a for a in actions if a.depth is not None and a.depth >= floor_depth]
# A floor that leaves nothing is a floor worth ignoring: the loop below
# still has to produce a turn, and its own truncation path is the honest
# way to handle a single action larger than the whole budget.
if kept:
windowed = kept
else:
floor_depth = None
included_actions: list[models.Action] = []
spent = 0
oldest_truncated = False
for action in reversed(windowed):
# Budget against the text as it appears in the prompt, which includes
# the state block when this adventure tracks world state.
rendered = _history_text(action)
tokens = count_tokens(rendered) + count_tokens(SEPARATOR)
if spent + tokens > history_budget:
if not included_actions:
# Even the newest action alone is over budget: hard-truncate it.
included_actions.append(
models.Action(
adventure_id=action.adventure_id,
type=action.type,
text=truncate_to_last_tokens(action.text, history_budget),
)
)
oldest_truncated = True
break
included_actions.append(action)
spent += tokens
included_actions.reverse()
# ----- Assemble the story text, with the author's note near the end -----
# Append each AI turn's state block again. The app strips it before storage,
# and the recent history has to show the model the pattern to follow.
texts = [_history_text(a) for a in included_actions]
note_sections: list[Section] = []
if authors_note:
pos = max(0, len(texts) - AUTHORS_NOTE_DEPTH)
before, after = texts[:pos], texts[pos:]
if before:
note_sections.append(Section("history", SEPARATOR.join(before)))
note_sections.append(Section("authors_note", authors_note))
note_sections.append(Section("recent_history", SEPARATOR.join(after)))
else:
note_sections.append(Section("history", SEPARATOR.join(texts)))
# The live sections, ordered from least to most volatile. See the comment
# where they are built. They go below the history so that the history stays
# cached, and above the final sections so that those stay last.
#
# M7 inserts the retrieved knowledge between the lore and the memories, in
# ascending authority: Inspiration, then Reference, then imported Canon,
# then the story's own memories, and the current authoritative state last of
# all. A model weights what it read most recently, so the section it reads
# last is the one that settles a conflict — which is the ordering
# `knowledge.classes.KNOWLEDGE_RULE` states in words. Both are needed. C05
# is not satisfied by section order alone, and a stated order the layout
# contradicts is worse than either.
for live in (
summary_section,
lore_section,
*reversed(knowledge_sections),
memories_section,
world_state_section,
):
if live is not None:
note_sections.append(live)
if front_memory:
note_sections.append(Section("front_memory", front_memory))
# Place the length hint just above the emit reminder, which keeps the last
# position. The length budget applies to the narration, and the reminder
# applies to the block that follows it, so this is also the order in which
# the model acts.
note_sections.append(Section("length_hint", length_note))
# A correction for the previous turn sits directly above the reminder to
# emit a block, which is the instruction it modifies.
if refusal_note:
note_sections.append(Section("state_refusals", refusal_note))
# The emit rule sits in the system block, far from where the model
# generates text, so repeat it last where it has the most effect.
note_sections.append(Section("state_reminder", narrative.extract.EMIT_REMINDER))
story_sections = [s for s in note_sections if s.text]
system_text = SEPARATOR.join(s.text for s in system_sections if s.text)
story_text = SEPARATOR.join(s.text for s in story_sections)
all_sections = [s for s in system_sections if s.text] + story_sections
report = {
"sections": [
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
],
"prompt": {"system": system_text, "story": story_text},
# M6: the numbers the reader needs to answer "how much did each part
# cost, and what was left for the reply?" (F04, F05). `available` is
# what the history was actually allowed to spend after everything
# protected was subtracted.
"tokens": {
"total": count_tokens(system_text) + count_tokens(story_text),
"budget": budget,
"configured_budget": settings.context_token_budget,
"output_reserve": output_reserve,
"protected": reserved,
"available_for_history": available,
"history_spent": spent,
},
# M11: what the server was found to accept, and how. `verified` false
# means nobody could check — the prompt was built to the configured
# budget and may be larger than the runtime will read. This travels in
# the stored snapshot, so a turn taken against an unverified window is
# identifiable afterwards rather than indistinguishable from a safe one.
"window": {
"verified": (window.verified if window is not None else False),
"tokens": (window.tokens if window is not None else None),
"source": (window.source if window is not None else contextwindow.UNKNOWN),
"model_max": (window.model_max if window is not None else None),
"detail": (window.detail if window is not None else "not checked"),
# `enforceable`, not `verified`: an operator-declared window caps
# the prompt exactly as a server-reported one does, and a turn built
# against it *was* capped. `verified` and `source` above still say
# which kind of answer produced the number.
"capped": (
window is not None
and window.enforceable
and window.tokens < settings.context_token_budget
),
},
"cards": card_records,
"memories": memory_bank,
# M6: which summary was used, and which stretch of story it covers, so
# "what history did that summary cover?" is answerable from the record
# rather than by guessing (F05, F06).
"summary": summaries.provenance(summary_row),
# M6: whether background derived work is currently failing for this
# campaign. A dead memory bank is visible here rather than only in a log
# nobody reads (F08).
"derived": derived.report(db, adventure.id) if db is not None else [],
# M7: every imported passage this turn was given — which source, which
# file, which class, which visibility, which passage, how it was found,
# what each path scored it, and what it cost — plus what was considered,
# what was set aside as redundant, and what there was no budget for.
#
# The rendered text travels in this record, not a reference to the chunk
# row it came from. That is what makes a historical turn's evidence
# survive the source being deleted
# (`IMPORTED-KNOWLEDGE-DESIGN.md` §49-50): the snapshot says what the
# narrator was actually shown, and it goes on saying it.
"knowledge": knowledge_inject.report(knowledge_plan),
"history": {
"included": len(included_actions),
# The count covers the whole story rather than the window fetched
# above. Insights reports how many of the total actions it
# included, so this number must be the real total.
"total": history.count(adventure, exclude_action_id),
"oldest_truncated": oldest_truncated,
# Where the window was cut, and how big a step it takes when it
# moves. Both are in `depth` units. `floor_depth` is null while the
# story still fits whole, which is also while every turn is a pure
# prefix extension of the last one. A reader comparing two turns can
# tell from these whether the prompt's prefix was preserved.
"floor_depth": floor_depth,
"trim_block": block,
},
"settings": {
"model": settings.model,
"api_mode": settings.api_mode,
"temperature": settings.temperature,
"max_output_tokens": settings.max_output_tokens,
},
}
return system_text, story_text, report