Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
587 lines
28 KiB
Python
587 lines
28 KiB
Python
"""Context assembly per AI Dungeon's memory system
|
|
(help.aidungeon.com/faq/the-memory-system):
|
|
|
|
[AI Instructions] always included
|
|
[Player Character] always included when the adventure has a persona
|
|
[Plot Essentials] always included (classic "Memory")
|
|
[Story Summary] always included (manual in Phase 3, auto in Phase 6)
|
|
[Used Memories] top-K memory-bank retrievals (Phase 6, when enabled)
|
|
[Triggered Story Cards] "World Lore: <entry>", conditional; first dropped when over budget
|
|
[Story history] newest actions that fit the remaining token budget
|
|
[Author's Note] injected AUTHORS_NOTE_DEPTH actions before the end of history
|
|
[Latest player action] (+ script frontMemory right after it, Phase 4)
|
|
|
|
The list above comes from AI Dungeon's design. The order does not. This module
|
|
emits every fixed section first and every changing section after the history,
|
|
because prompt caching bills on a shared prefix. A section that changes near the
|
|
top of the prompt re-prices everything below it. See the comments on the static
|
|
block and the live sections in `build_context`.
|
|
"""
|
|
|
|
from dataclasses import dataclass
|
|
|
|
import tiktoken
|
|
from sqlalchemy.orm import object_session
|
|
|
|
from .. import derived, models, narrative, summaries, worldstate
|
|
from . import encoding, history
|
|
|
|
AUTHORS_NOTE_DEPTH = 3 # actions from the end of history
|
|
CARD_BUDGET_SHARE = 0.4 # max share of non-reserved budget that story cards may take
|
|
NPC_WINDOW = 6 # actions of story searched for NPC trigger words ("in scene")
|
|
SEPARATOR = "\n\n"
|
|
|
|
# Output-length guidance. The endpoint enforces `max_output_tokens` as a hard
|
|
# limit, and it truncates the reply mid-sentence when the model reaches it. The
|
|
# state block is emitted last, so truncation removes it. Asking the model to
|
|
# finish inside the limit prevents the truncation.
|
|
LENGTH_HEADROOM = 50 # Tokens reserved from the cap for the state block.
|
|
# Models cannot count their own tokens, but they do follow a word budget, so the
|
|
# hint states a number of words. English prose averages 0.75 words per token.
|
|
WORDS_PER_TOKEN = 0.75
|
|
# Models regularly exceed a word budget, and the cap it protects is a hard
|
|
# limit. Aiming 10% below the real ceiling leaves room for that overshoot, so it
|
|
# does not consume the state block.
|
|
LENGTH_BUFFER = 0.90
|
|
MIN_LENGTH_HINT_WORDS = 40 # Below this, the hint adds nothing useful.
|
|
# A ceiling on its own gives one-sided guidance, and models respond to it
|
|
# differently. A verbose model treats it as a limit. A terse model has only the
|
|
# instruction to write as much as the moment needs, and it produces two
|
|
# paragraphs. Adding a floor turns the guidance into a range, so the same prompt
|
|
# produces a similar length from either model. The floor is a share of the
|
|
# ceiling so that it can never approach the ceiling.
|
|
LENGTH_FLOOR_SHARE = 0.35
|
|
# Below this word count, a floor means nothing, because a short turn is the
|
|
# correct turn at a tight cap. The wording used at a tight cap is also the
|
|
# wording that was measured to preserve the state block, so it is unchanged.
|
|
MIN_LENGTH_FLOOR_WORDS = 60
|
|
# The floor prevents a collapse to two paragraphs. It does not ask for an essay.
|
|
# At a 2400-token cap, the share alone would request a minimum of 555 words. A
|
|
# reader who wants longer turns can ask for them in the author's note.
|
|
MAX_LENGTH_FLOOR_WORDS = 300
|
|
|
|
|
|
# Built from the table vendored in `encoding.py`, not fetched: the upstream
|
|
# `tiktoken.get_encoding("cl100k_base")` downloads it on first use, and this
|
|
# is called on every turn.
|
|
# M6: added to the configured reply budget when reserving output space. It
|
|
# absorbs the section separators added after budgeting and the drift between
|
|
# this tokenizer and the serving model's. Fixed rather than proportional: what
|
|
# it covers does not grow with the size of the budget.
|
|
OUTPUT_SAFETY_MARGIN = 64
|
|
|
|
|
|
class ContextOverflow(RuntimeError):
|
|
"""Raised when protected context alone cannot fit in the token budget.
|
|
|
|
Protected means the narrator rules, the campaign canon, the authoritative
|
|
narrative state, the reader's own input, and the reserve for the reply
|
|
(`CONTEXT-AND-MEMORY.md` §30). None of those may be dropped to make room for
|
|
old prose, so when they do not fit there is no prompt to build and saying so
|
|
is the only honest answer.
|
|
"""
|
|
|
|
|
|
def _encoding() -> tiktoken.Encoding:
|
|
return encoding.get_encoding()
|
|
|
|
|
|
def count_tokens(text: str) -> int:
|
|
return len(_encoding().encode(text))
|
|
|
|
|
|
def truncate_to_last_tokens(text: str, budget: int) -> str:
|
|
tokens = _encoding().encode(text)
|
|
if len(tokens) <= budget:
|
|
return text
|
|
return _encoding().decode(tokens[-budget:])
|
|
|
|
|
|
@dataclass
|
|
class Section:
|
|
label: str
|
|
text: str
|
|
|
|
@property
|
|
def tokens(self) -> int:
|
|
return count_tokens(self.text)
|
|
|
|
|
|
def length_hint(max_output_tokens: int) -> str:
|
|
"""Ask for a turn that fits inside the output cap, stated as a word budget.
|
|
|
|
Returns an empty string when the cap is too small to state usefully. The
|
|
model can exceed the hint, so the hint earns its tokens only when there is
|
|
enough room for that overshoot to stay inside the cap.
|
|
"""
|
|
words = int((max_output_tokens - LENGTH_HEADROOM) * WORDS_PER_TOKEN * LENGTH_BUFFER)
|
|
if words < MIN_LENGTH_HINT_WORDS:
|
|
return ""
|
|
tail = " Finish the narration and append the state block well inside the limit."
|
|
|
|
# State the number as a ceiling, never as a budget. In measurements, the
|
|
# wording "keep this turn under about N words" read to the model as a target
|
|
# to fill. It raised the average from 174 words to 246 across five runs, and
|
|
# every hinted run was longer than every unhinted run. The hint therefore
|
|
# pushed turns toward the limit it exists to avoid. Naming the number as a
|
|
# limit, and adding that a typical turn is much shorter, held the average at
|
|
# 170 while still preserving the state block at tight caps.
|
|
floor = min(int(words * LENGTH_FLOOR_SHARE), MAX_LENGTH_FLOOR_WORDS)
|
|
if floor < MIN_LENGTH_FLOOR_WORDS:
|
|
return (
|
|
f"[Hard limit: this turn must not exceed {words} words. Write only as "
|
|
f"much as the moment needs — a typical turn is much shorter.{tail}]"
|
|
)
|
|
# Both numbers are bounds, and the wording is deliberately asymmetric. The
|
|
# ceiling uses "must not exceed", because the endpoint enforces it. The floor
|
|
# uses "should not stop short of". Neither reads as a target, which the
|
|
# measurement above shows is what matters. The clause that asks the model to
|
|
# prefer the lower end does the job the earlier wording did, which was to
|
|
# keep a verbose model away from the ceiling. It now has a number beneath it,
|
|
# so a terse model reading the same clause stops at the floor rather than at
|
|
# forty words.
|
|
return (
|
|
f"[Hard limit: this turn must not exceed {words} words, and it should not "
|
|
f"stop short of about {floor}. Prefer the lower end of that range unless "
|
|
f"the scene genuinely needs more.{tail}]"
|
|
)
|
|
|
|
|
|
def render_persona(adventure: models.Adventure) -> str:
|
|
"""Returns the Player Character section, or "" when there is no persona.
|
|
|
|
The three fields are independent. A name alone is enough, a description
|
|
alone is enough, and the wording holds together for either. Pronouns are
|
|
stated because the summarizer in `memorybank` writes about the protagonist
|
|
in the third person, and a model that has to infer a pronoun from a name
|
|
will sometimes infer wrongly and then repeat that error in every memory it
|
|
writes.
|
|
"""
|
|
name = adventure.persona_name.strip()
|
|
pronouns = adventure.persona_pronouns.strip()
|
|
desc = adventure.persona_desc.strip()
|
|
if not (name or desc):
|
|
return ""
|
|
head = f"You are {name}" if name else ""
|
|
if head and pronouns:
|
|
head += f" ({pronouns})"
|
|
# Joined with a space, not `SEPARATOR`: this is one short paragraph about
|
|
# one character, and a blank line inside it reads as two unrelated notes.
|
|
body = " ".join(part for part in (f"{head}." if head else "", desc) if part)
|
|
return f"Player character:\n{body}"
|
|
|
|
|
|
def _script_memory(adventure: models.Adventure) -> dict:
|
|
"""Script-provided memory overrides (populated by Phase 4 scripting)."""
|
|
state = adventure.script_state if isinstance(adventure.script_state, dict) else {}
|
|
memory = state.get("memory")
|
|
return memory if isinstance(memory, dict) else {}
|
|
|
|
|
|
def _history_text(action: models.Action) -> str:
|
|
"""Returns an AI turn as the model should see it in replayed history.
|
|
|
|
Replayed history is **prose only**. The protocol block is not reconstructed
|
|
into it, and the M5 corrective pass is why (review Finding 4).
|
|
|
|
Replaying the block was meant to teach the model the output format by
|
|
example. What it actually did was put a second, older account of the world
|
|
into the same prompt as the authoritative one, with nothing marking which
|
|
governed. A fact the reader had explicitly withdrawn through a manual
|
|
correction was dropped from the state section and then handed straight back
|
|
in the history section, as an accepted event, phrased exactly as the model
|
|
had first asserted it. C04 requires a correction to reach the narrator's
|
|
context; a correction the next prompt contradicts has not reached it.
|
|
|
|
Two other things were wrong with it. The blocks are implementation
|
|
metadata, not story, and every other consumer of stored text — memory,
|
|
summaries, export, the transcript — treats an action's text as prose. And a
|
|
turn's accepted events are a record of what was true *then*, which is
|
|
precisely what a later correction, retcon or invalidation revises.
|
|
|
|
The format instruction survives without the examples: `EMIT_RULE` carries a
|
|
worked example in the system block and `EMIT_REMINDER` repeats the demand
|
|
last, where recency is strongest.
|
|
"""
|
|
return action.text
|
|
|
|
|
|
def _memory_line(memory: dict) -> str:
|
|
"""One retrieved memory, marked with its authority (M6)."""
|
|
mark = " [inferred]" if memory.get("authority") == "heuristic" else ""
|
|
return f"-{mark} {memory['text']}"
|
|
|
|
|
|
def _canon_section(adventure: models.Adventure) -> str:
|
|
"""The campaign's own rules, rendered for the system block.
|
|
|
|
Canon is configuration (C01, J03): the campaign writes what is true and what
|
|
is forbidden, and both the prompt and the validator read the same field.
|
|
Putting it in the system block is what makes C01 a narration-time constraint
|
|
as well as a validation-time one — the model is told the rule rather than
|
|
only refused after breaking it.
|
|
"""
|
|
canon = adventure.campaign_canon
|
|
if not isinstance(canon, dict):
|
|
return ""
|
|
lines: list[str] = []
|
|
rules = canon.get("rules")
|
|
if isinstance(rules, list):
|
|
lines += [f"- {rule}" for rule in rules if isinstance(rule, str) and rule.strip()]
|
|
forbidden = canon.get("forbidden_status_changes")
|
|
if isinstance(forbidden, list):
|
|
for rule in forbidden:
|
|
if isinstance(rule, dict) and rule.get("from") and rule.get("to"):
|
|
lines.append(
|
|
f"- Nothing that is {rule['from']} can become {rule['to']}."
|
|
)
|
|
if not lines:
|
|
return ""
|
|
body = "\n".join(lines)
|
|
return f"Campaign canon (these are true and may not be contradicted):\n{body}"
|
|
|
|
|
|
def _visible_npcs(actions: list[models.Action], stat_schema: dict) -> dict[str, str]:
|
|
"""Returns the NPCs whose trigger words appear in the recent story.
|
|
|
|
These are the NPCs in scene, and the prompt includes stats for them only.
|
|
The result maps an NPC id to its display name.
|
|
|
|
`actions` holds only the most recent actions. See `NPC_WINDOW`.
|
|
"""
|
|
recent = SEPARATOR.join(a.text for a in actions).lower()
|
|
visible: dict[str, str] = {}
|
|
for npc_key, ndef in (stat_schema.get("npcs") or {}).items():
|
|
if not isinstance(ndef, dict):
|
|
continue
|
|
if any(trigger in recent for trigger in worldstate.npc_triggers(ndef, npc_key)):
|
|
visible[npc_key] = worldstate.npc_name(ndef, npc_key)
|
|
return visible
|
|
|
|
|
|
def match_cards(cards: list[models.StoryCard], window_text: str) -> list[dict]:
|
|
"""Returns one record per matched story card, naming the keyword that matched.
|
|
|
|
Matching follows AI Dungeon's rules. It ignores case, respects spaces, and
|
|
matches partial words, so "boat" matches "boats".
|
|
|
|
Public since Phase 18b: `memorybank.cast_brief` runs the same rule over the
|
|
block it is about to summarize, so that the summarizer is told who the
|
|
characters in that stretch of story are. One rule, one implementation.
|
|
"""
|
|
haystack = window_text.lower()
|
|
matched = []
|
|
for card in cards:
|
|
for key in (k.strip().lower() for k in card.keys.split(",")):
|
|
if key and key in haystack:
|
|
matched.append(
|
|
{"id": card.id, "name": card.name, "keyword": key, "entry": card.entry}
|
|
)
|
|
break
|
|
return matched
|
|
|
|
|
|
def build_context(
|
|
adventure: models.Adventure,
|
|
settings: models.Settings,
|
|
memory_bank: dict | None = None,
|
|
exclude_action_id: int | None = None,
|
|
) -> tuple[str, str, dict]:
|
|
"""Returns (system_text, story_text, context_report). `memory_bank` is the
|
|
result of memorybank.retrieve_memories (None when the bank is off);
|
|
`exclude_action_id` omits one action from the story (see history.py)."""
|
|
script_mem = _script_memory(adventure)
|
|
|
|
# ----- The static block, which is identical on every turn -----
|
|
# This ordering exists to reduce cost. Prompt caching matches a prefix. The
|
|
# endpoint reuses the prompt up to the first byte that differs from the
|
|
# previous request, and no further. A section that changes near the top
|
|
# therefore re-prices everything below it, and what sits below it is the
|
|
# story history, which is most of the prompt. Sections that change from turn
|
|
# to turn go after the history, among the live sections. Placing them there
|
|
# also gives them the most recency, which is why `EMIT_REMINDER` goes last.
|
|
system_sections: list[Section] = [Section("narrator", settings.narrator_prompt.strip())]
|
|
|
|
# RPG world state (Phase 12): the instructions for reporting changes. The
|
|
# live values go into a live section below. The guide derived from the
|
|
# schema and the emit rule do not change while the scenario is unchanged.
|
|
stat_schema = adventure.scenario.stat_schema if adventure.scenario else None
|
|
has_ws = worldstate.has_schema(stat_schema)
|
|
persona_name = adventure.persona_name.strip()
|
|
# M5: the typed-event protocol replaces the delta rule for every campaign,
|
|
# with or without an inherited stat schema. State is no longer an opt-in
|
|
# RPG layer — a story has entities, places and possessions whatever genre it
|
|
# is, so the rule is unconditional.
|
|
system_sections.append(Section("state_rule", narrative.extract.EMIT_RULE))
|
|
canon_text = _canon_section(adventure)
|
|
if canon_text:
|
|
system_sections.append(Section("campaign_canon", canon_text))
|
|
|
|
if isinstance(script_mem.get("context"), str) and script_mem["context"].strip():
|
|
system_sections.append(Section("script_context", script_mem["context"].strip()))
|
|
if adventure.ai_instructions.strip():
|
|
system_sections.append(Section("ai_instructions", adventure.ai_instructions.strip()))
|
|
# Phase 18. This sits in the static block because only the user can edit it,
|
|
# so it never changes mid-story and stays inside the cached prefix. It is
|
|
# emitted whether or not the adventure has an RPG layer: an adventure with
|
|
# no stats still has a protagonist, and that is the case the persona was
|
|
# added for.
|
|
persona_text = render_persona(adventure)
|
|
if persona_text:
|
|
system_sections.append(Section("persona", persona_text))
|
|
if adventure.memory.strip():
|
|
system_sections.append(
|
|
Section("plot_essentials", f"Plot essentials:\n{adventure.memory.strip()}")
|
|
)
|
|
|
|
# ----- Live sections, which hold everything that changes -----
|
|
# This code builds them here and places them after the history further down.
|
|
# They are ordered from least to most volatile, so a turn that changes only
|
|
# the fastest-moving section leaves the others cached. The summary is
|
|
# rewritten every few turns. Lore changes with the scene. The retrieved
|
|
# memories change on most turns, and the stat values change on nearly every
|
|
# turn. `world_lore` is added below, because the history window determines
|
|
# which cards trigger and that window is not known yet.
|
|
# M6: the summary the *current lineage* is entitled to, not whatever was
|
|
# written last. A summary is derived data anchored to the story it covers,
|
|
# so an Undo or a divergence makes an old one ineligible rather than
|
|
# leaking it into a story it does not describe (E03, `app/summaries.py`).
|
|
db = object_session(adventure)
|
|
summary_row = summaries.current(db, adventure) if db is not None else None
|
|
summary_text = summary_row.text.strip() if summary_row is not None else ""
|
|
summary_section = (
|
|
Section("story_summary", f"Story summary:\n{summary_text}")
|
|
if summary_text
|
|
else None
|
|
)
|
|
memories_section = None
|
|
if memory_bank and memory_bank.get("used"):
|
|
# M6: an inference must not read as a record. A heuristic memory is
|
|
# marked in the prompt itself, because the narrator decides what to
|
|
# treat as established from what it is shown, and an unlabelled guess
|
|
# sitting beside accepted history is how a guess becomes canon
|
|
# (`CONTEXT-AND-MEMORY.md` §14). Authoritative state changes still come
|
|
# only from the M5 event path, whatever a memory says.
|
|
lines_text = "\n".join(_memory_line(m) for m in memory_bank["used"])
|
|
memories_section = Section(
|
|
"used_memories",
|
|
"Memories from earlier in the story. Lines marked [inferred] are "
|
|
"interpretation, not established fact — do not treat them as "
|
|
f"settled truth:\n{lines_text}",
|
|
)
|
|
world_state_section = None
|
|
refusal_note = ""
|
|
# M5: the authoritative narrative state, as the model is shown it. Read from
|
|
# the campaign's live document, which head movement keeps pointed at the
|
|
# position being read — so an undone story is described by the state it had
|
|
# then, not by the state it reached later.
|
|
state_block = narrative.render.for_prompt(adventure.narrative_state)
|
|
if state_block:
|
|
world_state_section = Section("narrative_state", state_block)
|
|
# Corrections for the previous AI turn only. A refusal the model has
|
|
# already had one chance to fix is stale, and repeating it every turn
|
|
# would price a correction into the whole rest of the adventure.
|
|
recent = history.tail(adventure, NPC_WINDOW, exclude_action_id)
|
|
last_ai = next((a for a in reversed(recent) if a.type == "ai"), None)
|
|
if last_ai is not None:
|
|
refusal_note = narrative.extract.render_rejections(last_ai.state_rejections)
|
|
|
|
authors_note_text = adventure.authors_note.strip()
|
|
if isinstance(script_mem.get("authorsNote"), str) and script_mem["authorsNote"].strip():
|
|
authors_note_text = script_mem["authorsNote"].strip()
|
|
authors_note = f"[Author's note: {authors_note_text}]" if authors_note_text else ""
|
|
|
|
front_memory = ""
|
|
if isinstance(script_mem.get("frontMemory"), str):
|
|
front_memory = script_mem["frontMemory"].strip()
|
|
|
|
length_note = length_hint(settings.max_output_tokens)
|
|
|
|
# The live sections sit below the history, but they are still part of the
|
|
# prompt, so they still count against the budget. `world_lore` is the
|
|
# exception, because the code below budgets it out of `available`.
|
|
reserved = (
|
|
sum(s.tokens for s in system_sections)
|
|
+ sum(
|
|
s.tokens
|
|
for s in (summary_section, memories_section, world_state_section)
|
|
if s is not None
|
|
)
|
|
+ count_tokens(authors_note)
|
|
+ count_tokens(front_memory)
|
|
+ count_tokens(length_note)
|
|
+ count_tokens(narrative.extract.EMIT_REMINDER)
|
|
+ count_tokens(refusal_note)
|
|
)
|
|
|
|
# ----- M6: the output reserve, and what happens when it does not fit -----
|
|
#
|
|
# `context_token_budget` is the whole window the model is given, so the
|
|
# narrator's reply has to be subtracted from it before any history is
|
|
# chosen. Until M6 it was not: the builder spent the entire budget on input
|
|
# and left the reply to fit in whatever the endpoint had left, which is a
|
|
# truncated turn on a model whose window is the budget
|
|
# (`CONTEXT-AND-MEMORY.md` §32, acceptance test F04).
|
|
#
|
|
# The margin covers what is added after this arithmetic — the separators
|
|
# between sections, and the difference between our tokenizer's count and the
|
|
# serving model's. It is small and fixed rather than proportional, because
|
|
# what it absorbs does not scale with the budget.
|
|
output_reserve = max(0, settings.max_output_tokens) + OUTPUT_SAFETY_MARGIN
|
|
protected = reserved + output_reserve
|
|
if protected >= settings.context_token_budget:
|
|
# Failing here is the point. The alternative — carrying on with a token
|
|
# or two of history — builds a prompt that is known to overflow, and
|
|
# the reader gets a truncated reply with no explanation. §32: "fail
|
|
# gracefully if protected context alone is too large."
|
|
raise ContextOverflow(
|
|
f"The protected context needs {protected} tokens "
|
|
f"({reserved} of prompt plus {output_reserve} reserved for the "
|
|
f"reply) but the context budget is {settings.context_token_budget}. "
|
|
"Raise the context budget, lower the maximum reply length, or "
|
|
"shorten the campaign's canon, instructions and persona."
|
|
)
|
|
available = settings.context_token_budget - protected
|
|
|
|
# Only the newest actions can reach the prompt, because the code below
|
|
# either truncates the text to `available` tokens or stops at the budget.
|
|
# Fetch a window that is provably larger than that and no larger. Otherwise
|
|
# a long adventure reads its whole history on every turn and uses only the
|
|
# end of it.
|
|
actions = history.window_covering(
|
|
adventure, available, count_tokens, exclude_action_id
|
|
)
|
|
|
|
# ----- Story cards: triggered by recent story text (the window history could fill) -----
|
|
trigger_window = truncate_to_last_tokens(SEPARATOR.join(a.text for a in actions), available)
|
|
triggered = match_cards(adventure.story_cards, trigger_window)
|
|
|
|
card_budget = int(available * CARD_BUDGET_SHARE)
|
|
card_records = []
|
|
lore_lines: list[str] = []
|
|
used = 0
|
|
for match in triggered:
|
|
line = f"World Lore: {match['entry'].strip()}"
|
|
tokens = count_tokens(line)
|
|
included = used + tokens <= card_budget
|
|
if included:
|
|
lore_lines.append(line)
|
|
used += tokens
|
|
card_records.append(
|
|
{"id": match["id"], "name": match["name"], "keyword": match["keyword"],
|
|
"included": included}
|
|
)
|
|
lore_section = (
|
|
Section("world_lore", "\n".join(lore_lines)) if lore_lines else None
|
|
)
|
|
|
|
# ----- Story history: newest first until the remaining budget is spent -----
|
|
history_budget = available - used
|
|
included_actions: list[models.Action] = []
|
|
spent = 0
|
|
oldest_truncated = False
|
|
for action in reversed(actions):
|
|
# Budget against the text as it appears in the prompt, which includes
|
|
# the state block when this adventure tracks world state.
|
|
rendered = _history_text(action)
|
|
tokens = count_tokens(rendered) + count_tokens(SEPARATOR)
|
|
if spent + tokens > history_budget:
|
|
if not included_actions:
|
|
# Even the newest action alone is over budget: hard-truncate it.
|
|
included_actions.append(
|
|
models.Action(
|
|
adventure_id=action.adventure_id,
|
|
type=action.type,
|
|
text=truncate_to_last_tokens(action.text, history_budget),
|
|
)
|
|
)
|
|
oldest_truncated = True
|
|
break
|
|
included_actions.append(action)
|
|
spent += tokens
|
|
included_actions.reverse()
|
|
|
|
# ----- Assemble the story text, with the author's note near the end -----
|
|
# Append each AI turn's state block again. The app strips it before storage,
|
|
# and the recent history has to show the model the pattern to follow.
|
|
texts = [_history_text(a) for a in included_actions]
|
|
note_sections: list[Section] = []
|
|
if authors_note:
|
|
pos = max(0, len(texts) - AUTHORS_NOTE_DEPTH)
|
|
before, after = texts[:pos], texts[pos:]
|
|
if before:
|
|
note_sections.append(Section("history", SEPARATOR.join(before)))
|
|
note_sections.append(Section("authors_note", authors_note))
|
|
note_sections.append(Section("recent_history", SEPARATOR.join(after)))
|
|
else:
|
|
note_sections.append(Section("history", SEPARATOR.join(texts)))
|
|
# The live sections, ordered from least to most volatile. See the comment
|
|
# where they are built. They go below the history so that the history stays
|
|
# cached, and above the final sections so that those stay last.
|
|
for live in (summary_section, lore_section, memories_section, world_state_section):
|
|
if live is not None:
|
|
note_sections.append(live)
|
|
if front_memory:
|
|
note_sections.append(Section("front_memory", front_memory))
|
|
# Place the length hint just above the emit reminder, which keeps the last
|
|
# position. The length budget applies to the narration, and the reminder
|
|
# applies to the block that follows it, so this is also the order in which
|
|
# the model acts.
|
|
note_sections.append(Section("length_hint", length_note))
|
|
# A correction for the previous turn sits directly above the reminder to
|
|
# emit a block, which is the instruction it modifies.
|
|
if refusal_note:
|
|
note_sections.append(Section("state_refusals", refusal_note))
|
|
# The emit rule sits in the system block, far from where the model
|
|
# generates text, so repeat it last where it has the most effect.
|
|
note_sections.append(Section("state_reminder", narrative.extract.EMIT_REMINDER))
|
|
|
|
story_sections = [s for s in note_sections if s.text]
|
|
system_text = SEPARATOR.join(s.text for s in system_sections if s.text)
|
|
story_text = SEPARATOR.join(s.text for s in story_sections)
|
|
|
|
all_sections = [s for s in system_sections if s.text] + story_sections
|
|
report = {
|
|
"sections": [
|
|
{"label": s.label, "text": s.text, "tokens": s.tokens} for s in all_sections
|
|
],
|
|
"prompt": {"system": system_text, "story": story_text},
|
|
# M6: the numbers the reader needs to answer "how much did each part
|
|
# cost, and what was left for the reply?" (F04, F05). `available` is
|
|
# what the history was actually allowed to spend after everything
|
|
# protected was subtracted.
|
|
"tokens": {
|
|
"total": count_tokens(system_text) + count_tokens(story_text),
|
|
"budget": settings.context_token_budget,
|
|
"output_reserve": output_reserve,
|
|
"protected": reserved,
|
|
"available_for_history": available,
|
|
"history_spent": spent,
|
|
},
|
|
"cards": card_records,
|
|
"memories": memory_bank,
|
|
# M6: which summary was used, and which stretch of story it covers, so
|
|
# "what history did that summary cover?" is answerable from the record
|
|
# rather than by guessing (F05, F06).
|
|
"summary": summaries.provenance(summary_row),
|
|
# M6: whether background derived work is currently failing for this
|
|
# campaign. A dead memory bank is visible here rather than only in a log
|
|
# nobody reads (F08).
|
|
"derived": derived.report(db, adventure.id) if db is not None else [],
|
|
"history": {
|
|
"included": len(included_actions),
|
|
# The count covers the whole story rather than the window fetched
|
|
# above. Insights reports how many of the total actions it
|
|
# included, so this number must be the real total.
|
|
"total": history.count(adventure, exclude_action_id),
|
|
"oldest_truncated": oldest_truncated,
|
|
},
|
|
"settings": {
|
|
"model": settings.model,
|
|
"api_mode": settings.api_mode,
|
|
"temperature": settings.temperature,
|
|
"max_output_tokens": settings.max_output_tokens,
|
|
},
|
|
}
|
|
return system_text, story_text, report
|