Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
227 lines
9.4 KiB
Python
227 lines
9.4 KiB
Python
"""Prompt caching: the prompt has to start with the same bytes every turn.
|
|
|
|
Every endpoint that caches prompts caches a prefix. It reuses the request
|
|
up to the first byte that differs from last time, and no further. The cost
|
|
of a turn is therefore decided by layout: one section that changes each
|
|
turn, placed near the top, re-prices everything underneath it, and
|
|
underneath it is the story history, which makes up most of the prompt.
|
|
|
|
Three things must hold, and each is easy to undo by accident:
|
|
|
|
* The static block is byte-identical across turns. Adding a section that
|
|
moves (live stats, retrieved memories, a rewritten summary) to
|
|
`system_sections` is the mistake this file exists to catch.
|
|
* The sections that move sit after the history, but still before the tail
|
|
that is last for its own reasons: front memory, the length hint, and
|
|
`EMIT_REMINDER`, which is what keeps the state block emitted at all.
|
|
* Moving a section out of the system block does not drop it from the
|
|
token budget. It is still in the prompt.
|
|
|
|
This file also covers one request-level concern: reading back the usage
|
|
the endpoint reports, so the hit rate is measurable rather than assumed.
|
|
It covered OpenRouter upstream pinning too, until M2 removed cloud
|
|
provider support.
|
|
|
|
python -m pytest tests/test_prompt_caching.py -v
|
|
"""
|
|
import os
|
|
|
|
|
|
import pytest
|
|
|
|
from app import models, summaries, worldstate
|
|
from app import narrative
|
|
from app.context import builder
|
|
from app.database import Base, SessionLocal, engine
|
|
from app.providers.openai_compatible import OpenAICompatibleProvider
|
|
|
|
SCHEMA = {
|
|
"player": {"hp": {"min": 0, "max": 100, "initial": 100, "desc": "Health"}},
|
|
"world": {"alarm": {"min": 0, "max": 10, "initial": 0, "desc": "Alarm level"}},
|
|
}
|
|
|
|
|
|
def test_usage_is_recorded_from_a_final_chunk():
|
|
"""In a stream, the usage block arrives in a final chunk that carries
|
|
no choices, which is why it is read separately from the text
|
|
extraction."""
|
|
provider = OpenAICompatibleProvider("http://127.0.0.1:11434/v1", "m")
|
|
assert provider.last_usage is None
|
|
provider._record_usage({"choices": [{"delta": {"content": "hi"}}]})
|
|
assert provider.last_usage is None
|
|
provider._record_usage(
|
|
{"choices": [], "usage": {"prompt_tokens": 900,
|
|
"prompt_tokens_details": {"cached_tokens": 768}}}
|
|
)
|
|
assert provider.last_usage["prompt_tokens_details"]["cached_tokens"] == 768
|
|
|
|
|
|
def test_a_later_chunk_without_usage_does_not_erase_it():
|
|
provider = OpenAICompatibleProvider("http://127.0.0.1:11434/v1", "m")
|
|
provider._record_usage({"usage": {"prompt_tokens": 5}})
|
|
provider._record_usage({"choices": [{"delta": {"content": "x"}}]})
|
|
provider._record_usage({"usage": {}})
|
|
assert provider.last_usage == {"prompt_tokens": 5}
|
|
|
|
|
|
# ------------------------------------------------------------ prompt layout
|
|
|
|
NARRATIVE = {
|
|
"version": 1, "entities": {}, "possessions": {}, "threads": {},
|
|
"relationships": [], "scene": {},
|
|
"facts": [{"id": "f1", "predicate": "the lantern is lit", "status": "active"}],
|
|
}
|
|
|
|
|
|
def _with_hp(world_state, hp):
|
|
"""`world_state` is nested by group, and the JSON column only detects a
|
|
whole new object. Build a new one instead of mutating in place."""
|
|
return {**world_state, "player": {**world_state["player"], "hp": hp}}
|
|
|
|
|
|
@pytest.fixture()
|
|
def story():
|
|
Base.metadata.create_all(bind=engine)
|
|
db = SessionLocal()
|
|
user = models.User(is_guest=False, email="cache@example.com")
|
|
db.add(user)
|
|
db.flush()
|
|
settings = models.Settings(user_id=user.id, api_key="enc:dummy", model="m")
|
|
db.add(settings)
|
|
scenario = models.Scenario(
|
|
user_id=user.id, title="S", prompt="A road.", stat_schema=SCHEMA
|
|
)
|
|
db.add(scenario)
|
|
db.flush()
|
|
adventure = models.Adventure(
|
|
user_id=user.id, title="A", scenario_id=scenario.id, script_state={},
|
|
memory="The hero hunts bandits.",
|
|
ai_instructions="Write in second person.",
|
|
world_state=worldstate.instantiate(SCHEMA),
|
|
narrative_state=NARRATIVE,
|
|
# Phase 18. Set here so that every test in this file runs with a
|
|
# persona present: it is user-only, so it belongs in the static block,
|
|
# and this is the file that guards what may live there.
|
|
persona_name="Kaelen",
|
|
persona_pronouns="he/him",
|
|
persona_desc="A half-elf ranger.",
|
|
)
|
|
db.add(adventure)
|
|
db.flush()
|
|
for i in range(6):
|
|
db.add(models.Action(adventure_id=adventure.id,
|
|
type="ai" if i % 2 else "do",
|
|
text=f"[{i}] The road bends onward past the treeline."))
|
|
db.flush()
|
|
# M6: the summary is a row anchored to the story it covers, not a column.
|
|
# `build_context` reads whichever summary is eligible for the current head,
|
|
# so a test that wants one in the prompt has to record one.
|
|
summaries.record(db, adventure, "The hero left the village.")
|
|
db.commit()
|
|
db.expire_all()
|
|
adventure = db.get(models.Adventure, adventure.id)
|
|
settings = db.get(models.Settings, settings.id)
|
|
try:
|
|
yield db, adventure, settings
|
|
finally:
|
|
db.close()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
|
|
def test_changing_a_stat_leaves_the_static_block_untouched(story):
|
|
"""The whole point. Live values used to sit third from the top, so a
|
|
single point of damage re-priced the instructions, the plot, and the
|
|
history."""
|
|
db, adventure, settings = story
|
|
before, _, _ = builder.build_context(adventure, settings)
|
|
# M5 replaced the RPG stat block with the narrative-state block; the caching
|
|
# property is unchanged, and so is the test — moving the live state must not
|
|
# move the static prefix.
|
|
adventure.narrative_state = {**NARRATIVE, "facts": [
|
|
{"id": "f1", "predicate": "the lantern has gone out", "status": "active"}]}
|
|
db.commit()
|
|
after, story_text, _ = builder.build_context(adventure, settings)
|
|
assert before == after
|
|
assert "the lantern has gone out" in story_text, \
|
|
"the new value still has to reach the model"
|
|
|
|
|
|
def test_the_static_block_holds_the_things_that_do_not_move(story):
|
|
db, adventure, settings = story
|
|
system_text, story_text, _ = builder.build_context(adventure, settings)
|
|
for fixed in ("Write in second person.", "The hero hunts bandits.",
|
|
"You are Kaelen (he/him).", "A half-elf ranger."):
|
|
assert fixed in system_text
|
|
# The event vocabulary is derived from the allowlist, so it is fixed and
|
|
# belongs in the cached prefix. The state it describes is not fixed, and
|
|
# belongs to the story text.
|
|
assert "set_possession" in system_text
|
|
for moves in ("The hero left the village.", "the lantern is lit"):
|
|
assert moves not in system_text
|
|
assert moves in story_text
|
|
|
|
|
|
def test_volatile_sections_sit_after_the_history(story):
|
|
db, adventure, settings = story
|
|
_, story_text, _ = builder.build_context(adventure, settings)
|
|
history_at = story_text.index("[5] The road bends")
|
|
for label in ("Story summary:", "Established:"):
|
|
assert story_text.index(label) > history_at, label
|
|
|
|
|
|
def test_the_tail_stays_the_tail(story):
|
|
"""Front memory, the length hint, and `EMIT_REMINDER` are last for
|
|
reasons of their own, and the live sections must not have displaced
|
|
them."""
|
|
db, adventure, settings = story
|
|
_, story_text, report = builder.build_context(adventure, settings)
|
|
labels = [s["label"] for s in report["sections"]]
|
|
assert labels[-1] == "state_reminder"
|
|
assert labels[-2] == "length_hint"
|
|
assert labels.index("narrative_state") < labels.index("length_hint")
|
|
assert story_text.rstrip().endswith(narrative.extract.EMIT_REMINDER.rstrip())
|
|
|
|
|
|
def test_live_sections_are_still_charged_to_the_budget(story):
|
|
"""They moved out of `system_sections`, so it would be easy to stop
|
|
counting them in `reserved`. If that happened, the history, which is
|
|
budgeted with what is left over, would quietly overrun."""
|
|
db, adventure, settings = story
|
|
for i in range(6, 90):
|
|
db.add(models.Action(
|
|
adventure_id=adventure.id, type="do",
|
|
text=f"[{i}] " + "The road bends onward past the treeline. " * 6,
|
|
))
|
|
settings.context_token_budget = 4000
|
|
db.commit()
|
|
db.expire_all()
|
|
adventure = db.get(models.Adventure, adventure.id)
|
|
settings = db.get(models.Settings, settings.id)
|
|
|
|
_, _, lean = builder.build_context(adventure, settings)
|
|
summaries.record(db, adventure, "The hero left the village. " * 150)
|
|
db.commit()
|
|
_, _, fat = builder.build_context(adventure, settings)
|
|
|
|
assert fat["history"]["included"] < lean["history"]["included"], (
|
|
"a bigger summary has to leave less room for history"
|
|
)
|
|
|
|
|
|
def test_a_new_turn_only_appends_to_the_cached_prefix(story):
|
|
"""Playing on must extend the previous prompt, not rewrite it: the shared
|
|
prefix has to still contain the whole static block and the older history."""
|
|
db, adventure, settings = story
|
|
system_a, story_a, _ = builder.build_context(adventure, settings)
|
|
db.add(models.Action(adventure_id=adventure.id, type="do",
|
|
text="[6] You step into the clearing."))
|
|
db.commit()
|
|
db.expire_all()
|
|
adventure = db.get(models.Adventure, adventure.id)
|
|
system_b, story_b, _ = builder.build_context(adventure, settings)
|
|
|
|
assert system_a == system_b
|
|
shared = os.path.commonprefix([story_a, story_b])
|
|
assert "[0] The road bends" in shared
|
|
assert "[5] The road bends" in shared
|