Files
interactive-story/backend/tests/test_provider_wiring.py
T
JesseMarkowitzandClaude Opus 5 a6e9c7a32b M6: branch-safe context, summaries and long-term story memory
Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.

This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.

What was already correct, and was kept rather than rebuilt

  Memory lineage. Memories already carried (branch_id, depth) and retrieval
  already filtered through the capped-path clause; the ten-step negative control
  was measured passing against b7005e6 before any change here. M6 adds the
  regression tests that pin it, plus provenance and authority on the result.

Summary lineage — both halves

  A summary is a row carrying the coordinate of the last node it covers, and
  eligibility is the same head-capped lineage clause memories use. That alone
  was not enough: generation was seeded from adventures.story_summary, a
  campaign-global column with no lineage, so after a divergence the summariser
  was handed the abandoned line's prose and asked to update it. The row it
  produced was correctly anchored and therefore looked safe while its sentences
  described a story the reader had left.

  Generation is now seeded from summaries.current — the same question the
  context builder asks — so the input and the output are scoped by one rule.
  adventures.story_summary remains a reader-facing mirror for the Plot panel and
  the export bundle, kept in step when a summary is written and when the head
  moves, and nothing authoritative reads it.

Retrieval redundancy

  With a real embedding model, four near-identical memories crowded out the one
  distinctive clue, which survived only because the default memory_top_k is 5.
  Retrieval now drops a candidate that repeats one already chosen, never across
  authority classes, at a threshold measured against the configured embedding
  model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
  the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
  and are recorded as such.

Memory authority, budgeting, observability

  Memory.authority is accepted_story or heuristic, classified by the application
  and marked in the prompt; retrieval never writes state. The reply is reserved
  out of the context budget, and an impossible configuration fails clearly
  instead of overflowing. Each derived pass records ok/idle/failed per campaign,
  served by GET /adventures/{id}/derived and shown in Insights, so the M2
  failure — a dead memory bank with a green suite — is visible if it recurs.
  Provider-wiring tests mock no factory.

Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.

Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
2026-09-06 03:00:33 -04:00

226 lines
8.4 KiB
Python

"""M6: the provider construction path, exercised for real.
M2 shipped with the entire memory bank dead and the full suite green. The
summariser and the embedder were built from `Settings` attributes that had moved,
the resulting `AttributeError` was raised inside a fire-and-forget task, and
every memory test had stubbed the factories out — so nothing anywhere noticed
(`BUILD-MILESTONES.md`, note from M2).
These tests exist so that cannot happen twice. **Nothing here mocks a provider
factory.** They call the real factories with a real `Settings` row read back out
of the database, and assert that the configured values arrive at the object that
consumes them. A renamed or removed column fails here loudly instead of killing
the memory bank quietly.
Network is never touched: constructing a provider makes no request. The one test
that would make one is skipped unless a trusted-LAN endpoint is configured, and
it is reported separately from these.
python -m pytest tests/test_provider_wiring.py -v
"""
import os
import pytest
from app import endpoints, memorybank, models
from app.database import Base, SessionLocal, engine
from app.providers import OpenAICompatibleProvider, ProviderError
@pytest.fixture()
def settings():
"""A real Settings row, round-tripped through the database.
Round-tripping matters: a column that was renamed in the model but still
referenced by a factory fails on the read, which is the failure this module
is here to produce.
"""
Base.metadata.create_all(bind=engine)
db = SessionLocal()
user = models.User(is_guest=False, email="wiring@example.com")
db.add(user)
db.flush()
row = models.Settings(
user_id=user.id,
api_key="enc:dummy",
endpoint_url="http://127.0.0.1:11434/v1",
model="narrator-model",
summary_model="summariser-model",
embedding_model="embedding-model",
api_mode="chat",
model_timeout_seconds=123,
max_output_tokens=456,
context_token_budget=4096,
)
db.add(row)
db.commit()
row_id = row.id
db.close()
db = SessionLocal()
try:
yield db.get(models.Settings, row_id)
finally:
db.close()
Base.metadata.drop_all(bind=engine)
# ------------------------------------------------- the real construction path
def test_the_summary_provider_is_built_from_the_configured_values(settings):
"""Every value the summariser needs reaches the provider that uses it."""
provider = memorybank.summary_provider(settings)
assert isinstance(provider, OpenAICompatibleProvider)
assert provider.base_url == "http://127.0.0.1:11434/v1"
assert provider.model == "summariser-model"
assert provider.api_mode == "chat"
assert provider.read_timeout == 123
def test_the_summary_provider_falls_back_to_the_narrator_model(settings):
"""An empty summary model means "use the main one", not "use nothing"."""
settings.summary_model = ""
assert memorybank.summary_provider(settings).model == "narrator-model"
def test_the_embedding_provider_is_built_from_the_configured_values(settings):
provider = memorybank.embedding_provider(settings)
assert isinstance(provider, OpenAICompatibleProvider)
assert provider.base_url == "http://127.0.0.1:11434/v1"
assert provider.model == "embedding-model"
def test_the_narrator_provider_is_built_from_the_configured_values(settings):
"""The turn path builds its own provider; the same values have to reach it."""
provider = OpenAICompatibleProvider(
settings.endpoint_url, settings.model, settings.api_mode,
settings.model_timeout_seconds,
)
assert provider.base_url == "http://127.0.0.1:11434/v1"
assert provider.model == "narrator-model"
assert provider.api_mode == "chat"
assert provider.read_timeout == 123
@pytest.mark.parametrize("attribute", [
"endpoint_url", "model", "summary_model", "embedding_model", "api_mode",
"model_timeout_seconds", "max_output_tokens", "context_token_budget",
])
def test_every_settings_attribute_the_providers_read_still_exists(settings, attribute):
"""The named guard against M2's failure.
Each attribute here is one a factory or the context builder reads. If a
migration renames one, this fails by name instead of the memory bank dying
in a task nobody is watching.
"""
assert hasattr(settings, attribute), (
f"Settings.{attribute} is gone; something that builds a provider reads it"
)
def test_building_a_provider_makes_no_request(settings):
"""Construction is inert, so these tests are safe to run offline."""
import socket
def refuse(*args, **kwargs): # pragma: no cover - only runs on a failure
raise AssertionError("provider construction opened a socket")
real = socket.socket.connect
socket.socket.connect = refuse
try:
memorybank.summary_provider(settings)
memorybank.embedding_provider(settings)
finally:
socket.socket.connect = real
# ------------------------------------------------------- the endpoint policy
def test_the_embedding_path_enforces_the_same_endpoint_policy(settings):
"""M6 section 18. Embedding inputs are story text, and they go to the same
kind of endpoint under the same rule as a narrator prompt.
Asserted through the real `embed()` rather than by reading the source: a
check that exists but is not reached would pass a source inspection.
"""
import asyncio
provider = OpenAICompatibleProvider("https://api.openai.com/v1", "embedding-model")
with pytest.raises(ProviderError) as exc:
asyncio.run(provider.embed(["a line of someone's story"]))
assert "can't be used" in str(exc.value)
def test_the_policy_refuses_a_public_address_for_embeddings():
"""The rule is the address, not the name of the caller."""
assert endpoints.rejection_reason("https://api.openai.com/v1/embeddings")
assert endpoints.rejection_reason("http://8.8.8.8:11434/v1/embeddings")
# And permits the local endpoints the product is built for.
assert endpoints.rejection_reason("http://127.0.0.1:11434/v1/embeddings") is None
def test_no_cloud_or_remote_vector_service_is_configured_anywhere():
"""M6 section 18: no new network path. Checked against the source, because
the point is that no such code exists to be exercised."""
import pathlib
forbidden = (
"api.openai.com", "api.anthropic.com", "pinecone", "weaviate",
"qdrant", "chromadb", "cohere.ai", "huggingface.co/api",
)
root = pathlib.Path(__file__).resolve().parent.parent / "app"
offenders = []
for path in root.rglob("*.py"):
text = path.read_text()
for needle in forbidden:
# `endpoints.py` names cloud hosts in order to refuse them.
if needle in text and path.name != "endpoints.py":
offenders.append(f"{path.name}: {needle}")
assert not offenders, offenders
# ------------------------------------------------- the live endpoint, if any
@pytest.mark.skipif(
not os.environ.get("AIDND_TEST_ENDPOINT"),
reason="set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL to run against a real model",
)
def test_the_real_construction_path_reaches_a_real_endpoint():
"""The wiring test with the network attached.
Reported separately from the tests above: this one proves the constructed
provider can actually talk to the configured endpoint, which is the half a
unit test cannot show. It uses the ordinary endpoint policy — no TLS
weakening, no allowlist bypass.
"""
import asyncio
endpoint = os.environ["AIDND_TEST_ENDPOINT"]
model = os.environ.get("AIDND_TEST_EMBED_MODEL", "nomic-embed-text")
assert endpoints.rejection_reason(endpoint) is None, (
"the configured test endpoint is refused by the policy"
)
Base.metadata.create_all(bind=engine)
db = SessionLocal()
try:
user = models.User(is_guest=False, email="live@example.com")
db.add(user)
db.flush()
row = models.Settings(
user_id=user.id, api_key="enc:dummy", endpoint_url=endpoint,
model=os.environ.get("AIDND_TEST_MODEL", ""), embedding_model=model,
)
db.add(row)
db.commit()
provider = memorybank.embedding_provider(row)
vectors = asyncio.run(provider.embed(["Aldric hid the ledger."]))
finally:
db.close()
Base.metadata.drop_all(bind=engine)
assert len(vectors) == 1
assert len(vectors[0]) > 8, "the endpoint returned no usable vector"