Aligns the inherited AI-DnD memory and context foundation with the history,
authority and state model M3-M5 established. Long stories now reach the narrator
through a bounded, lineage-safe, inspectable context rather than a growing
transcript.
This commit includes the corrective work that followed the independent review in
planning/reports/M6-IMPLEMENTATION-REPORT.md. The first implementation reported
E03 as passing and it was not; the report records that history rather than
hiding it.
What was already correct, and was kept rather than rebuilt
Memory lineage. Memories already carried (branch_id, depth) and retrieval
already filtered through the capped-path clause; the ten-step negative control
was measured passing against b7005e6 before any change here. M6 adds the
regression tests that pin it, plus provenance and authority on the result.
Summary lineage — both halves
A summary is a row carrying the coordinate of the last node it covers, and
eligibility is the same head-capped lineage clause memories use. That alone
was not enough: generation was seeded from adventures.story_summary, a
campaign-global column with no lineage, so after a divergence the summariser
was handed the abandoned line's prose and asked to update it. The row it
produced was correctly anchored and therefore looked safe while its sentences
described a story the reader had left.
Generation is now seeded from summaries.current — the same question the
context builder asks — so the input and the output are scoped by one rule.
adventures.story_summary remains a reader-facing mirror for the Plot panel and
the export bundle, kept in step when a summary is written and when the head
moves, and nothing authoritative reads it.
Retrieval redundancy
With a real embedding model, four near-identical memories crowded out the one
distinctive clue, which survived only because the default memory_top_k is 5.
Retrieval now drops a candidate that repeats one already chosen, never across
authority classes, at a threshold measured against the configured embedding
model. The clue is retrieved at top_k 5, 4 and 3. Ranking itself is unchanged;
the further factors CONTEXT-AND-MEMORY §20 contemplates remain unimplemented
and are recorded as such.
Memory authority, budgeting, observability
Memory.authority is accepted_story or heuristic, classified by the application
and marked in the prompt; retrieval never writes state. The reply is reserved
out of the context budget, and an impossible configuration fails clearly
instead of overflowing. Each derived pass records ok/idle/failed per campaign,
served by GET /adventures/{id}/derived and shown in Insights, so the M2
failure — a dead memory bank with a green suite — is visible if it recurs.
Provider-wiring tests mock no factory.
Also: two pre-existing test-suite leaks fixed; two fixtures that stored one
vector in every memory now use distinct ones, so lineage assertions stay
readable alongside redundancy suppression.
Planning: CONTEXT-AND-MEMORY, TECHNICAL-DESIGN, DATA-MODEL, V1-ACCEPTANCE-TESTS,
BUILD-MILESTONES, VERSION and planning/README updated to describe what exists,
including that a valid E03 test must regenerate a summary after diverging. The
M5 report was rotated to planning/archive/milestone-reports/. No new ADR — every
choice implements a decision the package had already settled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PWU4gTfLYY6Qq9U7aa9Qw2
226 lines
8.4 KiB
Python
226 lines
8.4 KiB
Python
"""M6: the provider construction path, exercised for real.
|
|
|
|
M2 shipped with the entire memory bank dead and the full suite green. The
|
|
summariser and the embedder were built from `Settings` attributes that had moved,
|
|
the resulting `AttributeError` was raised inside a fire-and-forget task, and
|
|
every memory test had stubbed the factories out — so nothing anywhere noticed
|
|
(`BUILD-MILESTONES.md`, note from M2).
|
|
|
|
These tests exist so that cannot happen twice. **Nothing here mocks a provider
|
|
factory.** They call the real factories with a real `Settings` row read back out
|
|
of the database, and assert that the configured values arrive at the object that
|
|
consumes them. A renamed or removed column fails here loudly instead of killing
|
|
the memory bank quietly.
|
|
|
|
Network is never touched: constructing a provider makes no request. The one test
|
|
that would make one is skipped unless a trusted-LAN endpoint is configured, and
|
|
it is reported separately from these.
|
|
|
|
python -m pytest tests/test_provider_wiring.py -v
|
|
"""
|
|
|
|
import os
|
|
|
|
import pytest
|
|
|
|
from app import endpoints, memorybank, models
|
|
from app.database import Base, SessionLocal, engine
|
|
from app.providers import OpenAICompatibleProvider, ProviderError
|
|
|
|
|
|
@pytest.fixture()
|
|
def settings():
|
|
"""A real Settings row, round-tripped through the database.
|
|
|
|
Round-tripping matters: a column that was renamed in the model but still
|
|
referenced by a factory fails on the read, which is the failure this module
|
|
is here to produce.
|
|
"""
|
|
Base.metadata.create_all(bind=engine)
|
|
db = SessionLocal()
|
|
user = models.User(is_guest=False, email="wiring@example.com")
|
|
db.add(user)
|
|
db.flush()
|
|
row = models.Settings(
|
|
user_id=user.id,
|
|
api_key="enc:dummy",
|
|
endpoint_url="http://127.0.0.1:11434/v1",
|
|
model="narrator-model",
|
|
summary_model="summariser-model",
|
|
embedding_model="embedding-model",
|
|
api_mode="chat",
|
|
model_timeout_seconds=123,
|
|
max_output_tokens=456,
|
|
context_token_budget=4096,
|
|
)
|
|
db.add(row)
|
|
db.commit()
|
|
row_id = row.id
|
|
db.close()
|
|
|
|
db = SessionLocal()
|
|
try:
|
|
yield db.get(models.Settings, row_id)
|
|
finally:
|
|
db.close()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
|
|
# ------------------------------------------------- the real construction path
|
|
|
|
def test_the_summary_provider_is_built_from_the_configured_values(settings):
|
|
"""Every value the summariser needs reaches the provider that uses it."""
|
|
provider = memorybank.summary_provider(settings)
|
|
|
|
assert isinstance(provider, OpenAICompatibleProvider)
|
|
assert provider.base_url == "http://127.0.0.1:11434/v1"
|
|
assert provider.model == "summariser-model"
|
|
assert provider.api_mode == "chat"
|
|
assert provider.read_timeout == 123
|
|
|
|
|
|
def test_the_summary_provider_falls_back_to_the_narrator_model(settings):
|
|
"""An empty summary model means "use the main one", not "use nothing"."""
|
|
settings.summary_model = ""
|
|
assert memorybank.summary_provider(settings).model == "narrator-model"
|
|
|
|
|
|
def test_the_embedding_provider_is_built_from_the_configured_values(settings):
|
|
provider = memorybank.embedding_provider(settings)
|
|
|
|
assert isinstance(provider, OpenAICompatibleProvider)
|
|
assert provider.base_url == "http://127.0.0.1:11434/v1"
|
|
assert provider.model == "embedding-model"
|
|
|
|
|
|
def test_the_narrator_provider_is_built_from_the_configured_values(settings):
|
|
"""The turn path builds its own provider; the same values have to reach it."""
|
|
provider = OpenAICompatibleProvider(
|
|
settings.endpoint_url, settings.model, settings.api_mode,
|
|
settings.model_timeout_seconds,
|
|
)
|
|
assert provider.base_url == "http://127.0.0.1:11434/v1"
|
|
assert provider.model == "narrator-model"
|
|
assert provider.api_mode == "chat"
|
|
assert provider.read_timeout == 123
|
|
|
|
|
|
@pytest.mark.parametrize("attribute", [
|
|
"endpoint_url", "model", "summary_model", "embedding_model", "api_mode",
|
|
"model_timeout_seconds", "max_output_tokens", "context_token_budget",
|
|
])
|
|
def test_every_settings_attribute_the_providers_read_still_exists(settings, attribute):
|
|
"""The named guard against M2's failure.
|
|
|
|
Each attribute here is one a factory or the context builder reads. If a
|
|
migration renames one, this fails by name instead of the memory bank dying
|
|
in a task nobody is watching.
|
|
"""
|
|
assert hasattr(settings, attribute), (
|
|
f"Settings.{attribute} is gone; something that builds a provider reads it"
|
|
)
|
|
|
|
|
|
def test_building_a_provider_makes_no_request(settings):
|
|
"""Construction is inert, so these tests are safe to run offline."""
|
|
import socket
|
|
|
|
def refuse(*args, **kwargs): # pragma: no cover - only runs on a failure
|
|
raise AssertionError("provider construction opened a socket")
|
|
|
|
real = socket.socket.connect
|
|
socket.socket.connect = refuse
|
|
try:
|
|
memorybank.summary_provider(settings)
|
|
memorybank.embedding_provider(settings)
|
|
finally:
|
|
socket.socket.connect = real
|
|
|
|
|
|
# ------------------------------------------------------- the endpoint policy
|
|
|
|
def test_the_embedding_path_enforces_the_same_endpoint_policy(settings):
|
|
"""M6 section 18. Embedding inputs are story text, and they go to the same
|
|
kind of endpoint under the same rule as a narrator prompt.
|
|
|
|
Asserted through the real `embed()` rather than by reading the source: a
|
|
check that exists but is not reached would pass a source inspection.
|
|
"""
|
|
import asyncio
|
|
|
|
provider = OpenAICompatibleProvider("https://api.openai.com/v1", "embedding-model")
|
|
with pytest.raises(ProviderError) as exc:
|
|
asyncio.run(provider.embed(["a line of someone's story"]))
|
|
assert "can't be used" in str(exc.value)
|
|
|
|
|
|
def test_the_policy_refuses_a_public_address_for_embeddings():
|
|
"""The rule is the address, not the name of the caller."""
|
|
assert endpoints.rejection_reason("https://api.openai.com/v1/embeddings")
|
|
assert endpoints.rejection_reason("http://8.8.8.8:11434/v1/embeddings")
|
|
# And permits the local endpoints the product is built for.
|
|
assert endpoints.rejection_reason("http://127.0.0.1:11434/v1/embeddings") is None
|
|
|
|
|
|
def test_no_cloud_or_remote_vector_service_is_configured_anywhere():
|
|
"""M6 section 18: no new network path. Checked against the source, because
|
|
the point is that no such code exists to be exercised."""
|
|
import pathlib
|
|
|
|
forbidden = (
|
|
"api.openai.com", "api.anthropic.com", "pinecone", "weaviate",
|
|
"qdrant", "chromadb", "cohere.ai", "huggingface.co/api",
|
|
)
|
|
root = pathlib.Path(__file__).resolve().parent.parent / "app"
|
|
offenders = []
|
|
for path in root.rglob("*.py"):
|
|
text = path.read_text()
|
|
for needle in forbidden:
|
|
# `endpoints.py` names cloud hosts in order to refuse them.
|
|
if needle in text and path.name != "endpoints.py":
|
|
offenders.append(f"{path.name}: {needle}")
|
|
assert not offenders, offenders
|
|
|
|
|
|
# ------------------------------------------------- the live endpoint, if any
|
|
|
|
@pytest.mark.skipif(
|
|
not os.environ.get("AIDND_TEST_ENDPOINT"),
|
|
reason="set AIDND_TEST_ENDPOINT and AIDND_TEST_MODEL to run against a real model",
|
|
)
|
|
def test_the_real_construction_path_reaches_a_real_endpoint():
|
|
"""The wiring test with the network attached.
|
|
|
|
Reported separately from the tests above: this one proves the constructed
|
|
provider can actually talk to the configured endpoint, which is the half a
|
|
unit test cannot show. It uses the ordinary endpoint policy — no TLS
|
|
weakening, no allowlist bypass.
|
|
"""
|
|
import asyncio
|
|
|
|
endpoint = os.environ["AIDND_TEST_ENDPOINT"]
|
|
model = os.environ.get("AIDND_TEST_EMBED_MODEL", "nomic-embed-text")
|
|
assert endpoints.rejection_reason(endpoint) is None, (
|
|
"the configured test endpoint is refused by the policy"
|
|
)
|
|
Base.metadata.create_all(bind=engine)
|
|
db = SessionLocal()
|
|
try:
|
|
user = models.User(is_guest=False, email="live@example.com")
|
|
db.add(user)
|
|
db.flush()
|
|
row = models.Settings(
|
|
user_id=user.id, api_key="enc:dummy", endpoint_url=endpoint,
|
|
model=os.environ.get("AIDND_TEST_MODEL", ""), embedding_model=model,
|
|
)
|
|
db.add(row)
|
|
db.commit()
|
|
provider = memorybank.embedding_provider(row)
|
|
vectors = asyncio.run(provider.embed(["Aldric hid the ledger."]))
|
|
finally:
|
|
db.close()
|
|
Base.metadata.drop_all(bind=engine)
|
|
|
|
assert len(vectors) == 1
|
|
assert len(vectors[0]) > 8, "the endpoint returned no usable vector"
|