fix: surface the response body on 4xx so 403s are diagnosable

Media downloads logged "HTTP Error 403:" with no reason. That string is
curl_cffi's raise_for_status() format, "HTTP Error {code}: {reason}", and
HTTP/2 carries no reason phrase — so the message said nothing, and
_make_request threw the response body away. The provider's JSON `detail`
is the only explanation available for a refused asset.

- base._make_request: end non-retryable statuses with a ProviderError
  carrying the body's detail/error/message (redacted, truncated to 300
  chars) instead of a bare raise_for_status().
- media: bucket 403 as `forbidden` in the run summary, separately from
  `download-error` — "the asset is gone" and "we were refused" are
  different problems.
- utils.redact_secrets: match secret key names per word. Exact matching
  let access_token, api_key, and session-token through into logged
  bodies; "keywords"/"monkey"/"tokenizer" stay intact.
- tests/test_config.py: test_defaults depended on the absence of a local
  .env — load_config() calls load_dotenv(override=False), which restored
  the variable the test had just deleted. Stub dotenv discovery.

305 tests pass.
This commit is contained in:
JesseMarkowitz
2026-08-17 07:58:02 -04:00
parent 1f5a445ada
commit 395ea19ca8
8 changed files with 193 additions and 8 deletions
+18
View File
@@ -145,3 +145,21 @@ class TestFormatTokenStatus:
expiry = datetime.now(tz=timezone.utc) + timedelta(days=10, hours=12)
result = format_token_status("tok", expiry)
assert "10 days" in result
class TestRedactCompoundKeys:
"""Exact-match redaction let compound secret names through into logs."""
def test_compound_secret_keys_redacted(self):
result = redact_secrets(
{"access_token": "sk-abc", "api_key": "k1", "session-token": "s1"}
)
assert result == {
"access_token": "[REDACTED]",
"api_key": "[REDACTED]",
"session-token": "[REDACTED]",
}
def test_innocent_keys_containing_a_secret_word_kept(self):
result = redact_secrets({"keywords": ["a"], "monkey": "b", "tokenizer": "c"})
assert result == {"keywords": ["a"], "monkey": "b", "tokenizer": "c"}