tools: analyze whether media loss is age-based

"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
(![source](media/…) when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.

Offline, no token, no API calls.
This commit is contained in:
JesseMarkowitz
2026-08-17 08:45:21 -04:00
parent f40b25001a
commit 04191eed8c
6 changed files with 326 additions and 215 deletions
+7 -1
View File
@@ -156,9 +156,15 @@ def _classify_failure(error: ProviderError) -> str:
The buckets separate "the asset is gone" from "the asset is there but we
were refused" — different causes, different fixes, so lumping both into
download-error hides which one you have.
Note that a bare 403 from ChatGPT usually means *gone*, not *refused*:
its download endpoint reports a deleted upload as 403 Forbidden. The
provider confirms that against /files/{id} before raising, so a failure
that reaches here still saying "forbidden" is one where the asset really
does still exist — worth looking at, unlike an expired upload.
"""
detail = str(error.original).lower()
if "404" in detail or "not found" in detail:
if "404" in detail or "not found" in detail or "no longer exists" in detail:
return "expired-or-missing"
if "403" in detail or "forbidden" in detail:
return "forbidden"
+44 -1
View File
@@ -687,6 +687,15 @@ class ChatGPTProvider(BaseProvider):
signed ``download_url``; fetching that yields the bytes. Expired
assets (e.g. old generated images) 404 on the first hop.
A *deleted* asset does not 404 on the first hop — it answers
``403 {"detail":"Forbidden"}``, which reads like a permissions
problem and isn't one. Measured 2026-08-17 across 18 such assets:
every one returned ``404 {"detail":"File not found"}`` on
``/files/{id}`` while assets that downloaded fine returned 200 on
both, in the same session. So a 403 here is confirmed against the
metadata endpoint before it is reported, and a missing record is
called what it is rather than a refusal.
Returns:
(content_bytes, mime_type_or_None, file_name_or_None)
@@ -702,7 +711,21 @@ class ChatGPTProvider(BaseProvider):
ValueError(f"Unrecognised asset reference: {ref[:80]}"),
)
meta = self._make_request("GET", f"{BASE_URL}/files/{file_id}/download")
try:
meta = self._make_request("GET", f"{BASE_URL}/files/{file_id}/download")
except ProviderError as e:
if "HTTP 403" in str(e.original) and self._asset_record_missing(file_id):
raise ProviderError(
self.provider_name,
f"download_asset({file_id})",
RuntimeError(
"Asset no longer exists — HTTP 404 'File not found' on "
"/files/{id}. The upload was deleted or expired "
"server-side; it is not recoverable from ChatGPT."
),
) from e
raise
download_url = meta.get("download_url")
if not download_url:
raise ProviderError(
@@ -729,6 +752,26 @@ class ChatGPTProvider(BaseProvider):
)
return resp.content, mime, file_name
def _asset_record_missing(self, file_id: str) -> bool:
"""True if ``/files/{id}`` reports the asset gone.
Runs only on the 403 path, so it costs one extra request per failed
asset and none per successful one. Any other outcome — 200, a network
error, an unparseable body — returns False, leaving the original 403
to be reported as-is rather than guessing that it means deletion.
"""
try:
self._pace()
resp = self._session.request(
"GET", f"{BASE_URL}/files/{file_id}", timeout=REQUEST_TIMEOUT
)
except Exception as e: # noqa: BLE001 - a failed probe must not mask the 403
logger.debug(
"[chatgpt] Existence probe for %s failed: %s", file_id, e
)
return False
return resp.status_code == 404
# ------------------------------------------------------------------
# Normalization
# ------------------------------------------------------------------