close the media 403 investigation: the bytes are gone

The render check settled it. Scrolling the whole conversation found sets of
images that do NOT display in ChatGPT — a set of 6 and a set of 4 — and they
are exactly our failures: the 6 are the records that 404 outright, the 4 are
the refused attachment batch we dumped (image.png, image(1).png,
image(2).png + 1). Every set that displays downloaded fine.

So the 403 was never OpenAI withholding something. It is the same failure
their own UI hits. All 19 are unrecoverable.

What the investigation established, now recorded in media.py and the
changelog so it is not repeated:
  - uploads do not expire (36/36 sampled, 2025-09 → 2026-08, still live), so
    export cadence was never the variable
  - 7 records 404; 12 report state=ready with a library_file_id
  - /files/{id}/download mints the signed estuary/content URL the UI fetches,
    and refuses the survivors regardless of headers, Authorization, gizmo_id,
    conversation_id, Referer or namespace; the Library id is file_not_found
  - all survivors were created 2026-07-14 within minutes of each other yet
    appear in conversations predating that date — a Library migration that
    kept the metadata and lost the bytes

Dropping the nine one-off probes; their findings live in the code comments
and changelog, and git history has the scripts if they are ever wanted.
This commit is contained in:
JesseMarkowitz
2026-08-17 12:47:05 -04:00
parent d30a9510bb
commit dfa0645fba
11 changed files with 17 additions and 1627 deletions
+4 -1
View File
@@ -12,7 +12,10 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
### Changed
- Media download failures are bucketed as `forbidden` (403 — the asset exists, we were refused) separately from `download-error`, so the run summary distinguishes it from `expired-or-missing` (404).
- Media download failures are bucketed as `forbidden` (403 — the file record survives) separately from `download-error`, so the run summary distinguishes it from `expired-or-missing` (404).
### Notes
- **ChatGPT media 403s: investigated and closed (2026-08-17).** 19 images across 7 conversations would not download. They are unrecoverable, and not because of anything the exporter or the export schedule did. Findings, recorded so this is not re-litigated: uploads do **not** expire (36/36 sampled images from 2025-09 through 2026-08 are still live, so export cadence is not a factor); the failures split into 7 records that 404 outright and 12 that report `state: "ready"` with a `library_file_id`; `/files/{id}/download` is the endpoint that mints the signed `estuary/content` URL the web UI fetches, and it refuses the survivors with a bare 403 regardless of headers, `Authorization`, `gizmo_id`, `conversation_id`, Referer, or namespace, while the Library id is rejected as `file_not_found`. Decisively, those same images render blank in ChatGPT's own UI — nothing is being withheld from the exporter. All the survivors were created 2026-07-14 within minutes of each other yet appear in conversations predating that date, pointing at a Library migration that kept the metadata and lost the bytes.
## [0.8.0] - 2026-07-06
+13 -2
View File
@@ -160,8 +160,19 @@ def _classify_failure(error: ProviderError) -> str:
Note that a bare 403 from ChatGPT usually means *gone*, not *refused*:
its download endpoint reports a deleted upload as 403 Forbidden. The
provider confirms that against /files/{id} before raising, so a failure
that reaches here still saying "forbidden" is one where the asset really
does still exist — worth looking at, unlike an expired upload.
that reaches here still saying "forbidden" is one where the file record
survives.
``forbidden`` does NOT mean recoverable. Investigated exhaustively
2026-08-17 (19 assets): those records report ``state: "ready"`` and carry
a ``library_file_id``, yet /files/{id}/download — the endpoint that mints
the signed estuary/content URL the web UI itself fetches — refuses them,
and the Library id is rejected as ``file_not_found``. Decisively, the same
images render blank in ChatGPT's own UI. Nothing was withheld from us; the
bytes are gone and only the metadata survived, apparently from a Library
migration dated 2026-07-14. Do not spend another afternoon on it: no
header, scope, namespace or id form reaches these. The bucket stays
separate only because the two failures look different on the wire.
"""
detail = str(error.original).lower()
if "404" in detail or "not found" in detail or "no longer exists" in detail:
-175
View File
@@ -1,175 +0,0 @@
"""Offline: is the media loss age-based, source-based, or neither?
Answers the question the 403s raised — do I have to export within N days? —
from the exports already on disk. No API calls, no token, nothing to expire.
It works because the renderer records the outcome of every image in the
Markdown itself:
![user_upload](media/file_x.png) ← downloaded, still alive
> 🖼️ **Image attached** — `sediment://file_y`
(user_upload, content not preserved…) ← dead or never fetched
and the conversation's date is in its filename (YYYY-MM-DD_slug_id.md). So
grouping outcomes by month and by source distinguishes the hypotheses:
* Age-based expiry → old months all-dead, recent months all-alive, with a
clean cutoff between them.
* Source-based → user_upload dies while model_generated survives at
the same age.
* Neither → deaths scattered across months, or concentrated in a
few conversations while their neighbours survive.
IMPORTANT — what "saved" does and does not mean. It means the image was
downloaded by *some* run at *some* point, not that ChatGPT still holds it. An
image captured in June looks saved forever after, even if it died in July. So
a month with no losses shows the exports were timely; it cannot show the
assets survived. Use probe_survival_by_month.py to ask the API what is still
live today. The one signal here that is immune to this confound is a single
conversation that both kept and lost images: same age, same chat, opposite
outcomes means the cause is per-asset, not retention.
Run from the project root:
python tools/analyze_media_age.py
"""
import os
import re
import sys
from collections import defaultdict
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
DATE_RE = re.compile(r"(\d{4}-\d{2}-\d{2})_")
def main() -> None:
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
if not export_dir.is_dir():
print(f"exports dir not found: {export_dir}")
return
# month → source → {"saved": n, "dead": n}
stats: dict[str, dict[str, dict[str, int]]] = defaultdict(
lambda: defaultdict(lambda: {"saved": 0, "dead": 0})
)
# conversations that lost at least one image
losses: dict[str, dict[str, int]] = defaultdict(lambda: {"saved": 0, "dead": 0})
oldest_saved: dict[str, str] = {}
newest_dead: dict[str, str] = {}
oldest_dead: dict[str, str] = {}
for md in export_dir.rglob("*.md"):
date_match = DATE_RE.search(md.name)
if not date_match:
continue
date = date_match.group(1)
month = date[:7]
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
conv_key = f"{date} {md.stem}"
for source, _path in SAVED_RE.findall(text):
source = source or "unknown"
stats[month][source]["saved"] += 1
losses[conv_key]["saved"] += 1
if source not in oldest_saved or date < oldest_saved[source]:
oldest_saved[source] = date
for _ref, meta in DEAD_RE.findall(text):
source = meta.split(",")[0].strip() or "unknown"
stats[month][source]["dead"] += 1
losses[conv_key]["dead"] += 1
if source not in newest_dead or date > newest_dead[source]:
newest_dead[source] = date
if source not in oldest_dead or date < oldest_dead[source]:
oldest_dead[source] = date
if not stats:
print(f"no dated conversations with images found under {export_dir}")
return
sources = sorted({s for m in stats.values() for s in m})
print("=" * 78)
print("Image outcomes by conversation month")
print("=" * 78)
header = f"{'month':<9}"
for source in sources:
header += f" {source[:16]:>16} (saved/dead)"
print(header)
for month in sorted(stats):
row = f"{month:<9}"
for source in sources:
cell = stats[month].get(source, {"saved": 0, "dead": 0})
if cell["saved"] or cell["dead"]:
row += f" {cell['saved']:>8} / {cell['dead']:<17}"
else:
row += f" {'—':>8} {'':<17}"
print(row)
print()
print("=" * 78)
print("The age question")
print("=" * 78)
for source in sources:
old_s = oldest_saved.get(source)
old_d = oldest_dead.get(source)
new_d = newest_dead.get(source)
print(f" {source}:")
print(f" oldest still downloadable : {old_s or '—'}")
print(f" dead range : {old_d or '—'} … {new_d or '—'}")
if old_s and new_d and old_s < new_d:
print(
f" → an asset from {old_s} was captured while one from "
f"{new_d} was not."
)
print(
" This does NOT disprove expiry: 'captured' means some "
"run got it in time,"
)
print(
" not that it is still on ChatGPT today. Run "
"probe_survival_by_month.py to tell"
)
print(" those apart.")
elif old_s and old_d and old_s > old_d:
print(
f" → everything lost is older than everything captured "
f"(cutoff between {old_d} and {old_s}): consistent with expiry."
)
print()
print("=" * 78)
print("Conversations that lost images (clustering check)")
print("=" * 78)
lossy = {k: v for k, v in losses.items() if v["dead"]}
for conv, counts in sorted(lossy.items()):
print(f" {conv[:66]:<66} saved={counts['saved']:<4} dead={counts['dead']}")
total_dead = sum(v["dead"] for v in lossy.values())
print()
print(f" {len(lossy)} conversation(s) affected, {total_dead} image(s) lost")
mixed = [k for k, v in lossy.items() if v["saved"]]
if mixed:
print(
f" {len(mixed)} of them ALSO kept images — same conversation, same "
"age, different outcome:"
)
for conv in sorted(mixed):
print(f" {conv[:70]}")
print(" → whatever killed these is per-asset, not per-conversation.")
if __name__ == "__main__":
main()
-202
View File
@@ -1,202 +0,0 @@
"""Do the lost images sit on abandoned conversation branches?
Established so far: uploads do not expire (36/36 sampled images from
2025-09 through 2026-08 are still live), and the loss is per-asset — one
conversation on 2026-07-09 kept 38 images and lost 12. So something
distinguishes those 12 from their neighbours in the same chat.
Hypothesis: they are attached to messages that were edited or regenerated.
ChatGPT keeps the whole conversation as a tree; editing a message forks it,
leaving the superseded messages on an abandoned branch. The exporter walks
every node (chatgpt.py:963), so it exports abandoned branches too — and an
attachment that is no longer reachable from the live conversation is exactly
the kind of thing a backend would garbage-collect.
The test: fetch the raw conversation, compute the live branch by walking
parent links up from ``current_node``, then cross-tabulate every image
against (on the live branch?) x (still downloadable?).
on-branch alive + off-branch dead → confirmed, and it is not data loss:
those images belong to messages you
replaced.
dead on the live branch → hypothesis dead; the images are
genuinely missing from live history.
Run from the project root:
python tools/branch_check.py
"""
import os
import re
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider # noqa: E402
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
def find_mixed_conversations(export_dir: Path) -> list[tuple[Path, str, int, int]]:
"""Conversations that both kept and lost images — the informative ones."""
out = []
for md in export_dir.rglob("*.md"):
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
saved = len(SAVED_RE.findall(text))
dead = len(DEAD_RE.findall(text))
if not dead:
continue
match = CONV_ID_RE.search(text)
if not match:
continue
out.append((md, match.group(1), saved, dead))
return out
def live_branch(mapping: dict, current_node: str | None) -> set[str]:
"""Node IDs reachable by walking parent links up from current_node."""
live: set[str] = set()
node_id = current_node
while node_id and node_id in mapping and node_id not in live:
live.add(node_id)
node_id = mapping[node_id].get("parent")
return live
def image_refs(node: dict) -> list[str]:
"""asset_pointer refs on this node's message, if any."""
message = node.get("message") or {}
content = message.get("content") or {}
refs = []
if content.get("content_type") == "image_asset_pointer":
ref = content.get("asset_pointer")
if ref:
refs.append(ref)
for part in content.get("parts") or []:
if isinstance(part, dict) and part.get("content_type") == "image_asset_pointer":
ref = part.get("asset_pointer")
if ref:
refs.append(ref)
return refs
def still_live(provider, file_id: str) -> bool | None:
try:
provider._pace()
resp = provider._session.request(
"GET", f"{BASE_URL}/files/{file_id}", timeout=30
)
except Exception: # noqa: BLE001 - diagnostic
return None
if resp.status_code == 200:
return True
if resp.status_code == 404:
return False
return None
def main() -> None:
from src.providers.chatgpt import parse_asset_file_id
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
mixed = find_mixed_conversations(export_dir)
if not mixed:
print(f"no conversations with lost images found under {export_dir}")
return
provider = ChatGPTProvider()
grand = {"on_alive": 0, "on_dead": 0, "off_alive": 0, "off_dead": 0, "unknown": 0}
for md, conv_id, saved, dead in sorted(mixed, key=lambda r: -r[3]):
print("=" * 78)
print(f"{md.name}")
print(f" {saved} kept, {dead} lost conversation_id={conv_id}")
print("=" * 78)
try:
raw = provider.get_conversation(conv_id)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" could not fetch: {e}\n")
continue
mapping = raw.get("mapping") or {}
current = raw.get("current_node")
live_nodes = live_branch(mapping, current)
print(f" mapping nodes: {len(mapping)} live branch: {len(live_nodes)}")
if len(live_nodes) < len(mapping):
print(
f" → {len(mapping) - len(live_nodes)} node(s) are OFF the live "
"branch (edited or regenerated messages)"
)
else:
print(" → conversation is linear; nothing was edited or regenerated")
print()
rows = []
for node_id, node in mapping.items():
on_branch = node_id in live_nodes
for ref in image_refs(node):
file_id = parse_asset_file_id(ref) or ref
alive = still_live(provider, file_id)
rows.append((on_branch, alive, file_id))
if alive is None:
grand["unknown"] += 1
else:
key = ("on_" if on_branch else "off_") + ("alive" if alive else "dead")
grand[key] += 1
counts = {"on_alive": 0, "on_dead": 0, "off_alive": 0, "off_dead": 0, "unk": 0}
for on_branch, alive, _fid in rows:
if alive is None:
counts["unk"] += 1
else:
counts[("on_" if on_branch else "off_") + ("alive" if alive else "dead")] += 1
print(f" {'':<18}{'still live':>12}{'gone':>8}")
print(f" {'on live branch':<18}{counts['on_alive']:>12}{counts['on_dead']:>8}")
print(f" {'off (abandoned)':<18}{counts['off_alive']:>12}{counts['off_dead']:>8}")
if counts["unk"]:
print(f" ({counts['unk']} indeterminate)")
print()
for on_branch, alive, fid in sorted(rows, key=lambda r: (r[0], r[1] is not False)):
if alive is False:
where = "live branch" if on_branch else "ABANDONED branch"
print(f" LOST {fid[:40]:<42} {where}")
print()
print("=" * 78)
print("Verdict")
print("=" * 78)
print(
f" on live branch : {grand['on_alive']} live, {grand['on_dead']} gone"
)
print(
f" abandoned : {grand['off_alive']} live, {grand['off_dead']} gone"
)
if grand["off_dead"] and not grand["on_dead"]:
print()
print(" Every lost image is on an abandoned branch, and nothing on the")
print(" live conversation is missing. These are attachments to messages")
print(" you edited or regenerated — ChatGPT collects them once they are")
print(" no longer part of the conversation. Your live history is intact,")
print(" and no export schedule would have changed this.")
elif grand["on_dead"]:
print()
print(" Images are missing from the LIVE conversation — not explained by")
print(" editing. Real loss from current history; worth digging further.")
if __name__ == "__main__":
main()
-203
View File
@@ -1,203 +0,0 @@
"""The web UI serves images from /backend-api/estuary/content — can we mint that?
A URL copied from the browser looks like:
/backend-api/estuary/content?id=file_…&ts=496382&p=fs&cid=1&sig=…&v=0
Fetched with no cookies it returns 403 {"detail":"File stream access denied."},
so the signature is not a bypass — it still rides on the session. Two things
follow, and this probe checks both:
A. What does /files/{id}/download hand back for a file that WORKS? If its
download_url is one of these estuary URLs, then that endpoint is the
minting step, and a gizmo-scoped file failing there means we are refused
at exactly the point the signature would be issued. Then the only hope is
a different minting route.
B. Does the estuary namespace expose one? /backend-api/estuary/* is new to us
and was never probed — every earlier attempt used /backend-api/files/*.
Run from the project root:
python tools/estuary_probe.py
Optionally pass a browser URL for one of the UNREACHABLE images to check
whether the exporter's session can replay it:
python tools/estuary_probe.py "https://chatgpt.com/backend-api/estuary/content?id=…"
"""
import json
import os
import re
import sys
from pathlib import Path
from urllib.parse import parse_qs, urlparse
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
def sample_ids(export_dir: Path) -> tuple[str | None, list[str]]:
"""(one working file id, all failed file ids)."""
working = None
failed: list[str] = []
for md in export_dir.rglob("*.md"):
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
for ref, _meta in DEAD_RE.findall(text):
fid = parse_asset_file_id(ref)
if fid:
failed.append(fid)
if working is None:
for _src, rel in SAVED_RE.findall(text):
stem = Path(rel).stem
if stem.startswith("file"):
working = stem
break
return working, failed
def main() -> None:
pasted = sys.argv[1] if len(sys.argv) > 1 else None
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
working, failed = sample_ids(export_dir)
provider = ChatGPTProvider()
def call(label: str, url: str, headers: dict | None = None) -> None:
try:
provider._pace()
resp = provider._session.request("GET", url, headers=headers, timeout=30)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {label:<48} error {type(e).__name__}")
return
ctype = (resp.headers.get("content-type") or "").split(";")[0]
if resp.status_code == 200 and ctype.startswith("image/"):
print(f" {label:<48} 200 {ctype} {len(resp.content)}B ← IMAGE BYTES")
return
body = ""
try:
body = json.dumps(resp.json(), default=str)[:220]
except Exception: # noqa: BLE001 - diagnostic
body = (resp.text or "")[:160].replace("\n", " ")
print(f" {label:<48} {resp.status_code} {ctype}")
if body:
print(f" {body}")
# ── A. what a working file's download_url actually looks like ──────────
print("=" * 78)
print("A. The minting step, on a file that works")
print("=" * 78)
if working:
print(f" working file: {working}")
try:
provider._pace()
resp = provider._session.request(
"GET", f"{BASE_URL}/files/{working}/download", timeout=30
)
body = resp.json() if resp.status_code == 200 else {}
url = body.get("download_url") or ""
print(f" /files/{{id}}/download → {resp.status_code}")
if url:
parsed = urlparse(url)
print(f" host {parsed.netloc}")
print(f" path {parsed.path}")
params = parse_qs(parsed.query)
print(f" query {sorted(params)}")
if "estuary" in parsed.path:
print(" → SAME estuary route the browser uses.")
print(" So /files/{id}/download IS the minting step, and a")
print(" gizmo file's 403 is a refusal to mint. Look for")
print(" another minter below.")
else:
print(" → a different route from the browser's estuary URL;")
print(" the UI must mint its URLs somewhere else.")
except Exception as e: # noqa: BLE001 - diagnostic
print(f" could not check: {e}")
else:
print(" no working image found on disk to compare against")
# ── B. does the estuary namespace offer a route for refused files? ─────
print()
print("=" * 78)
print("B. The estuary namespace, on a refused file")
print("=" * 78)
if not failed:
print(" no failed images found")
else:
# The failures are two different classes and only one is interesting.
# A deleted file 404s on /files/{id}; a refused one answers 200. Testing
# a deleted file here proves nothing, so pick a refused one.
target = None
for candidate in failed:
try:
provider._pace()
meta = provider._session.request(
"GET", f"{BASE_URL}/files/{candidate}", timeout=30
)
except Exception: # noqa: BLE001 - diagnostic
continue
if meta.status_code == 200:
target = candidate
break
print(f" (skipping {candidate[:28]}… — deleted, {meta.status_code})")
if target is None:
print(" every failure is a deleted file; nothing in the refused class")
return
print(f"\n refused file: {target} (exists, download refused)\n")
for label, url in [
("estuary/files/{id}/download", f"{BASE_URL}/estuary/files/{target}/download"),
("estuary/files/{id}", f"{BASE_URL}/estuary/files/{target}"),
("estuary/content?id=", f"{BASE_URL}/estuary/content?id={target}"),
("estuary/content?id=&p=fs&cid=1&v=0", f"{BASE_URL}/estuary/content?id={target}&p=fs&cid=1&v=0"),
("estuary/{id}", f"{BASE_URL}/estuary/{target}"),
("estuary/download?id=", f"{BASE_URL}/estuary/download?id={target}"),
("files/{id}/download?v=0", f"{BASE_URL}/files/{target}/download?v=0"),
# Validation asked only for id, p and ts — never sig. Perhaps the
# signature is enforced elsewhere, or not at all for an owner.
("estuary/content (no sig)", f"{BASE_URL}/estuary/content?id={target}&p=fs&cid=1&v=0&ts=496382"),
]:
call(label, url)
# Does a working file's freshly minted URL serve if we swap in the
# refused id? If it does, the signature is not bound to the file.
if working:
try:
provider._pace()
resp = provider._session.request(
"GET", f"{BASE_URL}/files/{working}/download", timeout=30
)
minted = (resp.json() or {}).get("download_url") if resp.status_code == 200 else None
except Exception: # noqa: BLE001 - diagnostic
minted = None
if minted:
swapped = re.sub(r"id=file_[0-9a-f]+", f"id={target}", minted)
print()
call("working file's URL, refused id swapped in", swapped)
# ── C. replay a pasted URL through our session ─────────────────────────
if pasted:
print()
print("=" * 78)
print("C. Replaying the pasted URL through the exporter session")
print("=" * 78)
params = parse_qs(urlparse(pasted).query)
pasted_id = (params.get("id") or [""])[0]
print(f" id in URL: {pasted_id}")
print(f" that id is {'a REFUSED file' if pasted_id in failed else 'NOT one of the refused files'}")
call("pasted URL", pasted)
if __name__ == "__main__":
main()
-185
View File
@@ -1,185 +0,0 @@
"""How do you download a gizmo-scoped file?
metadata_diff.py found the discriminator: every refused image carries
``use_case: "gizmo"`` (ChatGPT's name for Projects and custom GPTs), while
everything that downloads is ``image_gen`` or ``multimodal``. The files are
healthy — ``state: "ready"`` — so this is not loss, it is the plain
/files/{id}/download endpoint declining to serve a project-scoped file.
So find the call that works. Two parts:
1. Dump the raw conversation node carrying one of these images. The part dict
may name the scope the download needs (a gizmo id, a file token, a
conversation id) — cheaper than guessing.
2. Try the plausible variants and print what each returns:
?gizmo_id=<each configured project> scope by query param
?conversation_id=<the conversation> scope by conversation
/gizmos/{gizmo}/files/{id}/download scope by path
Referer: the project URL scope by origin
Anything that returns 200 is the fix, and the exporter can adopt it.
Run from the project root:
python tools/gizmo_download_probe.py
"""
import json
import os
import re
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
from src.utils import redact_secrets # noqa: E402
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
def find_candidates(export_dir: Path) -> list[tuple[str, str, str]]:
"""(file_id, conversation_id, export filename) for images that failed."""
out = []
for md in export_dir.rglob("*.md"):
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
conv_match = CONV_ID_RE.search(text)
if not conv_match:
continue
for ref, _meta in DEAD_RE.findall(text):
file_id = parse_asset_file_id(ref)
if file_id:
out.append((file_id, conv_match.group(1), md.name))
return out
def is_gizmo_scoped(provider, file_id: str) -> bool:
provider._pace()
resp = provider._session.request("GET", f"{BASE_URL}/files/{file_id}", timeout=30)
if resp.status_code != 200:
return False
try:
return resp.json().get("use_case") == "gizmo"
except Exception: # noqa: BLE001 - diagnostic
return False
def dump_conversation_node(provider, conv_id: str, file_id: str) -> str | None:
"""Print the raw part carrying this asset. Returns a gizmo id if found."""
try:
raw = provider.get_conversation(conv_id)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" could not fetch conversation: {e}")
return None
gizmo_id = raw.get("gizmo_id") or raw.get("conversation_template_id")
print(f" conversation gizmo_id : {gizmo_id}")
print(f" conversation keys : {sorted(raw.keys())}")
for node_id, node in (raw.get("mapping") or {}).items():
message = node.get("message") or {}
content = message.get("content") or {}
candidates = []
if content.get("content_type") == "image_asset_pointer":
candidates.append(content)
for part in content.get("parts") or []:
if isinstance(part, dict):
candidates.append(part)
for part in candidates:
pointer = part.get("asset_pointer") or ""
if file_id in pointer:
print(f"\n --- raw part on node {node_id} ---")
print(json.dumps(redact_secrets(part), indent=4, default=str)[:1800])
meta = message.get("metadata") or {}
interesting = {
k: v for k, v in meta.items()
if any(t in k for t in ("gizmo", "file", "attach", "source"))
}
if interesting:
print(f"\n --- message.metadata (filtered) ---")
print(json.dumps(redact_secrets(interesting), indent=4, default=str)[:1200])
return gizmo_id
print(" (asset not found in the conversation mapping)")
return gizmo_id
def try_variants(provider, file_id: str, conv_id: str, gizmo_id: str | None) -> None:
project_ids = [p for p in provider._project_ids]
if gizmo_id and gizmo_id not in project_ids:
project_ids.insert(0, gizmo_id)
base = f"{BASE_URL}/files/{file_id}/download"
def show(label: str, url: str, headers: dict | None = None) -> None:
try:
provider._pace()
resp = provider._session.request("GET", url, headers=headers, timeout=30)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {label:<52} error {type(e).__name__}")
return
note = ""
if resp.status_code == 200:
try:
note = " ← WORKS" if resp.json().get("download_url") else " (200, no download_url)"
except Exception: # noqa: BLE001 - diagnostic
note = " ← WORKS (non-JSON)"
print(f" {label:<52} {resp.status_code}{note}")
print("\n --- variants ---")
show("baseline", base)
show("?conversation_id", f"{base}?conversation_id={conv_id}")
for pid in project_ids[:6]:
show(f"?gizmo_id={pid[:22]}…", f"{base}?gizmo_id={pid}")
for pid in project_ids[:2]:
show(
f"/gizmos/{pid[:16]}…/files/…/download",
f"{BASE_URL}/gizmos/{pid}/files/{file_id}/download",
)
if gizmo_id:
show(
"Referer: project URL",
base,
{"Referer": f"https://chatgpt.com/g/{gizmo_id}/project"},
)
def main() -> None:
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
candidates = find_candidates(export_dir)
if not candidates:
print(f"no failed images found under {export_dir}")
return
provider = ChatGPTProvider()
print(f"configured project ids: {len(provider._project_ids)}")
tested = 0
for file_id, conv_id, md_name in candidates:
if not is_gizmo_scoped(provider, file_id):
continue
print()
print("=" * 78)
print(f"{file_id}")
print(f" from {md_name}")
print("=" * 78)
gizmo_id = dump_conversation_node(provider, conv_id, file_id)
try_variants(provider, file_id, conv_id, gizmo_id)
tested += 1
if tested >= 2:
break
if not tested:
print("no gizmo-scoped failures found — nothing to probe")
if __name__ == "__main__":
main()
-192
View File
@@ -1,192 +0,0 @@
"""Can a refused image be fetched by its Library ID instead?
gizmo_download_probe dumped the message metadata and found what the asset
pointer never carried: each attachment has a second identity.
"id": "file_000000003454722f9481506b96aed510" ← refused
"library_file_id": "libfile_4eb82f478fe081919127e2eba9886e86"
"source": "local"
The exporter only ever knew the sediment id from the asset pointer, and asks
/files/{sediment_id}/download — which 403s for these. The Library is a
separate store with its own ids, so the natural reading is that we are asking
for a conversation-scoped copy of a file that now lives in the Library.
This walks the attachments in message.metadata to pair each refused sediment
id with its library_file_id, then tries the endpoints that could serve it.
Whatever returns a download_url is what the exporter should use for any
attachment carrying a library_file_id.
Run from the project root:
python tools/library_download_probe.py
"""
import json
import os
import re
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
def find_failed(export_dir: Path) -> dict[str, set[str]]:
"""conversation_id → {failed sediment file ids}."""
out: dict[str, set[str]] = {}
for md in export_dir.rglob("*.md"):
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
conv = CONV_ID_RE.search(text)
if not conv:
continue
ids = {
fid
for ref, _meta in DEAD_RE.findall(text)
if (fid := parse_asset_file_id(ref))
}
if ids:
out.setdefault(conv.group(1), set()).update(ids)
return out
def attachment_index(raw: dict) -> dict[str, dict]:
"""sediment file id → its attachment record from message.metadata."""
index: dict[str, dict] = {}
for node in (raw.get("mapping") or {}).values():
message = node.get("message") or {}
for att in (message.get("metadata") or {}).get("attachments") or []:
if isinstance(att, dict) and att.get("id"):
index[att["id"]] = att
return index
def main() -> None:
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
failed = find_failed(export_dir)
if not failed:
print(f"no failed images found under {export_dir}")
return
provider = ChatGPTProvider()
pairs: list[tuple[str, str, str]] = []
print("=" * 78)
print("Pairing refused assets with their Library IDs")
print("=" * 78)
for conv_id, file_ids in failed.items():
try:
raw = provider.get_conversation(conv_id)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {conv_id[:8]}… could not fetch: {e}")
continue
index = attachment_index(raw)
for file_id in sorted(file_ids):
att = index.get(file_id)
lib = (att or {}).get("library_file_id")
source = (att or {}).get("source")
print(f" {file_id[:34]}… library={str(lib)[:34]:<36} source={source}")
if lib:
pairs.append((file_id, lib, conv_id))
if not pairs:
print("\n No refused asset carries a library_file_id — different cause.")
return
file_id, lib_id, conv_id = pairs[0]
print()
print("=" * 78)
print(f"Endpoint hunt for {lib_id}")
print(f" (sediment id {file_id})")
print("=" * 78)
def show(label: str, url: str, method: str = "GET") -> bool:
"""Print status AND body. A 200 carrying an error envelope says what
the endpoint wants — printing only the keys threw that away."""
try:
provider._pace()
resp = provider._session.request(method, url, timeout=30)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {label:<50} error {type(e).__name__}")
return False
hit = False
try:
body = resp.json()
except Exception: # noqa: BLE001 - diagnostic
preview = (resp.text or "").strip()[:200]
if resp.status_code == 200 and resp.content:
print(f" {label:<50} 200 non-JSON ({len(resp.content)} bytes) ← BYTES?")
return True
print(f" {label:<50} {resp.status_code} {preview}")
return False
if isinstance(body, dict) and body.get("download_url"):
print(f" {label:<50} {resp.status_code} ← DOWNLOAD URL")
print(f" {json.dumps(body, default=str)[:300]}")
return True
print(f" {label:<50} {resp.status_code}")
print(f" {json.dumps(body, default=str)[:400]}")
return hit
candidates = [
# This one already returns 200 with an error envelope — read it.
("/files/{lib}/download", f"{BASE_URL}/files/{lib_id}/download"),
("/files/{lib}/download?conversation_id=", f"{BASE_URL}/files/{lib_id}/download?conversation_id={conv_id}"),
("/files/{lib}/download (POST)", f"{BASE_URL}/files/{lib_id}/download"),
# How does the UI enumerate the Library? A listing tends to reveal both
# the id form it uses and the route that serves bytes.
("/library/files?limit=3", f"{BASE_URL}/library/files?limit=3"),
("/library?limit=3", f"{BASE_URL}/library?limit=3"),
("/files?limit=3", f"{BASE_URL}/files?limit=3"),
("/my_files?limit=3", f"{BASE_URL}/my_files?limit=3"),
# Content routes that serve the asset rather than a signed URL.
("/files/{lib}/content", f"{BASE_URL}/files/{lib_id}/content"),
("/content?asset_pointer=sediment://{sediment}", f"{BASE_URL}/content?asset_pointer=sediment://{file_id}"),
]
winners = []
for label, url in candidates:
method = "POST" if "(POST)" in label else "GET"
if show(label, url, method):
winners.append((label, url))
# If a second file behaves differently, the first one was the anomaly.
if len(pairs) > 1:
_other_file, other_lib, _other_conv = pairs[1]
print()
print(f" --- second sample: {other_lib} ---")
show("/files/{lib}/download", f"{BASE_URL}/files/{other_lib}/download")
print()
print("=" * 78)
print("Result")
print("=" * 78)
if winners:
print(" Served by:")
for label, url in winners:
print(f" {label}")
print()
print(" → The exporter can pair each asset_pointer with the")
print(" library_file_id in message.metadata.attachments and fetch")
print(f" {len(pairs)} otherwise-unreachable image(s) this way.")
else:
print(" None of these served the file. The Library ID is real but the")
print(" route is elsewhere — next step is watching what chatgpt.com")
print(" itself requests when it renders one of these images.")
if __name__ == "__main__":
main()
-176
View File
@@ -1,176 +0,0 @@
"""What distinguishes an image ChatGPT refuses to serve from one it serves?
branch_check.py turned up something the earlier probing missed: of the images
that failed to download, only some are actually gone. The rest answer
/files/{id} with 200 and full metadata, and refuse only /download. Three
distinct states:
gone 404 on /files/{id} — deleted, unrecoverable
refused 200 on /files/{id}, 403 on download — exists, will not serve
working 200 on both — fine
"refused" is the interesting one, because those files still exist and may be
recoverable. Since the metadata comes back for them, the discriminator can be
read straight off: fetch metadata for every file in each state and compare the
fields. A field that is constant within "refused" and different in "working"
is the cause.
Also re-checks /download now. If a file that was refused during the export
serves today, the failure was transient and a retry pass recovers it — a very
different fix from anything permanent.
Run from the project root:
python tools/metadata_diff.py
"""
import json
import os
import re
import sys
from collections import defaultdict
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider # noqa: E402
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
WORKING_SAMPLE = 6
def collect(export_dir: Path) -> tuple[list[str], list[str]]:
"""(ids that failed to download, ids that downloaded fine)."""
from src.providers.chatgpt import parse_asset_file_id
failed: list[str] = []
working: list[str] = []
for md in export_dir.rglob("*.md"):
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
for ref, _meta in DEAD_RE.findall(text):
file_id = parse_asset_file_id(ref)
if file_id:
failed.append(file_id)
for _source, rel in SAVED_RE.findall(text):
stem = Path(rel).stem
if stem.startswith("file"):
working.append(stem)
return failed, working
def probe(provider, file_id: str) -> tuple[str, dict]:
"""(state, metadata) for one file."""
provider._pace()
meta_resp = provider._session.request(
"GET", f"{BASE_URL}/files/{file_id}", timeout=30
)
if meta_resp.status_code == 404:
return "gone", {}
if meta_resp.status_code != 200:
return f"meta-{meta_resp.status_code}", {}
try:
meta = meta_resp.json()
except Exception: # noqa: BLE001 - diagnostic
meta = {}
provider._pace()
dl_resp = provider._session.request(
"GET", f"{BASE_URL}/files/{file_id}/download", timeout=30
)
state = "working" if dl_resp.status_code == 200 else f"refused-{dl_resp.status_code}"
return state, meta
def main() -> None:
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
failed, working = collect(export_dir)
if not failed:
print(f"no failed images found under {export_dir}")
return
# Sample working files across the archive rather than all of them.
step = max(1, len(working) // WORKING_SAMPLE)
working_sample = working[::step][:WORKING_SAMPLE]
provider = ChatGPTProvider()
by_state: dict[str, list[tuple[str, dict]]] = defaultdict(list)
print("=" * 78)
print(f"Probing {len(failed)} failed + {len(working_sample)} working files")
print("=" * 78)
for file_id in failed + working_sample:
state, meta = probe(provider, file_id)
by_state[state].append((file_id, meta))
print(f" {file_id[:40]:<42} {state}")
print()
print("=" * 78)
print("States")
print("=" * 78)
for state, entries in sorted(by_state.items()):
print(f" {state:<16} {len(entries)}")
print()
if any(s.startswith("refused") for s in by_state) is False:
print(" No file is in the 'refused' state right now.")
print(" → Every previously-failed file that still exists now serves.")
print(" The export-time 403s were TRANSIENT; a retry pass recovers them.")
print()
print("=" * 78)
print("Metadata field comparison")
print("=" * 78)
states = [s for s in by_state if by_state[s] and any(m for _i, m in by_state[s])]
all_fields: set[str] = set()
for state in states:
for _fid, meta in by_state[state]:
all_fields.update(meta.keys())
for field in sorted(all_fields):
line = f" {field:<24}"
distinct_per_state = []
for state in sorted(states):
values = {
json.dumps(meta.get(field), default=str)[:28]
for _fid, meta in by_state[state]
if meta
}
shown = ", ".join(sorted(values)[:3])
if len(values) > 3:
shown += f" (+{len(values) - 3} more)"
distinct_per_state.append((state, values, shown))
line += f" [{state}] {shown}"
# Flag fields that cleanly separate the states.
value_sets = [v for _s, v, _sh in distinct_per_state]
if len(value_sets) > 1 and all(
not (a & b) for i, a in enumerate(value_sets) for b in value_sets[i + 1:]
):
line += " ← DISCRIMINATOR"
print(line)
print()
print("=" * 78)
print("One full record per state")
print("=" * 78)
for state in sorted(states):
fid, meta = next(((f, m) for f, m in by_state[state] if m), (None, None))
if not meta:
continue
print(f" --- {state} ({fid}) ---")
for key, value in sorted(meta.items()):
print(f" {key:<24} {json.dumps(value, default=str)[:90]}")
print()
if __name__ == "__main__":
main()
-192
View File
@@ -1,192 +0,0 @@
"""The browser mints signatures for files we get 403 on. What does it send that we don't?
Established:
* The images DO render in ChatGPT, so a valid signature exists for them.
* /files/{id}/download is the minting endpoint — a working file's
download_url is the same estuary/content URL the browser uses.
* Signatures are per-file: swapping an id into a working URL gives
"Invalid signature or expired URL".
So the difference is in the request, not the route. Our call sends an
Authorization: Bearer header plus session cookies; a browser leans on cookies
and adds client headers (oai-device-id, oai-client-version, a conversation
Referer) that a gated endpoint may check. This walks those variants against a
file that is currently refused, and prints which combination mints a URL.
Any 200 carrying a download_url is the answer, and the exporter adopts those
headers.
Run from the project root:
python tools/mint_headers_probe.py
"""
import json
import os
import re
import sys
import uuid
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
def find_failed(export_dir: Path) -> list[tuple[str, str]]:
"""(file_id, conversation_id) for every image that failed to download."""
out = []
for md in export_dir.rglob("*.md"):
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
conv = CONV_ID_RE.search(text)
if not conv:
continue
for ref, _meta in DEAD_RE.findall(text):
fid = parse_asset_file_id(ref)
if fid:
out.append((fid, conv.group(1)))
return out
def main() -> None:
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
failed = find_failed(export_dir)
if not failed:
print(f"no failed images found under {export_dir}")
return
provider = ChatGPTProvider()
session = provider._session
# Pick one that still exists — a deleted file 403s no matter what we send.
target = conv_id = None
for fid, cid in failed:
provider._pace()
meta = session.request("GET", f"{BASE_URL}/files/{fid}", timeout=30)
if meta.status_code == 200:
target, conv_id = fid, cid
break
if not target:
print("every failure is a deleted file — nothing to test")
return
url = f"{BASE_URL}/files/{target}/download"
print("=" * 78)
print(f"Minting attempts for {target}")
print(f" conversation {conv_id}")
print("=" * 78)
def attempt(
label: str,
*,
add: dict | None = None,
drop: tuple[str, ...] = (),
target_url: str | None = None,
) -> bool:
removed = {}
for header in drop:
for key in list(session.headers.keys()):
if key.lower() == header.lower():
removed[key] = session.headers.pop(key)
try:
provider._pace()
resp = session.request(
"GET", target_url or url, headers=add or None, timeout=30
)
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {label:<46} error {type(e).__name__}")
return False
finally:
session.headers.update(removed)
try:
body = resp.json()
except Exception: # noqa: BLE001 - diagnostic
body = {}
if resp.status_code == 200 and isinstance(body, dict) and body.get("download_url"):
print(f" {label:<46} 200 ← MINTED")
print(f" {str(body['download_url'])[:110]}")
return True
detail = body.get("detail") if isinstance(body, dict) else None
print(f" {label:<46} {resp.status_code} {json.dumps(detail, default=str)[:60] if detail else ''}")
return False
device_id = str(uuid.uuid4())
winners = []
checks = [
("baseline (bearer + cookies)", {}, ()),
("no Authorization (cookies only)", {}, ("Authorization",)),
("Referer: the conversation", {"Referer": f"https://chatgpt.com/c/{conv_id}"}, ()),
("oai-device-id", {"oai-device-id": device_id}, ()),
("oai-language", {"oai-language": "en-US"}, ()),
(
"browser-ish set",
{
"Referer": f"https://chatgpt.com/c/{conv_id}",
"oai-device-id": device_id,
"oai-language": "en-US",
"sec-fetch-site": "same-origin",
"sec-fetch-mode": "cors",
"sec-fetch-dest": "empty",
},
(),
),
(
"browser-ish, no Authorization",
{
"Referer": f"https://chatgpt.com/c/{conv_id}",
"oai-device-id": device_id,
"oai-language": "en-US",
},
("Authorization",),
),
("Accept: image/*", {"Accept": "image/avif,image/webp,*/*"}, ()),
]
for label, add, drop in checks:
if attempt(label, add=add, drop=drop):
winners.append(label)
# Same call with the conversation named in the query string.
if attempt(
"?conversation_id=",
target_url=f"{url}?conversation_id={conv_id}",
):
winners.append("?conversation_id=")
print()
print("=" * 78)
print("Result")
print("=" * 78)
if winners:
print(" Minted by:")
for label in winners:
print(f" {label}")
print()
print(" → Adopt those headers in ChatGPTProvider and the 11 refused")
print(" images become downloadable.")
else:
print(" No header combination minted a URL.")
print()
print(" The browser is getting its signature some other way. Find it:")
print(" 1. Open the conversation in ChatGPT with DevTools → Network.")
print(" 2. Press Ctrl+F (search inside requests) and search for")
print(f" {target[5:22]}")
print(" 3. The hit that is NOT the estuary/content request is the")
print(" call that mints the signature — send me its URL, method,")
print(" and (if POST) its request body.")
if __name__ == "__main__":
main()
-153
View File
@@ -1,153 +0,0 @@
"""Are already-downloaded images still alive on ChatGPT, or did we just catch them in time?
analyze_media_age.py cannot answer this. It reports whether an image was ever
captured — and an image downloaded by a run back in June looks "saved" forever
after, whether or not ChatGPT still holds it today. So a month with zero losses
proves the exports were timely, not that the assets survived.
This closes that gap: take images that ARE on disk, grouped by the month of
their conversation, and ask the API whether each still exists right now.
Old months still 200 → assets do not expire with age. Export cadence is not
what saved them, and shortening it would not have
saved the July losses either.
Old months now 404 → assets DO die with age; the old months are on disk
only because a run reached them in time. Cadence is
the whole ballgame, and the July losses are the first
ones we simply arrived too late for.
Liveness is checked with /files/{id}, which answers a missing record with a
clean 404 (verified 2026-08-17); /files/{id}/download reports the same state
as a misleading 403.
Run from the project root:
python tools/probe_survival_by_month.py [samples_per_month]
"""
import os
import re
import sys
from collections import defaultdict
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
from src.providers.chatgpt import BASE_URL, ChatGPTProvider # noqa: E402
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
DATE_RE = re.compile(r"(\d{4}-\d{2}-\d{2})_")
DEFAULT_SAMPLES = 4
def collect_saved_ids(export_dir: Path) -> dict[str, list[tuple[str, str]]]:
"""month → [(file_id, source)] for images already downloaded to disk."""
by_month: dict[str, list[tuple[str, str]]] = defaultdict(list)
for md in export_dir.rglob("*.md"):
date_match = DATE_RE.search(md.name)
if not date_match:
continue
month = date_match.group(1)[:7]
try:
text = md.read_text(encoding="utf-8", errors="replace")
except OSError:
continue
for source, rel_path in SAVED_RE.findall(text):
file_id = Path(rel_path).stem
if file_id.startswith("file"):
by_month[month].append((file_id, source or "unknown"))
return by_month
def main() -> None:
samples = DEFAULT_SAMPLES
if len(sys.argv) > 1:
try:
samples = max(1, int(sys.argv[1]))
except ValueError:
pass
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
by_month = collect_saved_ids(export_dir)
if not by_month:
print(f"no downloaded images found under {export_dir}")
return
provider = ChatGPTProvider()
print("=" * 78)
print(f"Are on-disk images still live on ChatGPT? ({samples} sampled per month)")
print("=" * 78)
print(f"{'month':<9} {'alive':>6} {'gone':>6} {'err':>5} sampled sources")
totals = {"alive": 0, "gone": 0, "err": 0}
verdict_rows: list[tuple[str, int, int]] = []
for month in sorted(by_month):
entries = by_month[month]
# Spread the sample across the month rather than taking the first few
# from one conversation.
step = max(1, len(entries) // samples)
picked = entries[::step][:samples]
alive = gone = err = 0
sources: list[str] = []
for file_id, source in picked:
sources.append(source[:4])
try:
provider._pace()
resp = provider._session.request(
"GET", f"{BASE_URL}/files/{file_id}", timeout=30
)
except Exception: # noqa: BLE001 - diagnostic
err += 1
continue
if resp.status_code == 200:
alive += 1
elif resp.status_code == 404:
gone += 1
else:
err += 1
totals["alive"] += alive
totals["gone"] += gone
totals["err"] += err
verdict_rows.append((month, alive, gone))
print(
f"{month:<9} {alive:>6} {gone:>6} {err:>5} "
f"{','.join(sources)} (of {len(entries)} on disk)"
)
print()
print(f"total: {totals['alive']} alive, {totals['gone']} gone, {totals['err']} error")
print()
print("=" * 78)
print("Verdict")
print("=" * 78)
old_gone = [m for m, _a, g in verdict_rows[:-2] if g]
old_alive = [m for m, a, _g in verdict_rows[:-2] if a]
if old_gone and not old_alive:
print(" Every sampled older image is GONE from ChatGPT.")
print(" → Uploads expire. Your old exports survive only because a run")
print(" reached them in time. Export cadence directly determines")
print(" what you can still save.")
elif old_alive and not old_gone:
print(" Every sampled older image is STILL LIVE on ChatGPT.")
print(" → Uploads do not expire with age. The July losses are")
print(" something else, and exporting sooner would not have")
print(" prevented them.")
else:
print(" Mixed: some older images alive, some gone.")
print(" → Not a clean expiry. Compare the 'sampled sources' column and")
print(" the per-month rates above; the cause is likely per-asset.")
if __name__ == "__main__":
main()
-146
View File
@@ -1,146 +0,0 @@
"""Can the exporter reuse an image URL copied from the browser?
The endpoint hunt bottomed out: the Library id is rejected outright
(file_not_found from GetDownloadLinkError), and /backend-api/content wants
signed query params (id, ts, p, and in practice a signature) that cannot be
guessed. The web UI has those values, so the remaining question is whether a
URL taken from the UI works outside the browser.
Paste the request URL DevTools shows for one of the unreachable images:
python tools/try_pasted_url.py "https://chatgpt.com/backend-api/content?id=…&ts=…&p=…&sig=…"
It fetches that URL three ways, and the pattern of results says what to build:
works with our session, not bare → the signature is fine but the request
needs auth; the exporter can mint these
itself if we find what returns them.
works bare too → the URL is self-authenticating; whatever
produced it is the endpoint we need.
works in neither → the URL is bound to the browser session
(or already expired — check ts), so the
exporter cannot reuse it as-is.
Nothing is written to disk unless --save is passed.
"""
import sys
from pathlib import Path
from urllib.parse import parse_qs, urlparse
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from dotenv import load_dotenv
load_dotenv()
def describe(url: str) -> None:
parsed = urlparse(url)
print(f" host : {parsed.netloc}")
print(f" path : {parsed.path}")
params = parse_qs(parsed.query)
print(" query :")
for key, values in params.items():
value = values[0] if values else ""
shown = value if len(value) <= 48 else f"{value[:45]}…"
print(f" {key:<12} {shown}")
def main() -> None:
args = [a for a in sys.argv[1:] if a != "--save"]
save = "--save" in sys.argv
if not args:
print(__doc__)
return
url = args[0]
print("=" * 78)
print("The URL")
print("=" * 78)
describe(url)
print()
print("=" * 78)
print("Fetch attempts")
print("=" * 78)
def report(label: str, resp) -> bytes | None:
ctype = (resp.headers.get("content-type") or "").split(";")[0]
size = len(resp.content or b"")
verdict = ""
if resp.status_code == 200 and ctype.startswith("image/"):
verdict = " ← IMAGE BYTES"
elif resp.status_code == 200:
verdict = f" 200 but {ctype or 'unknown type'}"
print(f" {label:<34} {resp.status_code} {ctype} {size}B{verdict}")
if resp.status_code == 200 and ctype.startswith("image/"):
return resp.content
if resp.status_code != 200:
preview = (resp.text or "")[:160].replace("\n", " ")
if preview:
print(f" {preview}")
return None
image: bytes | None = None
# 1. Bare request, no cookies, no auth headers.
try:
from curl_cffi import requests as curl_requests
bare = curl_requests.Session(impersonate="chrome120")
image = report("bare (no auth)", bare.get(url, timeout=30)) or image
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {'bare (no auth)':<34} error {type(e).__name__}: {e}")
# 2. Through the exporter's authenticated session.
try:
from src.providers.chatgpt import ChatGPTProvider
provider = ChatGPTProvider()
provider._pace()
image = report(
"exporter session", provider._session.request("GET", url, timeout=30)
) or image
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {'exporter session':<34} error {type(e).__name__}: {e}")
provider = None
# 3. Authenticated session without the Authorization header, in case the
# signature and the bearer token conflict.
if provider is not None:
try:
saved = provider._session.headers.pop("Authorization", None)
provider._pace()
image = report(
"session, no Authorization",
provider._session.request("GET", url, timeout=30),
) or image
if saved:
provider._session.headers["Authorization"] = saved
except Exception as e: # noqa: BLE001 - diagnostic
print(f" {'session, no Authorization':<34} error {type(e).__name__}")
print()
print("=" * 78)
print("Verdict")
print("=" * 78)
if image:
print(f" Served {len(image)} bytes of image data.")
print(" → The exporter CAN fetch these. Next: find what mints the")
print(" signed params, so it can build the URL itself.")
if save:
out = Path("pasted_url_result.bin")
out.write_bytes(image)
print(f" Saved to {out}")
else:
print(" (pass --save to write the bytes out)")
else:
print(" No attempt returned image bytes.")
print(" → Either the URL is bound to the browser session, or its ts has")
print(" expired. Re-copy a fresh URL and retry once; if it still")
print(" fails, these images are not reachable programmatically.")
if __name__ == "__main__":
main()