close the media 403 investigation: the bytes are gone
The render check settled it. Scrolling the whole conversation found sets of
images that do NOT display in ChatGPT — a set of 6 and a set of 4 — and they
are exactly our failures: the 6 are the records that 404 outright, the 4 are
the refused attachment batch we dumped (image.png, image(1).png,
image(2).png + 1). Every set that displays downloaded fine.
So the 403 was never OpenAI withholding something. It is the same failure
their own UI hits. All 19 are unrecoverable.
What the investigation established, now recorded in media.py and the
changelog so it is not repeated:
- uploads do not expire (36/36 sampled, 2025-09 → 2026-08, still live), so
export cadence was never the variable
- 7 records 404; 12 report state=ready with a library_file_id
- /files/{id}/download mints the signed estuary/content URL the UI fetches,
and refuses the survivors regardless of headers, Authorization, gizmo_id,
conversation_id, Referer or namespace; the Library id is file_not_found
- all survivors were created 2026-07-14 within minutes of each other yet
appear in conversations predating that date — a Library migration that
kept the metadata and lost the bytes
Dropping the nine one-off probes; their findings live in the code comments
and changelog, and git history has the scripts if they are ever wanted.
This commit is contained in:
+4
-1
@@ -12,7 +12,10 @@ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
|
|||||||
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
|
- **`tests/test_config.py::TestSessionLimiterConfig::test_defaults` depended on the developer's `.env`.** `load_config()` calls `load_dotenv(override=False)`, which re-populated the variable the test had just deleted — so it passed only on a machine with no `.env`. The test now stubs dotenv discovery.
|
||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
- Media download failures are bucketed as `forbidden` (403 — the asset exists, we were refused) separately from `download-error`, so the run summary distinguishes it from `expired-or-missing` (404).
|
- Media download failures are bucketed as `forbidden` (403 — the file record survives) separately from `download-error`, so the run summary distinguishes it from `expired-or-missing` (404).
|
||||||
|
|
||||||
|
### Notes
|
||||||
|
- **ChatGPT media 403s: investigated and closed (2026-08-17).** 19 images across 7 conversations would not download. They are unrecoverable, and not because of anything the exporter or the export schedule did. Findings, recorded so this is not re-litigated: uploads do **not** expire (36/36 sampled images from 2025-09 through 2026-08 are still live, so export cadence is not a factor); the failures split into 7 records that 404 outright and 12 that report `state: "ready"` with a `library_file_id`; `/files/{id}/download` is the endpoint that mints the signed `estuary/content` URL the web UI fetches, and it refuses the survivors with a bare 403 regardless of headers, `Authorization`, `gizmo_id`, `conversation_id`, Referer, or namespace, while the Library id is rejected as `file_not_found`. Decisively, those same images render blank in ChatGPT's own UI — nothing is being withheld from the exporter. All the survivors were created 2026-07-14 within minutes of each other yet appear in conversations predating that date, pointing at a Library migration that kept the metadata and lost the bytes.
|
||||||
|
|
||||||
## [0.8.0] - 2026-07-06
|
## [0.8.0] - 2026-07-06
|
||||||
|
|
||||||
|
|||||||
+13
-2
@@ -160,8 +160,19 @@ def _classify_failure(error: ProviderError) -> str:
|
|||||||
Note that a bare 403 from ChatGPT usually means *gone*, not *refused*:
|
Note that a bare 403 from ChatGPT usually means *gone*, not *refused*:
|
||||||
its download endpoint reports a deleted upload as 403 Forbidden. The
|
its download endpoint reports a deleted upload as 403 Forbidden. The
|
||||||
provider confirms that against /files/{id} before raising, so a failure
|
provider confirms that against /files/{id} before raising, so a failure
|
||||||
that reaches here still saying "forbidden" is one where the asset really
|
that reaches here still saying "forbidden" is one where the file record
|
||||||
does still exist — worth looking at, unlike an expired upload.
|
survives.
|
||||||
|
|
||||||
|
``forbidden`` does NOT mean recoverable. Investigated exhaustively
|
||||||
|
2026-08-17 (19 assets): those records report ``state: "ready"`` and carry
|
||||||
|
a ``library_file_id``, yet /files/{id}/download — the endpoint that mints
|
||||||
|
the signed estuary/content URL the web UI itself fetches — refuses them,
|
||||||
|
and the Library id is rejected as ``file_not_found``. Decisively, the same
|
||||||
|
images render blank in ChatGPT's own UI. Nothing was withheld from us; the
|
||||||
|
bytes are gone and only the metadata survived, apparently from a Library
|
||||||
|
migration dated 2026-07-14. Do not spend another afternoon on it: no
|
||||||
|
header, scope, namespace or id form reaches these. The bucket stays
|
||||||
|
separate only because the two failures look different on the wire.
|
||||||
"""
|
"""
|
||||||
detail = str(error.original).lower()
|
detail = str(error.original).lower()
|
||||||
if "404" in detail or "not found" in detail or "no longer exists" in detail:
|
if "404" in detail or "not found" in detail or "no longer exists" in detail:
|
||||||
|
|||||||
@@ -1,175 +0,0 @@
|
|||||||
"""Offline: is the media loss age-based, source-based, or neither?
|
|
||||||
|
|
||||||
Answers the question the 403s raised — do I have to export within N days? —
|
|
||||||
from the exports already on disk. No API calls, no token, nothing to expire.
|
|
||||||
|
|
||||||
It works because the renderer records the outcome of every image in the
|
|
||||||
Markdown itself:
|
|
||||||
|
|
||||||
 ← downloaded, still alive
|
|
||||||
> 🖼️ **Image attached** — `sediment://file_y`
|
|
||||||
(user_upload, content not preserved…) ← dead or never fetched
|
|
||||||
|
|
||||||
and the conversation's date is in its filename (YYYY-MM-DD_slug_id.md). So
|
|
||||||
grouping outcomes by month and by source distinguishes the hypotheses:
|
|
||||||
|
|
||||||
* Age-based expiry → old months all-dead, recent months all-alive, with a
|
|
||||||
clean cutoff between them.
|
|
||||||
* Source-based → user_upload dies while model_generated survives at
|
|
||||||
the same age.
|
|
||||||
* Neither → deaths scattered across months, or concentrated in a
|
|
||||||
few conversations while their neighbours survive.
|
|
||||||
|
|
||||||
IMPORTANT — what "saved" does and does not mean. It means the image was
|
|
||||||
downloaded by *some* run at *some* point, not that ChatGPT still holds it. An
|
|
||||||
image captured in June looks saved forever after, even if it died in July. So
|
|
||||||
a month with no losses shows the exports were timely; it cannot show the
|
|
||||||
assets survived. Use probe_survival_by_month.py to ask the API what is still
|
|
||||||
live today. The one signal here that is immune to this confound is a single
|
|
||||||
conversation that both kept and lost images: same age, same chat, opposite
|
|
||||||
outcomes means the cause is per-asset, not retention.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/analyze_media_age.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from collections import defaultdict
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
DATE_RE = re.compile(r"(\d{4}-\d{2}-\d{2})_")
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
if not export_dir.is_dir():
|
|
||||||
print(f"exports dir not found: {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
# month → source → {"saved": n, "dead": n}
|
|
||||||
stats: dict[str, dict[str, dict[str, int]]] = defaultdict(
|
|
||||||
lambda: defaultdict(lambda: {"saved": 0, "dead": 0})
|
|
||||||
)
|
|
||||||
# conversations that lost at least one image
|
|
||||||
losses: dict[str, dict[str, int]] = defaultdict(lambda: {"saved": 0, "dead": 0})
|
|
||||||
oldest_saved: dict[str, str] = {}
|
|
||||||
newest_dead: dict[str, str] = {}
|
|
||||||
oldest_dead: dict[str, str] = {}
|
|
||||||
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
date_match = DATE_RE.search(md.name)
|
|
||||||
if not date_match:
|
|
||||||
continue
|
|
||||||
date = date_match.group(1)
|
|
||||||
month = date[:7]
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
|
|
||||||
conv_key = f"{date} {md.stem}"
|
|
||||||
|
|
||||||
for source, _path in SAVED_RE.findall(text):
|
|
||||||
source = source or "unknown"
|
|
||||||
stats[month][source]["saved"] += 1
|
|
||||||
losses[conv_key]["saved"] += 1
|
|
||||||
if source not in oldest_saved or date < oldest_saved[source]:
|
|
||||||
oldest_saved[source] = date
|
|
||||||
|
|
||||||
for _ref, meta in DEAD_RE.findall(text):
|
|
||||||
source = meta.split(",")[0].strip() or "unknown"
|
|
||||||
stats[month][source]["dead"] += 1
|
|
||||||
losses[conv_key]["dead"] += 1
|
|
||||||
if source not in newest_dead or date > newest_dead[source]:
|
|
||||||
newest_dead[source] = date
|
|
||||||
if source not in oldest_dead or date < oldest_dead[source]:
|
|
||||||
oldest_dead[source] = date
|
|
||||||
|
|
||||||
if not stats:
|
|
||||||
print(f"no dated conversations with images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
sources = sorted({s for m in stats.values() for s in m})
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print("Image outcomes by conversation month")
|
|
||||||
print("=" * 78)
|
|
||||||
header = f"{'month':<9}"
|
|
||||||
for source in sources:
|
|
||||||
header += f" {source[:16]:>16} (saved/dead)"
|
|
||||||
print(header)
|
|
||||||
for month in sorted(stats):
|
|
||||||
row = f"{month:<9}"
|
|
||||||
for source in sources:
|
|
||||||
cell = stats[month].get(source, {"saved": 0, "dead": 0})
|
|
||||||
if cell["saved"] or cell["dead"]:
|
|
||||||
row += f" {cell['saved']:>8} / {cell['dead']:<17}"
|
|
||||||
else:
|
|
||||||
row += f" {'—':>8} {'':<17}"
|
|
||||||
print(row)
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("The age question")
|
|
||||||
print("=" * 78)
|
|
||||||
for source in sources:
|
|
||||||
old_s = oldest_saved.get(source)
|
|
||||||
old_d = oldest_dead.get(source)
|
|
||||||
new_d = newest_dead.get(source)
|
|
||||||
print(f" {source}:")
|
|
||||||
print(f" oldest still downloadable : {old_s or '—'}")
|
|
||||||
print(f" dead range : {old_d or '—'} … {new_d or '—'}")
|
|
||||||
if old_s and new_d and old_s < new_d:
|
|
||||||
print(
|
|
||||||
f" → an asset from {old_s} was captured while one from "
|
|
||||||
f"{new_d} was not."
|
|
||||||
)
|
|
||||||
print(
|
|
||||||
" This does NOT disprove expiry: 'captured' means some "
|
|
||||||
"run got it in time,"
|
|
||||||
)
|
|
||||||
print(
|
|
||||||
" not that it is still on ChatGPT today. Run "
|
|
||||||
"probe_survival_by_month.py to tell"
|
|
||||||
)
|
|
||||||
print(" those apart.")
|
|
||||||
elif old_s and old_d and old_s > old_d:
|
|
||||||
print(
|
|
||||||
f" → everything lost is older than everything captured "
|
|
||||||
f"(cutoff between {old_d} and {old_s}): consistent with expiry."
|
|
||||||
)
|
|
||||||
print()
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print("Conversations that lost images (clustering check)")
|
|
||||||
print("=" * 78)
|
|
||||||
lossy = {k: v for k, v in losses.items() if v["dead"]}
|
|
||||||
for conv, counts in sorted(lossy.items()):
|
|
||||||
print(f" {conv[:66]:<66} saved={counts['saved']:<4} dead={counts['dead']}")
|
|
||||||
total_dead = sum(v["dead"] for v in lossy.values())
|
|
||||||
print()
|
|
||||||
print(f" {len(lossy)} conversation(s) affected, {total_dead} image(s) lost")
|
|
||||||
mixed = [k for k, v in lossy.items() if v["saved"]]
|
|
||||||
if mixed:
|
|
||||||
print(
|
|
||||||
f" {len(mixed)} of them ALSO kept images — same conversation, same "
|
|
||||||
"age, different outcome:"
|
|
||||||
)
|
|
||||||
for conv in sorted(mixed):
|
|
||||||
print(f" {conv[:70]}")
|
|
||||||
print(" → whatever killed these is per-asset, not per-conversation.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,202 +0,0 @@
|
|||||||
"""Do the lost images sit on abandoned conversation branches?
|
|
||||||
|
|
||||||
Established so far: uploads do not expire (36/36 sampled images from
|
|
||||||
2025-09 through 2026-08 are still live), and the loss is per-asset — one
|
|
||||||
conversation on 2026-07-09 kept 38 images and lost 12. So something
|
|
||||||
distinguishes those 12 from their neighbours in the same chat.
|
|
||||||
|
|
||||||
Hypothesis: they are attached to messages that were edited or regenerated.
|
|
||||||
ChatGPT keeps the whole conversation as a tree; editing a message forks it,
|
|
||||||
leaving the superseded messages on an abandoned branch. The exporter walks
|
|
||||||
every node (chatgpt.py:963), so it exports abandoned branches too — and an
|
|
||||||
attachment that is no longer reachable from the live conversation is exactly
|
|
||||||
the kind of thing a backend would garbage-collect.
|
|
||||||
|
|
||||||
The test: fetch the raw conversation, compute the live branch by walking
|
|
||||||
parent links up from ``current_node``, then cross-tabulate every image
|
|
||||||
against (on the live branch?) x (still downloadable?).
|
|
||||||
|
|
||||||
on-branch alive + off-branch dead → confirmed, and it is not data loss:
|
|
||||||
those images belong to messages you
|
|
||||||
replaced.
|
|
||||||
dead on the live branch → hypothesis dead; the images are
|
|
||||||
genuinely missing from live history.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/branch_check.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider # noqa: E402
|
|
||||||
|
|
||||||
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
|
|
||||||
|
|
||||||
|
|
||||||
def find_mixed_conversations(export_dir: Path) -> list[tuple[Path, str, int, int]]:
|
|
||||||
"""Conversations that both kept and lost images — the informative ones."""
|
|
||||||
out = []
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
saved = len(SAVED_RE.findall(text))
|
|
||||||
dead = len(DEAD_RE.findall(text))
|
|
||||||
if not dead:
|
|
||||||
continue
|
|
||||||
match = CONV_ID_RE.search(text)
|
|
||||||
if not match:
|
|
||||||
continue
|
|
||||||
out.append((md, match.group(1), saved, dead))
|
|
||||||
return out
|
|
||||||
|
|
||||||
|
|
||||||
def live_branch(mapping: dict, current_node: str | None) -> set[str]:
|
|
||||||
"""Node IDs reachable by walking parent links up from current_node."""
|
|
||||||
live: set[str] = set()
|
|
||||||
node_id = current_node
|
|
||||||
while node_id and node_id in mapping and node_id not in live:
|
|
||||||
live.add(node_id)
|
|
||||||
node_id = mapping[node_id].get("parent")
|
|
||||||
return live
|
|
||||||
|
|
||||||
|
|
||||||
def image_refs(node: dict) -> list[str]:
|
|
||||||
"""asset_pointer refs on this node's message, if any."""
|
|
||||||
message = node.get("message") or {}
|
|
||||||
content = message.get("content") or {}
|
|
||||||
refs = []
|
|
||||||
if content.get("content_type") == "image_asset_pointer":
|
|
||||||
ref = content.get("asset_pointer")
|
|
||||||
if ref:
|
|
||||||
refs.append(ref)
|
|
||||||
for part in content.get("parts") or []:
|
|
||||||
if isinstance(part, dict) and part.get("content_type") == "image_asset_pointer":
|
|
||||||
ref = part.get("asset_pointer")
|
|
||||||
if ref:
|
|
||||||
refs.append(ref)
|
|
||||||
return refs
|
|
||||||
|
|
||||||
|
|
||||||
def still_live(provider, file_id: str) -> bool | None:
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{file_id}", timeout=30
|
|
||||||
)
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
return None
|
|
||||||
if resp.status_code == 200:
|
|
||||||
return True
|
|
||||||
if resp.status_code == 404:
|
|
||||||
return False
|
|
||||||
return None
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
from src.providers.chatgpt import parse_asset_file_id
|
|
||||||
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
mixed = find_mixed_conversations(export_dir)
|
|
||||||
if not mixed:
|
|
||||||
print(f"no conversations with lost images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
grand = {"on_alive": 0, "on_dead": 0, "off_alive": 0, "off_dead": 0, "unknown": 0}
|
|
||||||
|
|
||||||
for md, conv_id, saved, dead in sorted(mixed, key=lambda r: -r[3]):
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"{md.name}")
|
|
||||||
print(f" {saved} kept, {dead} lost conversation_id={conv_id}")
|
|
||||||
print("=" * 78)
|
|
||||||
|
|
||||||
try:
|
|
||||||
raw = provider.get_conversation(conv_id)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" could not fetch: {e}\n")
|
|
||||||
continue
|
|
||||||
|
|
||||||
mapping = raw.get("mapping") or {}
|
|
||||||
current = raw.get("current_node")
|
|
||||||
live_nodes = live_branch(mapping, current)
|
|
||||||
print(f" mapping nodes: {len(mapping)} live branch: {len(live_nodes)}")
|
|
||||||
if len(live_nodes) < len(mapping):
|
|
||||||
print(
|
|
||||||
f" → {len(mapping) - len(live_nodes)} node(s) are OFF the live "
|
|
||||||
"branch (edited or regenerated messages)"
|
|
||||||
)
|
|
||||||
else:
|
|
||||||
print(" → conversation is linear; nothing was edited or regenerated")
|
|
||||||
print()
|
|
||||||
|
|
||||||
rows = []
|
|
||||||
for node_id, node in mapping.items():
|
|
||||||
on_branch = node_id in live_nodes
|
|
||||||
for ref in image_refs(node):
|
|
||||||
file_id = parse_asset_file_id(ref) or ref
|
|
||||||
alive = still_live(provider, file_id)
|
|
||||||
rows.append((on_branch, alive, file_id))
|
|
||||||
if alive is None:
|
|
||||||
grand["unknown"] += 1
|
|
||||||
else:
|
|
||||||
key = ("on_" if on_branch else "off_") + ("alive" if alive else "dead")
|
|
||||||
grand[key] += 1
|
|
||||||
|
|
||||||
counts = {"on_alive": 0, "on_dead": 0, "off_alive": 0, "off_dead": 0, "unk": 0}
|
|
||||||
for on_branch, alive, _fid in rows:
|
|
||||||
if alive is None:
|
|
||||||
counts["unk"] += 1
|
|
||||||
else:
|
|
||||||
counts[("on_" if on_branch else "off_") + ("alive" if alive else "dead")] += 1
|
|
||||||
|
|
||||||
print(f" {'':<18}{'still live':>12}{'gone':>8}")
|
|
||||||
print(f" {'on live branch':<18}{counts['on_alive']:>12}{counts['on_dead']:>8}")
|
|
||||||
print(f" {'off (abandoned)':<18}{counts['off_alive']:>12}{counts['off_dead']:>8}")
|
|
||||||
if counts["unk"]:
|
|
||||||
print(f" ({counts['unk']} indeterminate)")
|
|
||||||
print()
|
|
||||||
|
|
||||||
for on_branch, alive, fid in sorted(rows, key=lambda r: (r[0], r[1] is not False)):
|
|
||||||
if alive is False:
|
|
||||||
where = "live branch" if on_branch else "ABANDONED branch"
|
|
||||||
print(f" LOST {fid[:40]:<42} {where}")
|
|
||||||
print()
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print("Verdict")
|
|
||||||
print("=" * 78)
|
|
||||||
print(
|
|
||||||
f" on live branch : {grand['on_alive']} live, {grand['on_dead']} gone"
|
|
||||||
)
|
|
||||||
print(
|
|
||||||
f" abandoned : {grand['off_alive']} live, {grand['off_dead']} gone"
|
|
||||||
)
|
|
||||||
if grand["off_dead"] and not grand["on_dead"]:
|
|
||||||
print()
|
|
||||||
print(" Every lost image is on an abandoned branch, and nothing on the")
|
|
||||||
print(" live conversation is missing. These are attachments to messages")
|
|
||||||
print(" you edited or regenerated — ChatGPT collects them once they are")
|
|
||||||
print(" no longer part of the conversation. Your live history is intact,")
|
|
||||||
print(" and no export schedule would have changed this.")
|
|
||||||
elif grand["on_dead"]:
|
|
||||||
print()
|
|
||||||
print(" Images are missing from the LIVE conversation — not explained by")
|
|
||||||
print(" editing. Real loss from current history; worth digging further.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,203 +0,0 @@
|
|||||||
"""The web UI serves images from /backend-api/estuary/content — can we mint that?
|
|
||||||
|
|
||||||
A URL copied from the browser looks like:
|
|
||||||
|
|
||||||
/backend-api/estuary/content?id=file_…&ts=496382&p=fs&cid=1&sig=…&v=0
|
|
||||||
|
|
||||||
Fetched with no cookies it returns 403 {"detail":"File stream access denied."},
|
|
||||||
so the signature is not a bypass — it still rides on the session. Two things
|
|
||||||
follow, and this probe checks both:
|
|
||||||
|
|
||||||
A. What does /files/{id}/download hand back for a file that WORKS? If its
|
|
||||||
download_url is one of these estuary URLs, then that endpoint is the
|
|
||||||
minting step, and a gizmo-scoped file failing there means we are refused
|
|
||||||
at exactly the point the signature would be issued. Then the only hope is
|
|
||||||
a different minting route.
|
|
||||||
|
|
||||||
B. Does the estuary namespace expose one? /backend-api/estuary/* is new to us
|
|
||||||
and was never probed — every earlier attempt used /backend-api/files/*.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/estuary_probe.py
|
|
||||||
|
|
||||||
Optionally pass a browser URL for one of the UNREACHABLE images to check
|
|
||||||
whether the exporter's session can replay it:
|
|
||||||
|
|
||||||
python tools/estuary_probe.py "https://chatgpt.com/backend-api/estuary/content?id=…"
|
|
||||||
"""
|
|
||||||
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from pathlib import Path
|
|
||||||
from urllib.parse import parse_qs, urlparse
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
|
|
||||||
|
|
||||||
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
|
|
||||||
|
|
||||||
def sample_ids(export_dir: Path) -> tuple[str | None, list[str]]:
|
|
||||||
"""(one working file id, all failed file ids)."""
|
|
||||||
working = None
|
|
||||||
failed: list[str] = []
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
for ref, _meta in DEAD_RE.findall(text):
|
|
||||||
fid = parse_asset_file_id(ref)
|
|
||||||
if fid:
|
|
||||||
failed.append(fid)
|
|
||||||
if working is None:
|
|
||||||
for _src, rel in SAVED_RE.findall(text):
|
|
||||||
stem = Path(rel).stem
|
|
||||||
if stem.startswith("file"):
|
|
||||||
working = stem
|
|
||||||
break
|
|
||||||
return working, failed
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
pasted = sys.argv[1] if len(sys.argv) > 1 else None
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
working, failed = sample_ids(export_dir)
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
|
|
||||||
def call(label: str, url: str, headers: dict | None = None) -> None:
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request("GET", url, headers=headers, timeout=30)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {label:<48} error {type(e).__name__}")
|
|
||||||
return
|
|
||||||
ctype = (resp.headers.get("content-type") or "").split(";")[0]
|
|
||||||
if resp.status_code == 200 and ctype.startswith("image/"):
|
|
||||||
print(f" {label:<48} 200 {ctype} {len(resp.content)}B ← IMAGE BYTES")
|
|
||||||
return
|
|
||||||
body = ""
|
|
||||||
try:
|
|
||||||
body = json.dumps(resp.json(), default=str)[:220]
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
body = (resp.text or "")[:160].replace("\n", " ")
|
|
||||||
print(f" {label:<48} {resp.status_code} {ctype}")
|
|
||||||
if body:
|
|
||||||
print(f" {body}")
|
|
||||||
|
|
||||||
# ── A. what a working file's download_url actually looks like ──────────
|
|
||||||
print("=" * 78)
|
|
||||||
print("A. The minting step, on a file that works")
|
|
||||||
print("=" * 78)
|
|
||||||
if working:
|
|
||||||
print(f" working file: {working}")
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{working}/download", timeout=30
|
|
||||||
)
|
|
||||||
body = resp.json() if resp.status_code == 200 else {}
|
|
||||||
url = body.get("download_url") or ""
|
|
||||||
print(f" /files/{{id}}/download → {resp.status_code}")
|
|
||||||
if url:
|
|
||||||
parsed = urlparse(url)
|
|
||||||
print(f" host {parsed.netloc}")
|
|
||||||
print(f" path {parsed.path}")
|
|
||||||
params = parse_qs(parsed.query)
|
|
||||||
print(f" query {sorted(params)}")
|
|
||||||
if "estuary" in parsed.path:
|
|
||||||
print(" → SAME estuary route the browser uses.")
|
|
||||||
print(" So /files/{id}/download IS the minting step, and a")
|
|
||||||
print(" gizmo file's 403 is a refusal to mint. Look for")
|
|
||||||
print(" another minter below.")
|
|
||||||
else:
|
|
||||||
print(" → a different route from the browser's estuary URL;")
|
|
||||||
print(" the UI must mint its URLs somewhere else.")
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" could not check: {e}")
|
|
||||||
else:
|
|
||||||
print(" no working image found on disk to compare against")
|
|
||||||
|
|
||||||
# ── B. does the estuary namespace offer a route for refused files? ─────
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("B. The estuary namespace, on a refused file")
|
|
||||||
print("=" * 78)
|
|
||||||
if not failed:
|
|
||||||
print(" no failed images found")
|
|
||||||
else:
|
|
||||||
# The failures are two different classes and only one is interesting.
|
|
||||||
# A deleted file 404s on /files/{id}; a refused one answers 200. Testing
|
|
||||||
# a deleted file here proves nothing, so pick a refused one.
|
|
||||||
target = None
|
|
||||||
for candidate in failed:
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
meta = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{candidate}", timeout=30
|
|
||||||
)
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
continue
|
|
||||||
if meta.status_code == 200:
|
|
||||||
target = candidate
|
|
||||||
break
|
|
||||||
print(f" (skipping {candidate[:28]}… — deleted, {meta.status_code})")
|
|
||||||
|
|
||||||
if target is None:
|
|
||||||
print(" every failure is a deleted file; nothing in the refused class")
|
|
||||||
return
|
|
||||||
print(f"\n refused file: {target} (exists, download refused)\n")
|
|
||||||
for label, url in [
|
|
||||||
("estuary/files/{id}/download", f"{BASE_URL}/estuary/files/{target}/download"),
|
|
||||||
("estuary/files/{id}", f"{BASE_URL}/estuary/files/{target}"),
|
|
||||||
("estuary/content?id=", f"{BASE_URL}/estuary/content?id={target}"),
|
|
||||||
("estuary/content?id=&p=fs&cid=1&v=0", f"{BASE_URL}/estuary/content?id={target}&p=fs&cid=1&v=0"),
|
|
||||||
("estuary/{id}", f"{BASE_URL}/estuary/{target}"),
|
|
||||||
("estuary/download?id=", f"{BASE_URL}/estuary/download?id={target}"),
|
|
||||||
("files/{id}/download?v=0", f"{BASE_URL}/files/{target}/download?v=0"),
|
|
||||||
# Validation asked only for id, p and ts — never sig. Perhaps the
|
|
||||||
# signature is enforced elsewhere, or not at all for an owner.
|
|
||||||
("estuary/content (no sig)", f"{BASE_URL}/estuary/content?id={target}&p=fs&cid=1&v=0&ts=496382"),
|
|
||||||
]:
|
|
||||||
call(label, url)
|
|
||||||
|
|
||||||
# Does a working file's freshly minted URL serve if we swap in the
|
|
||||||
# refused id? If it does, the signature is not bound to the file.
|
|
||||||
if working:
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{working}/download", timeout=30
|
|
||||||
)
|
|
||||||
minted = (resp.json() or {}).get("download_url") if resp.status_code == 200 else None
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
minted = None
|
|
||||||
if minted:
|
|
||||||
swapped = re.sub(r"id=file_[0-9a-f]+", f"id={target}", minted)
|
|
||||||
print()
|
|
||||||
call("working file's URL, refused id swapped in", swapped)
|
|
||||||
|
|
||||||
# ── C. replay a pasted URL through our session ─────────────────────────
|
|
||||||
if pasted:
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("C. Replaying the pasted URL through the exporter session")
|
|
||||||
print("=" * 78)
|
|
||||||
params = parse_qs(urlparse(pasted).query)
|
|
||||||
pasted_id = (params.get("id") or [""])[0]
|
|
||||||
print(f" id in URL: {pasted_id}")
|
|
||||||
print(f" that id is {'a REFUSED file' if pasted_id in failed else 'NOT one of the refused files'}")
|
|
||||||
call("pasted URL", pasted)
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,185 +0,0 @@
|
|||||||
"""How do you download a gizmo-scoped file?
|
|
||||||
|
|
||||||
metadata_diff.py found the discriminator: every refused image carries
|
|
||||||
``use_case: "gizmo"`` (ChatGPT's name for Projects and custom GPTs), while
|
|
||||||
everything that downloads is ``image_gen`` or ``multimodal``. The files are
|
|
||||||
healthy — ``state: "ready"`` — so this is not loss, it is the plain
|
|
||||||
/files/{id}/download endpoint declining to serve a project-scoped file.
|
|
||||||
|
|
||||||
So find the call that works. Two parts:
|
|
||||||
|
|
||||||
1. Dump the raw conversation node carrying one of these images. The part dict
|
|
||||||
may name the scope the download needs (a gizmo id, a file token, a
|
|
||||||
conversation id) — cheaper than guessing.
|
|
||||||
|
|
||||||
2. Try the plausible variants and print what each returns:
|
|
||||||
?gizmo_id=<each configured project> scope by query param
|
|
||||||
?conversation_id=<the conversation> scope by conversation
|
|
||||||
/gizmos/{gizmo}/files/{id}/download scope by path
|
|
||||||
Referer: the project URL scope by origin
|
|
||||||
|
|
||||||
Anything that returns 200 is the fix, and the exporter can adopt it.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/gizmo_download_probe.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
|
|
||||||
from src.utils import redact_secrets # noqa: E402
|
|
||||||
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
|
|
||||||
|
|
||||||
|
|
||||||
def find_candidates(export_dir: Path) -> list[tuple[str, str, str]]:
|
|
||||||
"""(file_id, conversation_id, export filename) for images that failed."""
|
|
||||||
out = []
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
conv_match = CONV_ID_RE.search(text)
|
|
||||||
if not conv_match:
|
|
||||||
continue
|
|
||||||
for ref, _meta in DEAD_RE.findall(text):
|
|
||||||
file_id = parse_asset_file_id(ref)
|
|
||||||
if file_id:
|
|
||||||
out.append((file_id, conv_match.group(1), md.name))
|
|
||||||
return out
|
|
||||||
|
|
||||||
|
|
||||||
def is_gizmo_scoped(provider, file_id: str) -> bool:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request("GET", f"{BASE_URL}/files/{file_id}", timeout=30)
|
|
||||||
if resp.status_code != 200:
|
|
||||||
return False
|
|
||||||
try:
|
|
||||||
return resp.json().get("use_case") == "gizmo"
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
return False
|
|
||||||
|
|
||||||
|
|
||||||
def dump_conversation_node(provider, conv_id: str, file_id: str) -> str | None:
|
|
||||||
"""Print the raw part carrying this asset. Returns a gizmo id if found."""
|
|
||||||
try:
|
|
||||||
raw = provider.get_conversation(conv_id)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" could not fetch conversation: {e}")
|
|
||||||
return None
|
|
||||||
|
|
||||||
gizmo_id = raw.get("gizmo_id") or raw.get("conversation_template_id")
|
|
||||||
print(f" conversation gizmo_id : {gizmo_id}")
|
|
||||||
print(f" conversation keys : {sorted(raw.keys())}")
|
|
||||||
|
|
||||||
for node_id, node in (raw.get("mapping") or {}).items():
|
|
||||||
message = node.get("message") or {}
|
|
||||||
content = message.get("content") or {}
|
|
||||||
candidates = []
|
|
||||||
if content.get("content_type") == "image_asset_pointer":
|
|
||||||
candidates.append(content)
|
|
||||||
for part in content.get("parts") or []:
|
|
||||||
if isinstance(part, dict):
|
|
||||||
candidates.append(part)
|
|
||||||
for part in candidates:
|
|
||||||
pointer = part.get("asset_pointer") or ""
|
|
||||||
if file_id in pointer:
|
|
||||||
print(f"\n --- raw part on node {node_id} ---")
|
|
||||||
print(json.dumps(redact_secrets(part), indent=4, default=str)[:1800])
|
|
||||||
meta = message.get("metadata") or {}
|
|
||||||
interesting = {
|
|
||||||
k: v for k, v in meta.items()
|
|
||||||
if any(t in k for t in ("gizmo", "file", "attach", "source"))
|
|
||||||
}
|
|
||||||
if interesting:
|
|
||||||
print(f"\n --- message.metadata (filtered) ---")
|
|
||||||
print(json.dumps(redact_secrets(interesting), indent=4, default=str)[:1200])
|
|
||||||
return gizmo_id
|
|
||||||
print(" (asset not found in the conversation mapping)")
|
|
||||||
return gizmo_id
|
|
||||||
|
|
||||||
|
|
||||||
def try_variants(provider, file_id: str, conv_id: str, gizmo_id: str | None) -> None:
|
|
||||||
project_ids = [p for p in provider._project_ids]
|
|
||||||
if gizmo_id and gizmo_id not in project_ids:
|
|
||||||
project_ids.insert(0, gizmo_id)
|
|
||||||
|
|
||||||
base = f"{BASE_URL}/files/{file_id}/download"
|
|
||||||
|
|
||||||
def show(label: str, url: str, headers: dict | None = None) -> None:
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request("GET", url, headers=headers, timeout=30)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {label:<52} error {type(e).__name__}")
|
|
||||||
return
|
|
||||||
note = ""
|
|
||||||
if resp.status_code == 200:
|
|
||||||
try:
|
|
||||||
note = " ← WORKS" if resp.json().get("download_url") else " (200, no download_url)"
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
note = " ← WORKS (non-JSON)"
|
|
||||||
print(f" {label:<52} {resp.status_code}{note}")
|
|
||||||
|
|
||||||
print("\n --- variants ---")
|
|
||||||
show("baseline", base)
|
|
||||||
show("?conversation_id", f"{base}?conversation_id={conv_id}")
|
|
||||||
for pid in project_ids[:6]:
|
|
||||||
show(f"?gizmo_id={pid[:22]}…", f"{base}?gizmo_id={pid}")
|
|
||||||
for pid in project_ids[:2]:
|
|
||||||
show(
|
|
||||||
f"/gizmos/{pid[:16]}…/files/…/download",
|
|
||||||
f"{BASE_URL}/gizmos/{pid}/files/{file_id}/download",
|
|
||||||
)
|
|
||||||
if gizmo_id:
|
|
||||||
show(
|
|
||||||
"Referer: project URL",
|
|
||||||
base,
|
|
||||||
{"Referer": f"https://chatgpt.com/g/{gizmo_id}/project"},
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
candidates = find_candidates(export_dir)
|
|
||||||
if not candidates:
|
|
||||||
print(f"no failed images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
print(f"configured project ids: {len(provider._project_ids)}")
|
|
||||||
|
|
||||||
tested = 0
|
|
||||||
for file_id, conv_id, md_name in candidates:
|
|
||||||
if not is_gizmo_scoped(provider, file_id):
|
|
||||||
continue
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"{file_id}")
|
|
||||||
print(f" from {md_name}")
|
|
||||||
print("=" * 78)
|
|
||||||
gizmo_id = dump_conversation_node(provider, conv_id, file_id)
|
|
||||||
try_variants(provider, file_id, conv_id, gizmo_id)
|
|
||||||
tested += 1
|
|
||||||
if tested >= 2:
|
|
||||||
break
|
|
||||||
|
|
||||||
if not tested:
|
|
||||||
print("no gizmo-scoped failures found — nothing to probe")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,192 +0,0 @@
|
|||||||
"""Can a refused image be fetched by its Library ID instead?
|
|
||||||
|
|
||||||
gizmo_download_probe dumped the message metadata and found what the asset
|
|
||||||
pointer never carried: each attachment has a second identity.
|
|
||||||
|
|
||||||
"id": "file_000000003454722f9481506b96aed510" ← refused
|
|
||||||
"library_file_id": "libfile_4eb82f478fe081919127e2eba9886e86"
|
|
||||||
"source": "local"
|
|
||||||
|
|
||||||
The exporter only ever knew the sediment id from the asset pointer, and asks
|
|
||||||
/files/{sediment_id}/download — which 403s for these. The Library is a
|
|
||||||
separate store with its own ids, so the natural reading is that we are asking
|
|
||||||
for a conversation-scoped copy of a file that now lives in the Library.
|
|
||||||
|
|
||||||
This walks the attachments in message.metadata to pair each refused sediment
|
|
||||||
id with its library_file_id, then tries the endpoints that could serve it.
|
|
||||||
Whatever returns a download_url is what the exporter should use for any
|
|
||||||
attachment carrying a library_file_id.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/library_download_probe.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
|
|
||||||
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
|
|
||||||
|
|
||||||
|
|
||||||
def find_failed(export_dir: Path) -> dict[str, set[str]]:
|
|
||||||
"""conversation_id → {failed sediment file ids}."""
|
|
||||||
out: dict[str, set[str]] = {}
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
conv = CONV_ID_RE.search(text)
|
|
||||||
if not conv:
|
|
||||||
continue
|
|
||||||
ids = {
|
|
||||||
fid
|
|
||||||
for ref, _meta in DEAD_RE.findall(text)
|
|
||||||
if (fid := parse_asset_file_id(ref))
|
|
||||||
}
|
|
||||||
if ids:
|
|
||||||
out.setdefault(conv.group(1), set()).update(ids)
|
|
||||||
return out
|
|
||||||
|
|
||||||
|
|
||||||
def attachment_index(raw: dict) -> dict[str, dict]:
|
|
||||||
"""sediment file id → its attachment record from message.metadata."""
|
|
||||||
index: dict[str, dict] = {}
|
|
||||||
for node in (raw.get("mapping") or {}).values():
|
|
||||||
message = node.get("message") or {}
|
|
||||||
for att in (message.get("metadata") or {}).get("attachments") or []:
|
|
||||||
if isinstance(att, dict) and att.get("id"):
|
|
||||||
index[att["id"]] = att
|
|
||||||
return index
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
failed = find_failed(export_dir)
|
|
||||||
if not failed:
|
|
||||||
print(f"no failed images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
pairs: list[tuple[str, str, str]] = []
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print("Pairing refused assets with their Library IDs")
|
|
||||||
print("=" * 78)
|
|
||||||
for conv_id, file_ids in failed.items():
|
|
||||||
try:
|
|
||||||
raw = provider.get_conversation(conv_id)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {conv_id[:8]}… could not fetch: {e}")
|
|
||||||
continue
|
|
||||||
index = attachment_index(raw)
|
|
||||||
for file_id in sorted(file_ids):
|
|
||||||
att = index.get(file_id)
|
|
||||||
lib = (att or {}).get("library_file_id")
|
|
||||||
source = (att or {}).get("source")
|
|
||||||
print(f" {file_id[:34]}… library={str(lib)[:34]:<36} source={source}")
|
|
||||||
if lib:
|
|
||||||
pairs.append((file_id, lib, conv_id))
|
|
||||||
|
|
||||||
if not pairs:
|
|
||||||
print("\n No refused asset carries a library_file_id — different cause.")
|
|
||||||
return
|
|
||||||
|
|
||||||
file_id, lib_id, conv_id = pairs[0]
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"Endpoint hunt for {lib_id}")
|
|
||||||
print(f" (sediment id {file_id})")
|
|
||||||
print("=" * 78)
|
|
||||||
|
|
||||||
def show(label: str, url: str, method: str = "GET") -> bool:
|
|
||||||
"""Print status AND body. A 200 carrying an error envelope says what
|
|
||||||
the endpoint wants — printing only the keys threw that away."""
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request(method, url, timeout=30)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {label:<50} error {type(e).__name__}")
|
|
||||||
return False
|
|
||||||
|
|
||||||
hit = False
|
|
||||||
try:
|
|
||||||
body = resp.json()
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
preview = (resp.text or "").strip()[:200]
|
|
||||||
if resp.status_code == 200 and resp.content:
|
|
||||||
print(f" {label:<50} 200 non-JSON ({len(resp.content)} bytes) ← BYTES?")
|
|
||||||
return True
|
|
||||||
print(f" {label:<50} {resp.status_code} {preview}")
|
|
||||||
return False
|
|
||||||
|
|
||||||
if isinstance(body, dict) and body.get("download_url"):
|
|
||||||
print(f" {label:<50} {resp.status_code} ← DOWNLOAD URL")
|
|
||||||
print(f" {json.dumps(body, default=str)[:300]}")
|
|
||||||
return True
|
|
||||||
|
|
||||||
print(f" {label:<50} {resp.status_code}")
|
|
||||||
print(f" {json.dumps(body, default=str)[:400]}")
|
|
||||||
return hit
|
|
||||||
|
|
||||||
candidates = [
|
|
||||||
# This one already returns 200 with an error envelope — read it.
|
|
||||||
("/files/{lib}/download", f"{BASE_URL}/files/{lib_id}/download"),
|
|
||||||
("/files/{lib}/download?conversation_id=", f"{BASE_URL}/files/{lib_id}/download?conversation_id={conv_id}"),
|
|
||||||
("/files/{lib}/download (POST)", f"{BASE_URL}/files/{lib_id}/download"),
|
|
||||||
# How does the UI enumerate the Library? A listing tends to reveal both
|
|
||||||
# the id form it uses and the route that serves bytes.
|
|
||||||
("/library/files?limit=3", f"{BASE_URL}/library/files?limit=3"),
|
|
||||||
("/library?limit=3", f"{BASE_URL}/library?limit=3"),
|
|
||||||
("/files?limit=3", f"{BASE_URL}/files?limit=3"),
|
|
||||||
("/my_files?limit=3", f"{BASE_URL}/my_files?limit=3"),
|
|
||||||
# Content routes that serve the asset rather than a signed URL.
|
|
||||||
("/files/{lib}/content", f"{BASE_URL}/files/{lib_id}/content"),
|
|
||||||
("/content?asset_pointer=sediment://{sediment}", f"{BASE_URL}/content?asset_pointer=sediment://{file_id}"),
|
|
||||||
]
|
|
||||||
|
|
||||||
winners = []
|
|
||||||
for label, url in candidates:
|
|
||||||
method = "POST" if "(POST)" in label else "GET"
|
|
||||||
if show(label, url, method):
|
|
||||||
winners.append((label, url))
|
|
||||||
|
|
||||||
# If a second file behaves differently, the first one was the anomaly.
|
|
||||||
if len(pairs) > 1:
|
|
||||||
_other_file, other_lib, _other_conv = pairs[1]
|
|
||||||
print()
|
|
||||||
print(f" --- second sample: {other_lib} ---")
|
|
||||||
show("/files/{lib}/download", f"{BASE_URL}/files/{other_lib}/download")
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("Result")
|
|
||||||
print("=" * 78)
|
|
||||||
if winners:
|
|
||||||
print(" Served by:")
|
|
||||||
for label, url in winners:
|
|
||||||
print(f" {label}")
|
|
||||||
print()
|
|
||||||
print(" → The exporter can pair each asset_pointer with the")
|
|
||||||
print(" library_file_id in message.metadata.attachments and fetch")
|
|
||||||
print(f" {len(pairs)} otherwise-unreachable image(s) this way.")
|
|
||||||
else:
|
|
||||||
print(" None of these served the file. The Library ID is real but the")
|
|
||||||
print(" route is elsewhere — next step is watching what chatgpt.com")
|
|
||||||
print(" itself requests when it renders one of these images.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,176 +0,0 @@
|
|||||||
"""What distinguishes an image ChatGPT refuses to serve from one it serves?
|
|
||||||
|
|
||||||
branch_check.py turned up something the earlier probing missed: of the images
|
|
||||||
that failed to download, only some are actually gone. The rest answer
|
|
||||||
/files/{id} with 200 and full metadata, and refuse only /download. Three
|
|
||||||
distinct states:
|
|
||||||
|
|
||||||
gone 404 on /files/{id} — deleted, unrecoverable
|
|
||||||
refused 200 on /files/{id}, 403 on download — exists, will not serve
|
|
||||||
working 200 on both — fine
|
|
||||||
|
|
||||||
"refused" is the interesting one, because those files still exist and may be
|
|
||||||
recoverable. Since the metadata comes back for them, the discriminator can be
|
|
||||||
read straight off: fetch metadata for every file in each state and compare the
|
|
||||||
fields. A field that is constant within "refused" and different in "working"
|
|
||||||
is the cause.
|
|
||||||
|
|
||||||
Also re-checks /download now. If a file that was refused during the export
|
|
||||||
serves today, the failure was transient and a retry pass recovers it — a very
|
|
||||||
different fix from anything permanent.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/metadata_diff.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from collections import defaultdict
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider # noqa: E402
|
|
||||||
|
|
||||||
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
|
|
||||||
WORKING_SAMPLE = 6
|
|
||||||
|
|
||||||
|
|
||||||
def collect(export_dir: Path) -> tuple[list[str], list[str]]:
|
|
||||||
"""(ids that failed to download, ids that downloaded fine)."""
|
|
||||||
from src.providers.chatgpt import parse_asset_file_id
|
|
||||||
|
|
||||||
failed: list[str] = []
|
|
||||||
working: list[str] = []
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
for ref, _meta in DEAD_RE.findall(text):
|
|
||||||
file_id = parse_asset_file_id(ref)
|
|
||||||
if file_id:
|
|
||||||
failed.append(file_id)
|
|
||||||
for _source, rel in SAVED_RE.findall(text):
|
|
||||||
stem = Path(rel).stem
|
|
||||||
if stem.startswith("file"):
|
|
||||||
working.append(stem)
|
|
||||||
return failed, working
|
|
||||||
|
|
||||||
|
|
||||||
def probe(provider, file_id: str) -> tuple[str, dict]:
|
|
||||||
"""(state, metadata) for one file."""
|
|
||||||
provider._pace()
|
|
||||||
meta_resp = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{file_id}", timeout=30
|
|
||||||
)
|
|
||||||
if meta_resp.status_code == 404:
|
|
||||||
return "gone", {}
|
|
||||||
if meta_resp.status_code != 200:
|
|
||||||
return f"meta-{meta_resp.status_code}", {}
|
|
||||||
|
|
||||||
try:
|
|
||||||
meta = meta_resp.json()
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
meta = {}
|
|
||||||
|
|
||||||
provider._pace()
|
|
||||||
dl_resp = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{file_id}/download", timeout=30
|
|
||||||
)
|
|
||||||
state = "working" if dl_resp.status_code == 200 else f"refused-{dl_resp.status_code}"
|
|
||||||
return state, meta
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
failed, working = collect(export_dir)
|
|
||||||
if not failed:
|
|
||||||
print(f"no failed images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
# Sample working files across the archive rather than all of them.
|
|
||||||
step = max(1, len(working) // WORKING_SAMPLE)
|
|
||||||
working_sample = working[::step][:WORKING_SAMPLE]
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
by_state: dict[str, list[tuple[str, dict]]] = defaultdict(list)
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"Probing {len(failed)} failed + {len(working_sample)} working files")
|
|
||||||
print("=" * 78)
|
|
||||||
for file_id in failed + working_sample:
|
|
||||||
state, meta = probe(provider, file_id)
|
|
||||||
by_state[state].append((file_id, meta))
|
|
||||||
print(f" {file_id[:40]:<42} {state}")
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("States")
|
|
||||||
print("=" * 78)
|
|
||||||
for state, entries in sorted(by_state.items()):
|
|
||||||
print(f" {state:<16} {len(entries)}")
|
|
||||||
print()
|
|
||||||
|
|
||||||
if any(s.startswith("refused") for s in by_state) is False:
|
|
||||||
print(" No file is in the 'refused' state right now.")
|
|
||||||
print(" → Every previously-failed file that still exists now serves.")
|
|
||||||
print(" The export-time 403s were TRANSIENT; a retry pass recovers them.")
|
|
||||||
print()
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print("Metadata field comparison")
|
|
||||||
print("=" * 78)
|
|
||||||
states = [s for s in by_state if by_state[s] and any(m for _i, m in by_state[s])]
|
|
||||||
all_fields: set[str] = set()
|
|
||||||
for state in states:
|
|
||||||
for _fid, meta in by_state[state]:
|
|
||||||
all_fields.update(meta.keys())
|
|
||||||
|
|
||||||
for field in sorted(all_fields):
|
|
||||||
line = f" {field:<24}"
|
|
||||||
distinct_per_state = []
|
|
||||||
for state in sorted(states):
|
|
||||||
values = {
|
|
||||||
json.dumps(meta.get(field), default=str)[:28]
|
|
||||||
for _fid, meta in by_state[state]
|
|
||||||
if meta
|
|
||||||
}
|
|
||||||
shown = ", ".join(sorted(values)[:3])
|
|
||||||
if len(values) > 3:
|
|
||||||
shown += f" (+{len(values) - 3} more)"
|
|
||||||
distinct_per_state.append((state, values, shown))
|
|
||||||
line += f" [{state}] {shown}"
|
|
||||||
# Flag fields that cleanly separate the states.
|
|
||||||
value_sets = [v for _s, v, _sh in distinct_per_state]
|
|
||||||
if len(value_sets) > 1 and all(
|
|
||||||
not (a & b) for i, a in enumerate(value_sets) for b in value_sets[i + 1:]
|
|
||||||
):
|
|
||||||
line += " ← DISCRIMINATOR"
|
|
||||||
print(line)
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("One full record per state")
|
|
||||||
print("=" * 78)
|
|
||||||
for state in sorted(states):
|
|
||||||
fid, meta = next(((f, m) for f, m in by_state[state] if m), (None, None))
|
|
||||||
if not meta:
|
|
||||||
continue
|
|
||||||
print(f" --- {state} ({fid}) ---")
|
|
||||||
for key, value in sorted(meta.items()):
|
|
||||||
print(f" {key:<24} {json.dumps(value, default=str)[:90]}")
|
|
||||||
print()
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,192 +0,0 @@
|
|||||||
"""The browser mints signatures for files we get 403 on. What does it send that we don't?
|
|
||||||
|
|
||||||
Established:
|
|
||||||
* The images DO render in ChatGPT, so a valid signature exists for them.
|
|
||||||
* /files/{id}/download is the minting endpoint — a working file's
|
|
||||||
download_url is the same estuary/content URL the browser uses.
|
|
||||||
* Signatures are per-file: swapping an id into a working URL gives
|
|
||||||
"Invalid signature or expired URL".
|
|
||||||
|
|
||||||
So the difference is in the request, not the route. Our call sends an
|
|
||||||
Authorization: Bearer header plus session cookies; a browser leans on cookies
|
|
||||||
and adds client headers (oai-device-id, oai-client-version, a conversation
|
|
||||||
Referer) that a gated endpoint may check. This walks those variants against a
|
|
||||||
file that is currently refused, and prints which combination mints a URL.
|
|
||||||
|
|
||||||
Any 200 carrying a download_url is the answer, and the exporter adopts those
|
|
||||||
headers.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/mint_headers_probe.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
import uuid
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider, parse_asset_file_id # noqa: E402
|
|
||||||
|
|
||||||
DEAD_RE = re.compile(r"🖼️ \*\*Image attached\*\* — `([^`]+)`\s*\(([^)]*)\)")
|
|
||||||
CONV_ID_RE = re.compile(r"^conversation_id:\s*(\S+)", re.MULTILINE)
|
|
||||||
|
|
||||||
|
|
||||||
def find_failed(export_dir: Path) -> list[tuple[str, str]]:
|
|
||||||
"""(file_id, conversation_id) for every image that failed to download."""
|
|
||||||
out = []
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
conv = CONV_ID_RE.search(text)
|
|
||||||
if not conv:
|
|
||||||
continue
|
|
||||||
for ref, _meta in DEAD_RE.findall(text):
|
|
||||||
fid = parse_asset_file_id(ref)
|
|
||||||
if fid:
|
|
||||||
out.append((fid, conv.group(1)))
|
|
||||||
return out
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
failed = find_failed(export_dir)
|
|
||||||
if not failed:
|
|
||||||
print(f"no failed images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
session = provider._session
|
|
||||||
|
|
||||||
# Pick one that still exists — a deleted file 403s no matter what we send.
|
|
||||||
target = conv_id = None
|
|
||||||
for fid, cid in failed:
|
|
||||||
provider._pace()
|
|
||||||
meta = session.request("GET", f"{BASE_URL}/files/{fid}", timeout=30)
|
|
||||||
if meta.status_code == 200:
|
|
||||||
target, conv_id = fid, cid
|
|
||||||
break
|
|
||||||
if not target:
|
|
||||||
print("every failure is a deleted file — nothing to test")
|
|
||||||
return
|
|
||||||
|
|
||||||
url = f"{BASE_URL}/files/{target}/download"
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"Minting attempts for {target}")
|
|
||||||
print(f" conversation {conv_id}")
|
|
||||||
print("=" * 78)
|
|
||||||
|
|
||||||
def attempt(
|
|
||||||
label: str,
|
|
||||||
*,
|
|
||||||
add: dict | None = None,
|
|
||||||
drop: tuple[str, ...] = (),
|
|
||||||
target_url: str | None = None,
|
|
||||||
) -> bool:
|
|
||||||
removed = {}
|
|
||||||
for header in drop:
|
|
||||||
for key in list(session.headers.keys()):
|
|
||||||
if key.lower() == header.lower():
|
|
||||||
removed[key] = session.headers.pop(key)
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = session.request(
|
|
||||||
"GET", target_url or url, headers=add or None, timeout=30
|
|
||||||
)
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {label:<46} error {type(e).__name__}")
|
|
||||||
return False
|
|
||||||
finally:
|
|
||||||
session.headers.update(removed)
|
|
||||||
|
|
||||||
try:
|
|
||||||
body = resp.json()
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
body = {}
|
|
||||||
if resp.status_code == 200 and isinstance(body, dict) and body.get("download_url"):
|
|
||||||
print(f" {label:<46} 200 ← MINTED")
|
|
||||||
print(f" {str(body['download_url'])[:110]}")
|
|
||||||
return True
|
|
||||||
detail = body.get("detail") if isinstance(body, dict) else None
|
|
||||||
print(f" {label:<46} {resp.status_code} {json.dumps(detail, default=str)[:60] if detail else ''}")
|
|
||||||
return False
|
|
||||||
|
|
||||||
device_id = str(uuid.uuid4())
|
|
||||||
winners = []
|
|
||||||
|
|
||||||
checks = [
|
|
||||||
("baseline (bearer + cookies)", {}, ()),
|
|
||||||
("no Authorization (cookies only)", {}, ("Authorization",)),
|
|
||||||
("Referer: the conversation", {"Referer": f"https://chatgpt.com/c/{conv_id}"}, ()),
|
|
||||||
("oai-device-id", {"oai-device-id": device_id}, ()),
|
|
||||||
("oai-language", {"oai-language": "en-US"}, ()),
|
|
||||||
(
|
|
||||||
"browser-ish set",
|
|
||||||
{
|
|
||||||
"Referer": f"https://chatgpt.com/c/{conv_id}",
|
|
||||||
"oai-device-id": device_id,
|
|
||||||
"oai-language": "en-US",
|
|
||||||
"sec-fetch-site": "same-origin",
|
|
||||||
"sec-fetch-mode": "cors",
|
|
||||||
"sec-fetch-dest": "empty",
|
|
||||||
},
|
|
||||||
(),
|
|
||||||
),
|
|
||||||
(
|
|
||||||
"browser-ish, no Authorization",
|
|
||||||
{
|
|
||||||
"Referer": f"https://chatgpt.com/c/{conv_id}",
|
|
||||||
"oai-device-id": device_id,
|
|
||||||
"oai-language": "en-US",
|
|
||||||
},
|
|
||||||
("Authorization",),
|
|
||||||
),
|
|
||||||
("Accept: image/*", {"Accept": "image/avif,image/webp,*/*"}, ()),
|
|
||||||
]
|
|
||||||
|
|
||||||
for label, add, drop in checks:
|
|
||||||
if attempt(label, add=add, drop=drop):
|
|
||||||
winners.append(label)
|
|
||||||
|
|
||||||
# Same call with the conversation named in the query string.
|
|
||||||
if attempt(
|
|
||||||
"?conversation_id=",
|
|
||||||
target_url=f"{url}?conversation_id={conv_id}",
|
|
||||||
):
|
|
||||||
winners.append("?conversation_id=")
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("Result")
|
|
||||||
print("=" * 78)
|
|
||||||
if winners:
|
|
||||||
print(" Minted by:")
|
|
||||||
for label in winners:
|
|
||||||
print(f" {label}")
|
|
||||||
print()
|
|
||||||
print(" → Adopt those headers in ChatGPTProvider and the 11 refused")
|
|
||||||
print(" images become downloadable.")
|
|
||||||
else:
|
|
||||||
print(" No header combination minted a URL.")
|
|
||||||
print()
|
|
||||||
print(" The browser is getting its signature some other way. Find it:")
|
|
||||||
print(" 1. Open the conversation in ChatGPT with DevTools → Network.")
|
|
||||||
print(" 2. Press Ctrl+F (search inside requests) and search for")
|
|
||||||
print(f" {target[5:22]}")
|
|
||||||
print(" 3. The hit that is NOT the estuary/content request is the")
|
|
||||||
print(" call that mints the signature — send me its URL, method,")
|
|
||||||
print(" and (if POST) its request body.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,153 +0,0 @@
|
|||||||
"""Are already-downloaded images still alive on ChatGPT, or did we just catch them in time?
|
|
||||||
|
|
||||||
analyze_media_age.py cannot answer this. It reports whether an image was ever
|
|
||||||
captured — and an image downloaded by a run back in June looks "saved" forever
|
|
||||||
after, whether or not ChatGPT still holds it today. So a month with zero losses
|
|
||||||
proves the exports were timely, not that the assets survived.
|
|
||||||
|
|
||||||
This closes that gap: take images that ARE on disk, grouped by the month of
|
|
||||||
their conversation, and ask the API whether each still exists right now.
|
|
||||||
|
|
||||||
Old months still 200 → assets do not expire with age. Export cadence is not
|
|
||||||
what saved them, and shortening it would not have
|
|
||||||
saved the July losses either.
|
|
||||||
Old months now 404 → assets DO die with age; the old months are on disk
|
|
||||||
only because a run reached them in time. Cadence is
|
|
||||||
the whole ballgame, and the July losses are the first
|
|
||||||
ones we simply arrived too late for.
|
|
||||||
|
|
||||||
Liveness is checked with /files/{id}, which answers a missing record with a
|
|
||||||
clean 404 (verified 2026-08-17); /files/{id}/download reports the same state
|
|
||||||
as a misleading 403.
|
|
||||||
|
|
||||||
Run from the project root:
|
|
||||||
|
|
||||||
python tools/probe_survival_by_month.py [samples_per_month]
|
|
||||||
"""
|
|
||||||
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
|
||||||
from collections import defaultdict
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
from src.providers.chatgpt import BASE_URL, ChatGPTProvider # noqa: E402
|
|
||||||
|
|
||||||
SAVED_RE = re.compile(r"!\[([^\]]*)\]\((media/[^)]+)\)")
|
|
||||||
DATE_RE = re.compile(r"(\d{4}-\d{2}-\d{2})_")
|
|
||||||
|
|
||||||
DEFAULT_SAMPLES = 4
|
|
||||||
|
|
||||||
|
|
||||||
def collect_saved_ids(export_dir: Path) -> dict[str, list[tuple[str, str]]]:
|
|
||||||
"""month → [(file_id, source)] for images already downloaded to disk."""
|
|
||||||
by_month: dict[str, list[tuple[str, str]]] = defaultdict(list)
|
|
||||||
for md in export_dir.rglob("*.md"):
|
|
||||||
date_match = DATE_RE.search(md.name)
|
|
||||||
if not date_match:
|
|
||||||
continue
|
|
||||||
month = date_match.group(1)[:7]
|
|
||||||
try:
|
|
||||||
text = md.read_text(encoding="utf-8", errors="replace")
|
|
||||||
except OSError:
|
|
||||||
continue
|
|
||||||
for source, rel_path in SAVED_RE.findall(text):
|
|
||||||
file_id = Path(rel_path).stem
|
|
||||||
if file_id.startswith("file"):
|
|
||||||
by_month[month].append((file_id, source or "unknown"))
|
|
||||||
return by_month
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
samples = DEFAULT_SAMPLES
|
|
||||||
if len(sys.argv) > 1:
|
|
||||||
try:
|
|
||||||
samples = max(1, int(sys.argv[1]))
|
|
||||||
except ValueError:
|
|
||||||
pass
|
|
||||||
|
|
||||||
export_dir = Path(os.getenv("EXPORT_DIR", "./exports")).expanduser()
|
|
||||||
by_month = collect_saved_ids(export_dir)
|
|
||||||
if not by_month:
|
|
||||||
print(f"no downloaded images found under {export_dir}")
|
|
||||||
return
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"Are on-disk images still live on ChatGPT? ({samples} sampled per month)")
|
|
||||||
print("=" * 78)
|
|
||||||
print(f"{'month':<9} {'alive':>6} {'gone':>6} {'err':>5} sampled sources")
|
|
||||||
|
|
||||||
totals = {"alive": 0, "gone": 0, "err": 0}
|
|
||||||
verdict_rows: list[tuple[str, int, int]] = []
|
|
||||||
|
|
||||||
for month in sorted(by_month):
|
|
||||||
entries = by_month[month]
|
|
||||||
# Spread the sample across the month rather than taking the first few
|
|
||||||
# from one conversation.
|
|
||||||
step = max(1, len(entries) // samples)
|
|
||||||
picked = entries[::step][:samples]
|
|
||||||
|
|
||||||
alive = gone = err = 0
|
|
||||||
sources: list[str] = []
|
|
||||||
for file_id, source in picked:
|
|
||||||
sources.append(source[:4])
|
|
||||||
try:
|
|
||||||
provider._pace()
|
|
||||||
resp = provider._session.request(
|
|
||||||
"GET", f"{BASE_URL}/files/{file_id}", timeout=30
|
|
||||||
)
|
|
||||||
except Exception: # noqa: BLE001 - diagnostic
|
|
||||||
err += 1
|
|
||||||
continue
|
|
||||||
if resp.status_code == 200:
|
|
||||||
alive += 1
|
|
||||||
elif resp.status_code == 404:
|
|
||||||
gone += 1
|
|
||||||
else:
|
|
||||||
err += 1
|
|
||||||
|
|
||||||
totals["alive"] += alive
|
|
||||||
totals["gone"] += gone
|
|
||||||
totals["err"] += err
|
|
||||||
verdict_rows.append((month, alive, gone))
|
|
||||||
print(
|
|
||||||
f"{month:<9} {alive:>6} {gone:>6} {err:>5} "
|
|
||||||
f"{','.join(sources)} (of {len(entries)} on disk)"
|
|
||||||
)
|
|
||||||
|
|
||||||
print()
|
|
||||||
print(f"total: {totals['alive']} alive, {totals['gone']} gone, {totals['err']} error")
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("Verdict")
|
|
||||||
print("=" * 78)
|
|
||||||
|
|
||||||
old_gone = [m for m, _a, g in verdict_rows[:-2] if g]
|
|
||||||
old_alive = [m for m, a, _g in verdict_rows[:-2] if a]
|
|
||||||
|
|
||||||
if old_gone and not old_alive:
|
|
||||||
print(" Every sampled older image is GONE from ChatGPT.")
|
|
||||||
print(" → Uploads expire. Your old exports survive only because a run")
|
|
||||||
print(" reached them in time. Export cadence directly determines")
|
|
||||||
print(" what you can still save.")
|
|
||||||
elif old_alive and not old_gone:
|
|
||||||
print(" Every sampled older image is STILL LIVE on ChatGPT.")
|
|
||||||
print(" → Uploads do not expire with age. The July losses are")
|
|
||||||
print(" something else, and exporting sooner would not have")
|
|
||||||
print(" prevented them.")
|
|
||||||
else:
|
|
||||||
print(" Mixed: some older images alive, some gone.")
|
|
||||||
print(" → Not a clean expiry. Compare the 'sampled sources' column and")
|
|
||||||
print(" the per-month rates above; the cause is likely per-asset.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
@@ -1,146 +0,0 @@
|
|||||||
"""Can the exporter reuse an image URL copied from the browser?
|
|
||||||
|
|
||||||
The endpoint hunt bottomed out: the Library id is rejected outright
|
|
||||||
(file_not_found from GetDownloadLinkError), and /backend-api/content wants
|
|
||||||
signed query params (id, ts, p, and in practice a signature) that cannot be
|
|
||||||
guessed. The web UI has those values, so the remaining question is whether a
|
|
||||||
URL taken from the UI works outside the browser.
|
|
||||||
|
|
||||||
Paste the request URL DevTools shows for one of the unreachable images:
|
|
||||||
|
|
||||||
python tools/try_pasted_url.py "https://chatgpt.com/backend-api/content?id=…&ts=…&p=…&sig=…"
|
|
||||||
|
|
||||||
It fetches that URL three ways, and the pattern of results says what to build:
|
|
||||||
|
|
||||||
works with our session, not bare → the signature is fine but the request
|
|
||||||
needs auth; the exporter can mint these
|
|
||||||
itself if we find what returns them.
|
|
||||||
works bare too → the URL is self-authenticating; whatever
|
|
||||||
produced it is the endpoint we need.
|
|
||||||
works in neither → the URL is bound to the browser session
|
|
||||||
(or already expired — check ts), so the
|
|
||||||
exporter cannot reuse it as-is.
|
|
||||||
|
|
||||||
Nothing is written to disk unless --save is passed.
|
|
||||||
"""
|
|
||||||
|
|
||||||
import sys
|
|
||||||
from pathlib import Path
|
|
||||||
from urllib.parse import parse_qs, urlparse
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
|
||||||
|
|
||||||
from dotenv import load_dotenv
|
|
||||||
|
|
||||||
load_dotenv()
|
|
||||||
|
|
||||||
|
|
||||||
def describe(url: str) -> None:
|
|
||||||
parsed = urlparse(url)
|
|
||||||
print(f" host : {parsed.netloc}")
|
|
||||||
print(f" path : {parsed.path}")
|
|
||||||
params = parse_qs(parsed.query)
|
|
||||||
print(" query :")
|
|
||||||
for key, values in params.items():
|
|
||||||
value = values[0] if values else ""
|
|
||||||
shown = value if len(value) <= 48 else f"{value[:45]}…"
|
|
||||||
print(f" {key:<12} {shown}")
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
args = [a for a in sys.argv[1:] if a != "--save"]
|
|
||||||
save = "--save" in sys.argv
|
|
||||||
if not args:
|
|
||||||
print(__doc__)
|
|
||||||
return
|
|
||||||
url = args[0]
|
|
||||||
|
|
||||||
print("=" * 78)
|
|
||||||
print("The URL")
|
|
||||||
print("=" * 78)
|
|
||||||
describe(url)
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("Fetch attempts")
|
|
||||||
print("=" * 78)
|
|
||||||
|
|
||||||
def report(label: str, resp) -> bytes | None:
|
|
||||||
ctype = (resp.headers.get("content-type") or "").split(";")[0]
|
|
||||||
size = len(resp.content or b"")
|
|
||||||
verdict = ""
|
|
||||||
if resp.status_code == 200 and ctype.startswith("image/"):
|
|
||||||
verdict = " ← IMAGE BYTES"
|
|
||||||
elif resp.status_code == 200:
|
|
||||||
verdict = f" 200 but {ctype or 'unknown type'}"
|
|
||||||
print(f" {label:<34} {resp.status_code} {ctype} {size}B{verdict}")
|
|
||||||
if resp.status_code == 200 and ctype.startswith("image/"):
|
|
||||||
return resp.content
|
|
||||||
if resp.status_code != 200:
|
|
||||||
preview = (resp.text or "")[:160].replace("\n", " ")
|
|
||||||
if preview:
|
|
||||||
print(f" {preview}")
|
|
||||||
return None
|
|
||||||
|
|
||||||
image: bytes | None = None
|
|
||||||
|
|
||||||
# 1. Bare request, no cookies, no auth headers.
|
|
||||||
try:
|
|
||||||
from curl_cffi import requests as curl_requests
|
|
||||||
|
|
||||||
bare = curl_requests.Session(impersonate="chrome120")
|
|
||||||
image = report("bare (no auth)", bare.get(url, timeout=30)) or image
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {'bare (no auth)':<34} error {type(e).__name__}: {e}")
|
|
||||||
|
|
||||||
# 2. Through the exporter's authenticated session.
|
|
||||||
try:
|
|
||||||
from src.providers.chatgpt import ChatGPTProvider
|
|
||||||
|
|
||||||
provider = ChatGPTProvider()
|
|
||||||
provider._pace()
|
|
||||||
image = report(
|
|
||||||
"exporter session", provider._session.request("GET", url, timeout=30)
|
|
||||||
) or image
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {'exporter session':<34} error {type(e).__name__}: {e}")
|
|
||||||
provider = None
|
|
||||||
|
|
||||||
# 3. Authenticated session without the Authorization header, in case the
|
|
||||||
# signature and the bearer token conflict.
|
|
||||||
if provider is not None:
|
|
||||||
try:
|
|
||||||
saved = provider._session.headers.pop("Authorization", None)
|
|
||||||
provider._pace()
|
|
||||||
image = report(
|
|
||||||
"session, no Authorization",
|
|
||||||
provider._session.request("GET", url, timeout=30),
|
|
||||||
) or image
|
|
||||||
if saved:
|
|
||||||
provider._session.headers["Authorization"] = saved
|
|
||||||
except Exception as e: # noqa: BLE001 - diagnostic
|
|
||||||
print(f" {'session, no Authorization':<34} error {type(e).__name__}")
|
|
||||||
|
|
||||||
print()
|
|
||||||
print("=" * 78)
|
|
||||||
print("Verdict")
|
|
||||||
print("=" * 78)
|
|
||||||
if image:
|
|
||||||
print(f" Served {len(image)} bytes of image data.")
|
|
||||||
print(" → The exporter CAN fetch these. Next: find what mints the")
|
|
||||||
print(" signed params, so it can build the URL itself.")
|
|
||||||
if save:
|
|
||||||
out = Path("pasted_url_result.bin")
|
|
||||||
out.write_bytes(image)
|
|
||||||
print(f" Saved to {out}")
|
|
||||||
else:
|
|
||||||
print(" (pass --save to write the bytes out)")
|
|
||||||
else:
|
|
||||||
print(" No attempt returned image bytes.")
|
|
||||||
print(" → Either the URL is bound to the browser session, or its ts has")
|
|
||||||
print(" expired. Re-copy a fresh URL and retry once; if it still")
|
|
||||||
print(" fails, these images are not reachable programmatically.")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
Reference in New Issue
Block a user