Commit Graph
6 Commits
Author SHA1 Message Date
JesseMarkowitz 03646009b9 tools: probe how to download a gizmo-scoped file
metadata_diff found the discriminator. Every refused image carries
use_case="gizmo" — ChatGPT's name for Projects and custom GPTs — while
everything that downloads is image_gen or multimodal. All are state=ready,
so nothing is damaged: the plain /files/{id}/download endpoint just will
not serve a project-scoped file.

Second clue: all 11 were created 2026-07-14 within minutes of each other,
yet appear in conversations dated 2025-11 through 2026-08, several
predating their own creation_time. Something on July 14 re-created them as
gizmo-scoped copies and repointed the conversations — which is also why
the losses start in July. Not a policy change, an event.

gizmo_download_probe.py dumps the raw conversation part carrying one of
these assets (it may simply name the scope the download wants) and then
tries the plausible calls — gizmo_id/conversation_id query params, a
/gizmos/{id}/files path, a project Referer. A 200 is the fix.
2026-08-17 09:22:48 -04:00
JesseMarkowitz a3ac279e39 tools: find what separates a refused image from a served one
branch_check killed the abandoned-branch hypothesis — every lost image is
on the live branch, and the 3 images that do sit on abandoned branches are
alive. But it turned up something the earlier probing missed by sampling
four IDs from one family and generalising: of the 19 failures, only 7 are
actually gone. The other 12 answer /files/{id} with 200 and full metadata
and refuse only /download. They exist, and may be recoverable.

Three states, then: gone (404), refused (200 meta + 403 download), working
(200 both). Since metadata comes back for the refused ones, the
discriminator can be read straight off — fetch it for every file in each
state and compare fields, flagging any field whose values never overlap
between states.

Also re-checks /download now: if a file refused during the export serves
today, those 403s were transient and a retry pass recovers them, which is
a completely different fix from anything permanent.
2026-08-17 09:18:32 -04:00
JesseMarkowitz 021a76c628 tools: test whether lost images sit on abandoned branches
Expiry is now ruled out: 36/36 sampled images from 2025-09 through
2026-08 are still live server-side, so nothing dies of age and export
cadence is not the variable. The loss is per-asset — 2026-07-09 kept 38
images and lost 12 in one conversation.

Next hypothesis: those images belong to messages that were edited or
regenerated. ChatGPT stores a conversation as a tree and editing forks it,
stranding the superseded messages on an abandoned branch. The exporter
walks every node (chatgpt.py:963), so it exports those branches too — and
an attachment unreachable from live history is a natural GC target.

branch_check.py fetches the raw conversation, derives the live branch by
walking parent links up from current_node, and cross-tabulates every image
against (on the live branch?) x (still downloadable?). If the lost images
are all off-branch and nothing on the live branch is missing, this is not
data loss at all — it is attachments to messages that were replaced.
2026-08-17 08:58:20 -04:00
JesseMarkowitz b9e8896b33 tools: distinguish "captured in time" from "still alive"
analyze_media_age.py claimed age was ruled out because an image from
2025-11 was saved while one from 2026-08 was not. That conclusion does not
follow. "Saved" means some earlier run downloaded it, not that ChatGPT
still holds it — an image captured in June looks saved forever after, even
if it died in July. A month with no losses shows the exports were timely,
not that the assets survived.

- probe_survival_by_month.py: sample images already on disk, grouped by
  their conversation's month, and ask /files/{id} whether each still
  exists today. Old months still 200 → no expiry, and cadence did not save
  them. Old months now 404 → uploads do expire and cadence is the whole
  ballgame.
- analyze_media_age.py: stop asserting the unsupported verdict; say what
  the number does and does not show, and point at the probe.

One signal there is immune to the confound and survives: 2026-07-09 kept
38 images and lost 12. Same conversation, same day, opposite outcomes —
no retention policy does that, so at least part of this is per-asset.
2026-08-17 08:53:54 -04:00
JesseMarkowitz 04191eed8c tools: analyze whether media loss is age-based
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
(![source](media/…) when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.

Offline, no token, no API calls.
2026-08-17 08:45:21 -04:00
JesseMarkowitz f40b25001a tools: add one-off probe for the media 403s
The improved error reporting landed, and the answer it produced is
{"detail":"Forbidden"} — generic, no reason. The cause has to be narrowed
by experiment instead, so collect the experiments in one script:

- classify the failed assets offline from the exported placeholders
  (user_upload vs model_generated) — no API call needed
- probe a known-good asset in the same session as a control, to rule the
  session in or out
- retry the 403 with ChatGPT-Account-Id, which the exporter never sends
  and which workspace-scoped resources can require
- compare /files/{id} against /files/{id}/download

Temporary: delete once the cause is known, or fold into `doctor` if the
check earns a permanent home.
2026-08-17 08:08:01 -04:00