Commit Graph
11 Commits
Author SHA1 Message Date
JesseMarkowitz b1da1df986 tools: probe a refused file, not a deleted one
Confirmed by the last run: a working file's download_url is
chatgpt.com/backend-api/estuary/content?cid&id&p&sig&ts&v — the same route
the browser uses. So /files/{id}/download is the minting endpoint, and a
gizmo file's 403 is a refusal to mint the signature. That is why no amount
of scoping helped; we were turned away at the only door that issues them.
Nothing else in the estuary namespace serves files: 404 across the board.

But section B tested file_000000001b3071f5…, one of the seven DELETED
files, because it took the first failure without checking its class. Those
results say nothing about the eleven recoverable ones.

- Pick the target by probing /files/{id} and taking one that answers 200,
  skipping (and naming) the deleted ones.
- Add two experiments: estuary/content with a ts but no sig, since
  validation asked only for id/p/ts and never sig; and a working file's
  freshly minted URL with the refused id swapped in, which shows whether
  the signature is bound to the file.
2026-08-17 10:57:52 -04:00
JesseMarkowitz 726f57bdf9 tools: probe the estuary namespace
A URL copied from the browser gave us the route we never tried:

    /backend-api/estuary/content?id=…&ts=496382&p=fs&cid=1&sig=…&v=0

Fetched with no cookies it returns 403 {"detail":"File stream access
denied."}, so the signature rides on the session rather than replacing it.
Every earlier probe lived under /backend-api/files/*; estuary/* is new
ground.

estuary_probe.py checks two things:

A. What /files/{id}/download hands back for a file that works. If its
   download_url is an estuary URL, that endpoint is the minting step, and a
   gizmo file's 403 is a refusal to mint — which is why no amount of
   scoping helped.
B. Whether the estuary namespace exposes a route that serves a refused
   file directly.

It also replays a pasted URL through the exporter's session, and says
plainly whether the id in that URL is one of the refused files — the two
captured so far were working images, so they showed the shape without
telling us whether the broken ones have a URL at all.
2026-08-17 10:50:23 -04:00
JesseMarkowitz 024bfde030 tools: test whether a browser image URL works from the exporter
The endpoint hunt is done guessing. Results:

  /files/{libfile}/download    200 {"error_code":"file_not_found",
                                    "error_type":"GetDownloadLinkError"}
                               — twice, on two different Library ids, so the
                                 Library id is simply not valid there
  /library/*                   404 across every shape
  POST on the download route   405
  /content?asset_pointer=      422 requiring query params id, ts, p

That last one is the find: /backend-api/content is a signed route, and the
web UI has the values we cannot compute. So the question shifts from "which
endpoint" to "is the UI's URL reusable outside the browser".

try_pasted_url.py takes a URL copied from DevTools and fetches it three
ways — bare, through the exporter's authenticated session, and with the
Authorization header removed in case bearer and signature conflict. The
pattern says whether the exporter can fetch these at all, and therefore
whether it is worth hunting for what mints the signature.

Before any of that: check whether the images still render in ChatGPT's own
UI. If they show broken there, file_not_found is literally true, these 11
are lost like the other 7, and no route exists to find.
2026-08-17 10:27:57 -04:00
JesseMarkowitz ffd01ebcf3 tools: print the error body the Library endpoint returns
The pairing came back perfectly clean, and it is the whole diagnosis:

    has library_file_id  → refused (11/11)
    no  library_file_id  → deleted (7/7)

And /files/{libfile_id}/download answered 200 — not 403, not 404 — with
{status, error_code, error_type, error_message}. That endpoint accepts the
Library id; it is returning an application-level error inside an HTTP 200.
The probe printed only the keys and dropped the message, which is the one
thing that says what the call is missing.

- show() now prints the full body for every response, and detects a
  non-JSON 200 as possible raw bytes.
- Added routes worth ruling in or out: the same call as POST, with a
  conversation_id, /files/{lib}/content, /content?asset_pointer=, and four
  Library listing endpoints — a listing usually reveals both the id form
  the UI uses and the route that actually serves bytes.
- Probes a second Library id too, so one odd file cannot mislead us.
2026-08-17 10:14:18 -04:00
JesseMarkowitz 3f82b35a55 tools: try fetching refused images by their Library ID
Dumping the raw message metadata found the identity the asset pointer never
carried:

    "id":              "file_000000003454722f9481506b96aed510"   ← refused
    "library_file_id": "libfile_4eb82f478fe081919127e2eba9886e86"
    "source":          "local"

The exporter only knows the sediment id from the asset_pointer and asks
/files/{sediment}/download, which 403s for these. The Library is a separate
store with its own ids, so we have been asking for the conversation-scoped
copy of a file that now lives in the Library. That also fits the 2026-07-14
creation-time cluster: a Library migration would mint exactly these new
records.

Neither gizmo_id, conversation_id, a /gizmos path nor a project Referer
helped (403/404), so scope was never the missing piece — identity was.

library_download_probe.py pairs each refused sediment id with its
library_file_id from message.metadata.attachments and tries the endpoints
that could serve it.
2026-08-17 09:33:05 -04:00
JesseMarkowitz 03646009b9 tools: probe how to download a gizmo-scoped file
metadata_diff found the discriminator. Every refused image carries
use_case="gizmo" — ChatGPT's name for Projects and custom GPTs — while
everything that downloads is image_gen or multimodal. All are state=ready,
so nothing is damaged: the plain /files/{id}/download endpoint just will
not serve a project-scoped file.

Second clue: all 11 were created 2026-07-14 within minutes of each other,
yet appear in conversations dated 2025-11 through 2026-08, several
predating their own creation_time. Something on July 14 re-created them as
gizmo-scoped copies and repointed the conversations — which is also why
the losses start in July. Not a policy change, an event.

gizmo_download_probe.py dumps the raw conversation part carrying one of
these assets (it may simply name the scope the download wants) and then
tries the plausible calls — gizmo_id/conversation_id query params, a
/gizmos/{id}/files path, a project Referer. A 200 is the fix.
2026-08-17 09:22:48 -04:00
JesseMarkowitz a3ac279e39 tools: find what separates a refused image from a served one
branch_check killed the abandoned-branch hypothesis — every lost image is
on the live branch, and the 3 images that do sit on abandoned branches are
alive. But it turned up something the earlier probing missed by sampling
four IDs from one family and generalising: of the 19 failures, only 7 are
actually gone. The other 12 answer /files/{id} with 200 and full metadata
and refuse only /download. They exist, and may be recoverable.

Three states, then: gone (404), refused (200 meta + 403 download), working
(200 both). Since metadata comes back for the refused ones, the
discriminator can be read straight off — fetch it for every file in each
state and compare fields, flagging any field whose values never overlap
between states.

Also re-checks /download now: if a file refused during the export serves
today, those 403s were transient and a retry pass recovers them, which is
a completely different fix from anything permanent.
2026-08-17 09:18:32 -04:00
JesseMarkowitz 021a76c628 tools: test whether lost images sit on abandoned branches
Expiry is now ruled out: 36/36 sampled images from 2025-09 through
2026-08 are still live server-side, so nothing dies of age and export
cadence is not the variable. The loss is per-asset — 2026-07-09 kept 38
images and lost 12 in one conversation.

Next hypothesis: those images belong to messages that were edited or
regenerated. ChatGPT stores a conversation as a tree and editing forks it,
stranding the superseded messages on an abandoned branch. The exporter
walks every node (chatgpt.py:963), so it exports those branches too — and
an attachment unreachable from live history is a natural GC target.

branch_check.py fetches the raw conversation, derives the live branch by
walking parent links up from current_node, and cross-tabulates every image
against (on the live branch?) x (still downloadable?). If the lost images
are all off-branch and nothing on the live branch is missing, this is not
data loss at all — it is attachments to messages that were replaced.
2026-08-17 08:58:20 -04:00
JesseMarkowitz b9e8896b33 tools: distinguish "captured in time" from "still alive"
analyze_media_age.py claimed age was ruled out because an image from
2025-11 was saved while one from 2026-08 was not. That conclusion does not
follow. "Saved" means some earlier run downloaded it, not that ChatGPT
still holds it — an image captured in June looks saved forever after, even
if it died in July. A month with no losses shows the exports were timely,
not that the assets survived.

- probe_survival_by_month.py: sample images already on disk, grouped by
  their conversation's month, and ask /files/{id} whether each still
  exists today. Old months still 200 → no expiry, and cadence did not save
  them. Old months now 404 → uploads do expire and cadence is the whole
  ballgame.
- analyze_media_age.py: stop asserting the unsupported verdict; say what
  the number does and does not show, and point at the probe.

One signal there is immune to the confound and survives: 2026-07-09 kept
38 images and lost 12. Same conversation, same day, opposite outcomes —
no retention policy does that, so at least part of this is per-asset.
2026-08-17 08:53:54 -04:00
JesseMarkowitz 04191eed8c tools: analyze whether media loss is age-based
"Do I have to export within N days?" is answerable from the exports
already on disk — the renderer records every image's outcome inline
(![source](media/…) when saved, a placeholder when not) and the
conversation date is in the filename. Group by month and source and the
hypotheses separate: a clean old/new cutoff means expiry, user_upload
dying at an age model_generated survives means the source matters, and
losses scattered through months that otherwise downloaded fine means
neither.

Offline, no token, no API calls.
2026-08-17 08:45:21 -04:00
JesseMarkowitz f40b25001a tools: add one-off probe for the media 403s
The improved error reporting landed, and the answer it produced is
{"detail":"Forbidden"} — generic, no reason. The cause has to be narrowed
by experiment instead, so collect the experiments in one script:

- classify the failed assets offline from the exported placeholders
  (user_upload vs model_generated) — no API call needed
- probe a known-good asset in the same session as a control, to rule the
  session in or out
- retry the 403 with ChatGPT-Account-Id, which the exporter never sends
  and which workspace-scoped resources can require
- compare /files/{id} against /files/{id}/download

Temporary: delete once the cause is known, or fold into `doctor` if the
check earns a permanent home.
2026-08-17 08:08:01 -04:00