get file size before downloading using HTTP header not matching with one retrieved from urlopen

About

Asked 5/7, 2014 at 9:27 Answered 5/7, 2014 at 10:12

why is the content-lenght different in case of using requests and urlopen(url).info()

>>> url = 'http://pymotw.com/2/urllib/index.html'

>>> requests.head(url).headers.get('content-length', None)
'8176'
>>> urllib.urlopen(url).info()['content-length']
'38227'
>>> len(requests.get(url).content)
38274

I was going to make a check for size of file in bytes to split the buffer to multiple threads based on Range in urllib2 but if I do not have the actual size of file in bytes it won't work..

only len(requests.get(url).content) gives 38274 which is closest but still not correct and moreover it is downloading the content which i didn't wanted.

Reseat answered 5/7, 2014 at 9:27 Comment(4)

It might be that the server does not correctly support HEAD. Or it might be that the server returns different things based on other headers sent (or not) by the respective methods (user-agent, cookies...). Try using curl -v url and curl -I, or any other method that sends the exact same request save for the HEAD instead of post, and check the results. – Jenson 5/7, 2014 at 9:31

Maybe first size is the compressed size? – Keelboat 5/7, 2014 at 9:35

@Keelboat : then how to get uncompressed size ? – Reseat 5/7, 2014 at 9:36

You can have a (very rough) estimate by looking at the compression method. For example, gzip has a compression ratio of 3:1 - 5:1 (source: superuser.com/questions/139253/…) – Keelboat 5/7, 2014 at 9:49

By default, requests will send 'Accept-Encoding': 'gzip' as part of the request headers, and the server will respond with the compressed content:

>>> r = requests.head('http://pymotw.com/2/urllib/index.html')
r>>> r.headers['content-encoding'], r.headers['content-length']
('gzip', '8201')

But, if you manually set the request headers, then you'll get the uncompressed content:

>>> r = requests.head('http://pymotw.com/2/urllib/index.html',headers={'Accept-Encoding': 'identity'})
>>> r.headers['content-length']
'38227'

Juarez answered 5/7, 2014 at 10:12 Comment(0)

Hot tags

Godot Unity Godot Help Programming Godot 4.X GUI GDScript 3D 2D Physics CSharp Godot 3.X VR XR Projects C++

Recommended topics

Hot tags