代码之家  ›  专栏  ›  技术社区  ›  BigBowl

在网页上抓取一个jpg文件,然后使用python保存

  •  1
  • BigBowl  · 技术社区  · 11 年前

    好吧,我正在尝试从Gucci网站上抓取jpg图像。以这个为例。

    http://www.gucci.com/images/ecommerce/styles_new/201501/web_full/277520_F4CYG_4080_001_web_full_new_theme.jpg

    我尝试了urllib。urlreserve,这不起作用,因为Gucci阻止了该功能。所以我想使用请求来抓取图像的源代码,然后将其写入.jpg文件。

    image = requests.get("http://www.gucci.com/images/ecommerce/styles_new/201501/web_full/277520_F4CYG_4080_001_web_full_new_theme.jpg").text.encode('utf-8')
    

    我对它进行了编码,因为如果我不这样做,它会一直告诉我gbk无法对字符串进行编码。

    然后:

    with open('1.jpg', 'wb') as f:
        f.write(image)
    

    看起来不错吧?但结果是无法打开jpg文件。没有图像!Windows告诉我jpg文件已损坏。

    问题可能是什么?

    1. 我在想,也许当我抓取图像时,我丢失了一些信息,或者一些字符被错误地抓取了。但我怎样才能找出哪一个呢?

    2. 我在想,也许有些信息通过编码丢失了。但如果我不编码,我甚至无法打印,更不用说将其写入文件。

    什么会出错?

    2 回复  |  直到 11 年前
        1
  •  1
  •   Sam    11 年前

    我不确定你使用 encode 。您不是在处理文本,而是在处理图像。您需要将响应作为二进制数据而不是文本访问,并使用图像处理函数而不是文本函数。试试看:

    from PIL import Image
    from io import BytesIO
    import requests
    
    response = requests.get("http://www.gucci.com/images/ecommerce/styles_new/201501/web_full/277520_F4CYG_4080_001_web_full_new_theme.jpg")
    bytes = BytesIO(response.content)
    image = Image.open(bytes)
    image.save("1.jpg")
    

    注意使用 response.content 而不是 response.text 。您需要安装PIL或枕头才能使用 Image 单元 BytesIO 包含在Python 3中。

    或者,您可以直接将数据保存到磁盘,而不必查看其中的内容:

    import requests
    response = requests.get("http://www.gucci.com/images/ecommerce/styles_new/201501/web_full/277520_F4CYG_4080_001_web_full_new_theme.jpg")
    with open('1.jpg','wb') as f:
        f.write(response.content)
    
        2
  •  0
  •   PM 2Ring    11 年前

    JPEG文件不是文本,而是二进制数据。所以你需要使用 request.content 属性来访问它。

    下面的代码还包括 get_headers() 功能,这在您浏览网站时很方便。

    import requests
    
    def get_headers(url):
        resp = requests.head(url)
        print("Status: %d" % resp.status_code)
        resp.raise_for_status()
        for t in resp.headers.items():
            print('%-16s : %s' % t)
    
    def download(url, fname):
        ''' Download url to fname '''
        print("Downloading '%s' to '%s'" % (url, fname))
        resp = requests.get(url)
        resp.raise_for_status()
        with open(fname, 'wb') as f:
            f.write(resp.content)
    
    def main():
        site = 'http://www.gucci.com/images/ecommerce/styles_new/201501/web_full/'
        basename = '277520_F4CYG_4080_001_web_full_new_theme.jpg'
        url = site + basename
        fname = 'qtest.jpg'
    
        try:
            #get_headers(url)
            download(url, fname)
        except requests.exceptions.HTTPError as e:
            print("%s '%s'" % (e, url))
    
    if __name__ == '__main__':
        main()
    

    我们称之为 .raise_for_status() 方法,以便 获取标头() 和 download() 如果出现问题,则引发异常;我们在中捕获异常 main() 并打印相关信息。

    推荐文章