python 爬蟲之requests爬取頁面圖片的url，並將圖片下載到本地

夢寐千尋發表於2019-06-12

原文網址 : https://www.cnblogs.com/hardykay/p/11009670.html

Python爬蟲

大家好我叫hardy

需求：爬取某個頁面，並把該頁面的圖片下載到本地

思考：

　　img標籤一個有多少種型別的src值？四種：1、以http開頭的網路連結。2、以“//”開頭網路地址。3、以“/”開頭絕對路徑。4、以“./”開頭相對路徑。當然還有其他型別，不過這個不做考慮，能力有限呀。

　　使用什麼工具？我用requests、xpth

　　都有那些步驟：1、爬取網頁

　　　　　　　　　　2、分析html並獲取img中的src的值

　　　　　　　　　　3、獲取圖片

　　　　　　　　　　4、儲存

具體實現

import requests
from lxml import etree
import time
import os
import re

requests = requests.session()

website_url = ''
website_name = ''

'''
爬取的頁面
'''
def html_url(url):
    try:
        head = set_headers()
        text = requests.get(url,headers=head)
        # print(text)
        html = etree.HTML(text.text)
        img = html.xpath('//img/@src')
        # 儲存圖片
        for src in img:
            src = auto_completion(src)
            file_path = save_image(src)
            if file_path == False:
                print('請求的圖片路徑出錯，url地址為：%s'%src)
            else :
                print('儲存圖片的地址為：%s'%file_path)
    except requests.exceptions.ConnectionError as e:
        print('網路地址無法訪問，請檢查')
        print(e)
    except requests.exceptions.RequestException as e:
        print('訪問異常：')
        print(e)


'''
儲存圖片
'''
def save_image(image_url):
    size = 0
    number = 0
    while size == 0:
        try:
            img_file = requests.get(image_url)
        except requests.exceptions.RequestException as e:
            raise e

        # 不是圖片跳過
        if check_image(img_file.headers['Content-Type']):
            return False
        file_path = image_path(img_file.headers)
        # 儲存
        with open(file_path, 'wb') as f:
            f.write(img_file.content)
        # 判斷是否正確儲存圖片
        size = os.path.getsize(file_path)
        if size == 0:
            os.remove(file_path)
        # 如果該圖片獲取超過十次則跳過
        number += 1
        if number >= 10:
            break
    return (file_path if (size > 0) else False)

'''
自動完成url的補充
'''
def auto_completion(url):
    global website_name,website_url
    #如果是http://或者https://開頭直接返回
    if re.match('http://|https://',url):
        return url
    elif re.match('//',url):
        if 'https://' in website_name:
            return 'https:'+url
        elif 'http://' in website_name:
            return 'http:' + url
    elif re.match('/',url):
        return website_name+url
    elif re.match('./'):
        return website_url+url[1::]

'''
圖片儲存的路徑
'''
def image_path(header):
    # 資料夾
    file_dir = './save_image/'
    if not os.path.exists(file_dir):
        os.makedirs(file_dir)
    # 檔名
    file_name = str(time.time())
    # 檔案字尾
    suffix = img_type(header)

    return file_dir + file_name + suffix


'''
獲取圖片字尾名
'''
def img_type(header):
    # 獲取檔案屬性
    image_attr = header['Content-Type']
    # 獲取字尾
    suffix = image_attr.split('/')[1]
    if suffix == 'jpeg':
        suffix = 'jpg'

    return '.' + suffix


# 檢查是否為圖片型別
def check_image(content_type):
    if 'image' in content_type:
        return False
    else:
        return True
#設定頭部
def set_headers():
    global website_name, website_url
    head = {
        'Host':website_name.split('//')[1],
        'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/73.0.3683.103 Safari/537.36',
    }
    return head



if __name__ == '__main__':

    #當前的url，不包含檔名的比如index.html，用來下載當前頁的頁面圖片(./)
    website_url = 'https://dig.chouti.com/all/hot/recent'
    #域名，用來下載"/"開頭的圖片地址
    #感興趣的朋友請幫我完善一下這個自動完成圖片url的補充
    website_name = 'https://dig.chouti.com'
    url = 'https://dig.chouti.com/all/hot/recent/1'
    html_url(url)

爬蟲 Scrapy框架爬取圖蟲圖片並下載
2018-08-27
爬蟲框架
網路爬蟲---從千圖網爬取圖片到本地
2019-09-03
爬蟲
node：爬蟲爬取網頁圖片
2019-02-16
爬蟲網頁
python入門012～使用requests爬取網路圖片並儲存到本地
2021-09-09
Python
Python爬蟲—爬取某網站圖片
2020-11-19
Python爬蟲網站
python爬蟲---網頁爬蟲，圖片爬蟲，文章爬蟲，Python爬蟲爬取新聞網站新聞
2019-01-04
Python爬蟲網頁網站
Java爬蟲批量爬取圖片
2021-09-24
Java爬蟲
Python爬蟲入門【5】：27270圖片爬取
2019-07-30
Python爬蟲
python爬取鬥圖啦表情包並下載到本地
2018-12-25
Python
python 爬蟲下載百度美女圖片
2024-04-18
Python爬蟲
爬蟲---xpath解析（爬取美女圖片）
2020-12-23
爬蟲
Python爬蟲實戰詳解：爬取圖片之家
2020-11-04
Python爬蟲
Python爬蟲——批次爬取douyin影片，下載到本地
2024-12-06
Python爬蟲
【python--爬蟲】千圖網高清背景圖片爬蟲
2019-05-21
Python爬蟲
使用Python爬蟲實現自動下載圖片
2021-09-11
Python爬蟲
新手爬蟲教程：Python爬取知乎文章中的圖片
2019-01-17
爬蟲Python
Python爬蟲新手教程：知乎文章圖片爬取器
2019-07-20
Python爬蟲
Python爬蟲遞迴呼叫爬取動漫美女圖片
2020-10-19
Python爬蟲遞迴
使用Scrapy爬取圖片入庫,並儲存在本地
2019-06-27
go語言實現簡單爬蟲獲取頁面圖片
2022-11-14
Go爬蟲
自學python網路爬蟲，從小白快速成長，分別實現靜態網頁爬取，下載meiztu中圖片；動態網頁爬取，下載burberry官網所有當季新品圖片。
2020-02-06
Python爬蟲網頁
[Python]爬蟲獲取知乎某個問題下所有圖片並去除水印
2021-09-20
Python爬蟲
Python應用開發——爬取網頁圖片
2022-09-21
Python網頁
前端js儲存頁面為圖片下載到本地
2020-10-27
前端JS
Java爬蟲爬取bing必應每日一圖背景圖下載到本地(HttpClient+Jsoup+Jackson)
2020-10-20
Java爬蟲HTTPclientJS
Python網路爬蟲2 - 爬取新浪微博使用者圖片
2018-04-10
Python爬蟲
Python爬蟲入門【4】：美空網未登入圖片爬取
2019-07-30
Python爬蟲
爬蟲Selenium+PhantomJS爬取動態網站圖片資訊（Python）
2018-03-24
爬蟲JS網站Python
python爬蟲系列(4.5-使用urllib模組方式下載圖片)
2018-11-09
Python爬蟲
Python資料爬蟲學習筆記（11）爬取千圖網圖片資料
2018-09-18
Python爬蟲筆記
AotucCrawler 快速爬取圖片
2021-11-25
實用爬蟲-03-爬取視訊教程課程名+連結+下載圖片
2018-10-29
爬蟲
簡單的爬蟲：爬取網站內容正文與圖片
2021-09-09
爬蟲網站
Python《必應bing桌面圖片爬取》
2020-12-26
Python
ReactPHP 爬蟲實戰：下載整個網站的圖片
2019-01-20
ReactPHP爬蟲網站
如何用Python爬蟲實現百度圖片自動下載？
2019-03-01
Python爬蟲
Python 爬蟲零基礎教程(1)：爬單個圖片
2024-03-13
Python爬蟲
爬取愛套圖網上的圖片
2018-03-28

python 爬蟲之requests爬取頁面圖片的url，並將圖片下載到本地

相關文章