Python爬蟲入門教程 11-100 行行網電子書多執行緒爬取

夢想橡皮擦發表於2018-12-25

原文網址 : https://flycode.co/archives/232648

行行網電子書多執行緒爬取-寫在前面

最近想找幾本電子書看看，就翻啊翻，然後呢，找到了一個叫做 周讀的網站，網站特別好，簡單清爽，書籍很多，而且開啟都是百度網盤可以直接下載，更新速度也還可以，於是乎，我給爬了。本篇文章學習即可，這麼好的分享網站，儘量不要去爬，影響人家訪問速度就不好了 http://www.ireadweek.com/ ,想要資料的，可以在我部落格下面評論，我發給你，QQ，郵箱，啥的都可以。

在這裡插入圖片描述

這個網站頁面邏輯特別簡單，我翻了翻書籍詳情頁面，就是下面這個樣子的，我們只需要迴圈生成這些頁面的連結，然後去爬就可以了，為了速度，我採用的多執行緒，你試試就可以了，想要爬取之後的資料，就在本篇部落格下面評論，不要搞壞別人伺服器。

http://www.ireadweek.com/index.php/bookInfo/11393.html
http://www.ireadweek.com/index.php/bookInfo/11.html
....

行行網電子書多執行緒爬取-擼程式碼

程式碼非常簡單，有我們們前面的教程做鋪墊，很少的程式碼就可以實現完整的功能了，最後把採集到的內容寫到 csv 檔案裡面，(csv 是啥，你百度一下就知道了) 這段程式碼是IO密集操作 我們採用aiohttp模組編寫。

第1步

拼接URL，開啟執行緒。

import requests

# 匯入協程模組
import asyncio
import aiohttp


headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/68.0.3440.106 Safari/537.36",
           "Host": "www.ireadweek.com",
           "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8"}

async def get_content(url):
    print("正在操作:{}".format(url))
    # 建立一個session 去獲取資料 
    async with aiohttp.ClientSession() as session:
        async with session.get(url,headers=headers,timeout=3) as res:
            if res.status == 200:
                source = await res.text()  # 等待獲取文字
                print(source)


if __name__ == `__main__`:
    url_format = "http://www.ireadweek.com/index.php/bookInfo/{}.html"
    full_urllist = [url_format.format(i) for i in range(1,11394)]  # 11394
    loop = asyncio.get_event_loop()
    tasks = [get_content(url) for url in full_urllist]
    results = loop.run_until_complete(asyncio.wait(tasks))

上面的程式碼可以同步開啟N多個執行緒，但是這樣子很容易造成別人的伺服器癱瘓，所以，我們必須要限制一下併發次數，下面的程式碼，你自己嘗試放到指定的位置吧。

sema = asyncio.Semaphore(5)
# 為避免爬蟲一次性請求次數太多，控制一下
async def x_get_source(url):
    with(await sema):
        await get_content(url)

第2步

處理抓取到的網頁原始碼，提取我們想要的元素，我新增了一個方法，採用lxml進行資料提取。

def async_content(tree):
    title = tree.xpath("//div[@class=`hanghang-za-title`]")[0].text
    # 如果頁面沒有資訊，直接返回即可
    if title == ``:
        return
    else:
        try:
            description = tree.xpath("//div[@class=`hanghang-shu-content-font`]")
            author = description[0].xpath("p[1]/text()")[0].replace("作者：","") if description[0].xpath("p[1]/text()")[0] is not None else None
            cate = description[0].xpath("p[2]/text()")[0].replace("分類：","") if description[0].xpath("p[2]/text()")[0] is not None else None
            douban = description[0].xpath("p[3]/text()")[0].replace("豆瓣評分：","") if description[0].xpath("p[3]/text()")[0] is not None else None
            # 這部分內容不明確，不做記錄
            #des = description[0].xpath("p[5]/text()")[0] if description[0].xpath("p[5]/text()")[0] is not None else None
            download = tree.xpath("//a[@class=`downloads`]")
        except Exception as e:
            print(title)
            return

    ls = [
        title,author,cate,douban,download[0].get(`href`)
    ]
    return ls

第3步

資料格式化之後，儲存到csv檔案，收工！

 print(data)
 with open(`hang.csv`, `a+`, encoding=`utf-8`) as fw:
     writer = csv.writer(fw)
     writer.writerow(data)
 print("插入成功！")

行行網電子書多執行緒爬取-執行程式碼，檢視結果

在這裡插入圖片描述

因為這個可能涉及到獲取別人伺服器重要資料了，程式碼不上傳github了，有需要的留言吧，我單獨傳送給你

Python爬蟲入門【10】：電子書多執行緒爬取
2019-07-31
Python爬蟲執行緒
Python爬蟲入門【9】：圖蟲網多執行緒爬取
2019-07-31
Python爬蟲執行緒
Python爬蟲入門教程 13-100 鬥圖啦表情包多執行緒爬取
2018-12-27
Python爬蟲執行緒
python爬蟲入門八：多程式/多執行緒
2019-01-07
Python爬蟲執行緒
python爬蟲學習01--電子書爬取
2020-07-13
Python爬蟲
python多執行緒爬蟲與單執行緒爬蟲效率效率對比
2021-03-19
Python執行緒爬蟲
Python《多執行緒併發爬蟲》
2020-12-12
Python執行緒爬蟲
Python爬蟲入門教程 2-100 妹子圖網站爬取
2018-12-13
Python爬蟲網站
Python爬蟲入門教程 50-100 Python3爬蟲爬取VIP視訊-Python爬蟲6操作
2019-02-14
Python爬蟲
Python爬蟲入門【3】：美空網資料爬取
2019-07-30
Python爬蟲
Python爬蟲入門教程 4-100 美空網未登入圖片爬取
2018-12-17
Python爬蟲
python多執行緒非同步爬蟲-Python非同步爬蟲試驗[Celery,gevent,requests]
2020-11-11
Python執行緒非同步爬蟲
Python爬蟲入門【5】：27270圖片爬取
2019-07-30
Python爬蟲
python爬蟲之多執行緒、多程式+程式碼示例
2020-08-26
Python爬蟲執行緒
簡易多執行緒爬蟲框架
2018-06-02
執行緒爬蟲框架
多執行緒爬蟲實現（上）
2018-05-26
執行緒爬蟲
Python爬蟲入門【4】：美空網未登入圖片爬取
2019-07-30
Python爬蟲
python爬蟲---網頁爬蟲，圖片爬蟲，文章爬蟲，Python爬蟲爬取新聞網站新聞
2019-01-04
Python爬蟲網頁網站
Python爬蟲教程-13-爬蟲使用cookie爬取登入後的頁面(人人網)（下）
2018-09-06
Python爬蟲Cookie
Python爬蟲教程-12-爬蟲使用cookie爬取登入後的頁面(人人網)（上）
2018-09-06
Python爬蟲Cookie
Python爬蟲入門
2020-11-30
Python爬蟲
Python爬蟲入門教程導航帖
2019-01-08
Python爬蟲
擼個爬蟲，爬取電影種子
2019-05-11
爬蟲
Python爬蟲入門【11】：半次元COS圖爬取
2019-07-31
Python爬蟲
資料提取方法-多程式多執行緒爬蟲
2020-11-16
執行緒爬蟲
Python使用多程式提高網路爬蟲的爬取速度
2019-02-01
Python爬蟲
如何使用python多執行緒有效爬取大量資料？
2021-09-11
Python執行緒
Python爬蟲教程+書籍分享
2018-11-29
Python爬蟲
Python爬蟲教程-17-ajax爬取例項（豆瓣電影）
2018-09-06
Python爬蟲
【爬蟲】python爬蟲從入門到放棄
2018-12-20
爬蟲Python
多執行緒爬取B站視訊
2020-10-13
執行緒
python-爬蟲入門
2024-09-22
Python爬蟲
Python 從入門到爬蟲極簡教程
2019-02-16
Python爬蟲
什麼是Python爬蟲？python爬蟲入門難嗎？
2021-12-27
Python爬蟲
Python網路爬蟲4 - scrapy入門
2018-05-29
Python爬蟲
python網路爬蟲_Python爬蟲：30個小時搞定Python網路爬蟲視訊教程
2020-10-21
Python爬蟲
Python爬蟲15--爬蟲遇上多執行緒，速度更上一層樓，爬取1000張圖片連一分鐘也不要！
2021-01-03
Python爬蟲執行緒
Python爬蟲入門教程 8-100 蜂鳥網圖片爬取之三
2018-12-20
Python爬蟲

Python爬蟲入門教程 11-100 行行網電子書多執行緒爬取

行行網電子書多執行緒爬取-寫在前面

行行網電子書多執行緒爬取-擼程式碼

第1步

第2步

第3步

行行網電子書多執行緒爬取-執行程式碼，檢視結果

相關文章