Python 超簡單爬取微博熱搜榜資料

pythondict發表於2020-05-13

原文網址 : https://learnku.com/articles/44486

微博的熱搜榜對於研究大眾的流量有非常大的價值。今天的教程就來說說如何爬取微博的熱搜榜。熱搜榜的連結是：

s.weibo.com/top/summary/

用瀏覽器瀏覽，發現在不登入的情況下也可以正常檢視，那就簡單多了。使用開發者工具(F12)檢視頁面邏輯，並拿到每條熱搜的CSS位置，方法如下：

Python 熱搜榜爬蟲

按照這個方法，拿到這個td標籤的selector是：
pl_top_realtimehot > table > tbody > tr:nth-child(3) > td.td-02
其中nth-child(3)指的是第三個tr標籤，因為這條熱搜是在第三名的位置上，但是我們要爬的是所有熱搜，因此:nth-child(3)可以去掉。還要注意的是 pl_top_realtimehot 是該標籤的id，id前需要加#號，最後變成：
#pl_top_realtimehot > table > tbody > tr > td.td-02

你可以自定義你想要爬的資訊，這裡我需要的資訊是：熱搜的連結及標題、熱搜的熱度。它們分別對應的CSS選擇器是：

連結及標題：#pl_top_realtimehot > table > tbody > tr > td.td-02 > a
熱度：#pl_top_realtimehot > table > tbody > tr > td.td-02 > span

值得注意的是連結及標題是在同一個地方，連結在a標籤的href屬性裡，標題在a的文字中，用beautifulsoup有辦法可以都拿到，請看後文程式碼。

現在這些資訊的位置我們都知道了，接下來可以開始編寫程式。預設你已經安裝好了python，並能使用cmd的pip，如果沒有的話請見這篇教程：python安裝。需要用到的python的包有：

BeautifulSoup4:
cmd/Terminal 安裝指令：

pip install beautifulsoup4

lxml解析器：
cmd/Terminal 安裝指令：

pip install lxml

lxml是python中的一個包，這個包中包含了將html文字轉成xml物件的工具，可以讓我們定位標籤的位置。而能用來識別xml物件中這些標籤的位置的包就是 Beautifulsoup4.

編寫程式碼：

# https://s.weibo.com/top/summary/
import requests
from bs4 import BeautifulSoup

if __name__ == "__main__":
    news = []
    # 新建陣列存放熱搜榜
    hot_url = 'https://s.weibo.com/top/summary/'
    # 熱搜榜連結
    r = requests.get(hot_url)
    # 向連結傳送get請求獲得頁面
    soup = BeautifulSoup(r.text, 'lxml')
    # 解析頁面

    urls_titles = soup.select('#pl_top_realtimehot > table > tbody > tr > td.td-02 > a')
    hotness = soup.select('#pl_top_realtimehot > table > tbody > tr > td.td-02 > span')

    for i in range(len(urls_titles)-1):
        hot_news = {}
        # 將資訊儲存到字典中
        hot_news['title'] = urls_titles[i+1].get_text()
        # get_text()獲得a標籤的文字
        hot_news['url'] = "https://s.weibo.com"+urls_titles[i]['href']
        # ['href']獲得a標籤的連結，並補全字首
        hot_news['hotness'] = hotness[i].get_text()
        # 獲得熱度文字
        news.append(hot_news) 
        # 字典追加到陣列中 

    print(news)

程式碼說明請看註釋，不過這樣做，我們僅僅是將結果儲存到陣列中，如下所示，其實不易觀看，我們下面將其儲存為csv檔案。

Python 熱搜榜爬蟲

    import datetime
    today = datetime.date.today()
    f = open('./熱搜榜-%s.csv'%(today), 'w', encoding='utf-8')
    for i in news:
        f.write(i['title'] + ',' + i['url'] + ','+ i['hotness'] + 'n')

效果如下，怎麼樣，是不是好看很多：

Python 微博熱搜榜爬蟲

完整程式碼如下：

# https://s.weibo.com/top/summary/
import requests
from bs4 import BeautifulSoup

if __name__ == "__main__":
    news = []
    # 新建陣列存放熱搜榜
    hot_url = 'https://s.weibo.com/top/summary/'
    # 熱搜榜連結
    r = requests.get(hot_url)
    # 向連結傳送get請求獲得頁面
    soup = BeautifulSoup(r.text, 'lxml')
    # 解析頁面

    urls_titles = soup.select('#pl_top_realtimehot > table > tbody > tr > td.td-02 > a')
    hotness = soup.select('#pl_top_realtimehot > table > tbody > tr > td.td-02 > span')

    for i in range(len(urls_titles)-1):
        hot_news = {}
        # 將資訊儲存到字典中
        hot_news['title'] = urls_titles[i+1].get_text()
        # get_text()獲得a標籤的文字
        hot_news['url'] = "https://s.weibo.com"+urls_titles[i]['href']
        # ['href']獲得a標籤的連結，並補全字首
        hot_news['hotness'] = hotness[i].get_text()
        # 獲得熱度文字
        news.append(hot_news)
        # 字典追加到陣列中

    print(news)

    import datetime
    today = datetime.date.today()
    f = open('./熱搜榜-%s.csv'%(today), 'w', encoding='utf-8')
    for i in news:
        f.write(i['title'] + ',' + i['url'] + ','+ i['hotness'] + 'n')

Python實用寶典 (pythondict.com)
不只是一個寶典
歡迎關注公眾號：Python實用寶典
原文來自Python實用寶典：Python 微博熱搜

Python 教程

本作品採用《CC 協議》，轉載必須註明作者和本文連結

Python實用寶典, pythondict.com

Python 超簡單爬取新浪微博資料 (高階版)
2020-05-16
Python
Python實現微博爬蟲，爬取新浪微博
2020-12-14
Python爬蟲
一個批次爬取微博資料的神器
2024-08-30
微博-指定話題當日資料爬取
2024-06-12
python爬蟲獲取百度熱搜
2024-06-15
Python爬蟲
上萬條資料撕開微博熱搜的真相！
2022-12-08
python實現微博個人主頁的資訊爬取
2021-01-03
Python
Python 教你動態展示微博熱搜排名變化
2020-03-11
Python
一個比微博熱搜更適合吃瓜的平臺——即時熱榜
2021-11-02
JB的Python之旅-爬蟲篇-新浪微博內容爬取
2018-06-30
Python爬蟲
爬取微博圖片資料存到Mysql中遇到的各種坑mysql儲存圖片爬取微博圖片
2019-02-16
MySql
2021上半年微博熱搜榜趨勢報告（附下載）
2021-08-30
python爬取qq音樂歌手排行熱度資料
2022-05-23
Python
爬蟲實戰（一）：爬取微博使用者資訊
2018-07-15
爬蟲
Python網路爬蟲2 - 爬取新浪微博使用者圖片
2018-04-10
Python爬蟲
用PYTHON爬蟲簡單爬取網路小說
2021-09-11
Python爬蟲
Golang+chromedp+goquery 簡單爬取動態資料
2021-03-05
GolangChrome
微博爬取長津湖博文及評論
2021-10-08
python 爬取 blessing skin 的簡單實現
2020-03-04
Python
Python爬蟲精簡步驟1 獲取資料
2020-02-17
Python爬蟲
selenium + xpath爬取csdn關於python的博文博主資訊
2020-12-19
Python
python itchat 爬取微信好友資訊
2018-06-02
Python
房產資料爬取、智慧財產權資料爬取、企業工商資料爬取、抖音直播間資料python爬蟲爬取
2024-07-11
Python爬蟲
Python：爬取疫情每日資料
2020-02-17
Python
python 爬蟲 mc 皮膚站 little skin 的簡單爬取
2019-08-02
Python爬蟲
爬蟲如何爬取貓眼電影TOP榜資料
2019-06-17
爬蟲
Android Kotlin retrofit2 網路請求學習獲取微博熱搜列表
2020-11-16
AndroidKotlin
【爬蟲+資料分析+資料視覺化】python資料分析全流程《2021胡潤百富榜》榜單資料！
2022-12-29
爬蟲視覺化Python
Python 爬取 baidu 股票市值資料
2019-02-16
PythonAI
Python爬取噹噹網APP資料
2020-10-21
PythonAPP
Python爬取CSDN部落格資料
2019-01-03
Python
使用 Python 爬取網站資料
2024-07-27
Python網站
Scrapy框架的使用之Scrapy爬取新浪微博
2018-05-23
框架
搜狗搜尋微信Python爬蟲案例
2022-04-04
Python爬蟲
python-爬蟲-css提取-寫入csv-爬取貓眼電影榜單
2023-04-05
Python爬蟲CSS
Python 超簡單玩轉微信自動回覆
2020-05-14
Python
python爬取股票資料並存到資料庫
2021-03-29
Python資料庫
利用 Python 爬取“工商祕密”微博，看看大家都在關注些什麼？
2020-12-21
Python

Python 超簡單爬取微博熱搜榜資料

相關文章