Python爬蟲包BeautifulSoup異常處理（二）

發布時間：2020-09-05 18:31:43 來源：腳本之家閱讀：192 作者：SuPhoebe 欄目：開發技術

面對網絡不穩定，頁面更新等問題，很可能出現程序異常的問題，所以我們要對程序進行一些異常處理。大家可能覺得處理異常是一個比較麻煩的活，但在面對復雜網頁和任務的時候，無疑成為一個很好的代碼習慣。

網頁‘404'、‘500'等問題

try:
    html = urlopen('http://www.pmcaff.com/2221')
  except HTTPError as e:
    print(e)

返回的是空網頁

if html is None:
    print('沒有找到網頁')

目標標簽在網頁中缺失

try:
    #不存在的標簽
    content = bsObj.nonExistingTag.anotherTag 
  except AttributeError as e:
    print('沒有找到你想要的標簽')
  else:
    if content == None:
      print('沒有找到你想要的標簽')
    else:
      print(content)

實例

if sys.version_info[0] == 2:
  from urllib2 import urlopen # Python 2
  from urllib2 import HTTPError
else:
  from urllib.request import urlopen # Python3
  from urllib.error import HTTPError
from bs4 import BeautifulSoup
import sys


def getTitle(url):
  try:
    html = urlopen(url)
  except HTTPError as e:
    print(e)
    return None
  try:
    bsObj = BeautifulSoup(html.read())
    title = bsObj.body.h2
  except AttributeError as e:
    return None
  return title

title = getTitle("http://www.pythonscraping.com/exercises/exercise1.html")
if title == None:
  print("Title could not be found")
else:
  print(title)

以上全部為本篇文章的全部內容，希望對大家的學習有所幫助，也希望大家多多支持億速云。

向AI問一下細節

91超碰碰碰碰久久久久久综合_超碰av人澡人澡人澡人澡人掠_国产黄大片在线观看画质优化_txt小说免费全本

Python爬蟲包BeautifulSoup異常處理（二）

猜你喜歡

91超碰碰碰碰久久久久久综合_超碰av人澡人澡人澡人澡人掠_国产黄大片在线观看画质优化_txt小说免费全本

Python爬蟲包BeautifulSoup異常處理（二）

猜你喜歡

最新資訊

相關推薦

相關標簽