python网页请求超时_python爬虫多次请求超时的几种重试方法(6种)

第一种方法

headers = Dict()

url = 'https://www.baidu.com'

try:

proxies = None

response = requests.get(url, headers=headers, verify=False, proxies=None, timeout=3)

except:

# logdebug('requests failed one time')

try:

proxies = None

response = requests.get(url, headers=headers, verify=False, proxies=None, timeout=3)

except:

# logdebug('requests failed two time')

print('requests failed two time')

总结：代码比较冗余，重试try的次数越多，代码行数越多，但是打印日志比较方便

第二种方法

def requestDemo(url，):

headers = Dict()

trytimes = 3 # 重试的次数

for i in range(trytimes):

try:

proxies = None

response = requests.get(url, headers=headers, verify=False, proxies=None, timeout=3)

# 注意此处也可能是302等状态码

if response.status_code == 200:

break

except:

# logdebug(f'requests failed {i}time')

print(f'requests failed {i} time')

总结：遍历代码明显比第一个简化了很多，打印日志也方便

第三种方法

def requestDemo(url， times=1):

headers = Dict()

try:

proxies = None

response = requests.get(url, headers=headers, verify=False, proxies=None, timeout=3)

html = response.text()

# todo 此处处理代码正常逻辑

pass

return html

except:

# logdebug(f'requests failed {i}time')

trytimes = 3 # 重试的次数

if times < trytimes:

times += 1

return requestDemo(url， times)

return 'out of maxtimes'

总结：迭代显得比较高大上，中间处理代码时有其它错误照样可以进行重试；缺点不太好理解，容易出错，另外try包含的内容过多时，对代码运行速度不利。

第四种方法

@retry(3) # 重试的次数 3

def requestDemo(url):

headers = Dict()

proxies = None

response = requests.get(url, headers=headers, verify=False, proxies=None, timeout=3)

html = response.text()

# todo 此处处理代码正常逻辑

pass

return html

def retry(times):

def wrapper(func):

def inner_wrapper(*args, **kwargs):

i = 0

while i < times:

try:

print(i)

return func(*args, **kwargs)

except:

# 此处打印日志 func.__name__ 为say函数

print("logdebug: {}()".format(func.__name__))

i += 1

return inner_wrapper

return wrapper

总结：装饰器优点多种函数复用，使用十分方便

第五种方法

#!/usr/bin/python

# -*-coding='utf-8' -*-

import requests

import time

import json

from lxml import etree

import warnings

warnings.filterwarnings("ignore")

def get_xiaomi():

try:

# for n in range(5): # 重试5次

# print("第"+str(n)+"次")

for a in range(5): # 重试5次

print(a)

url = "https://www.mi.com/"

headers = {

"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3",

"Accept-Encoding": "gzip, deflate, br",

"Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",

"Connection": "keep-alive",

# "Cookie": "xmuuid=XMGUEST-D80D9CE0-910B-11EA-8EE0-3131E8FF9940; Hm_lvt_c3e3e8b3ea48955284516b186acf0f4e=1588929065; XM_agreement=0; pageid=81190ccc4d52f577; lastsource=www.baidu.com; mstuid=1588929065187_5718; log_code=81190ccc4d52f577-e0f893c4337cbe4d|https%3A%2F%2Fwww.mi.com%2F; Hm_lpvt_c3e3e8b3ea48955284516b186acf0f4e=1588929099; mstz=||1156285732.7|||; xm_vistor=1588929065187_5718_1588929065187-1588929100964",

"Host": "www.mi.com",

"Upgrade-Insecure-Requests": "1",

"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/75.0.3770.90 Safari/537.36"

}

response = requests.get(url,headers=headers,timeout=10,verify=False)

html = etree.HTML(response.text)

# print(html)

result = etree.tostring(html)

# print(result)

print(result.decode("utf-8"))

title = html.xpath('//head/title/text()')[0]

print("title==",title)

if "左左" in title:

# print(response.status_code)

# if response.status_code ==200:

break

return title

except:

result = "异常"

return result

if __name__ == '__main__':

print(get_xiaomi())

第六种方法

Python重试模块retrying

# 设置最大重试次数

@retry(stop_max_attempt_number=5)

def get_proxies(self):

r = requests.get('代理地址')

print('正在获取')

raise Exception("异常")

print('获取到最新代理 = %s' % r.text)

params = dict()

if r and r.status_code == 200:

proxy = str(r.content, encoding='utf-8')

params['http'] = 'http://' + proxy

params['https'] = 'https://' + proxy

# 设置方法的最大延迟时间，默认为100毫秒(是执行这个方法重试的总时间)

@retry(stop_max_attempt_number=5,stop_max_delay=50)

# 通过设置为50，我们会发现，任务并没有执行5次才结束！

# 添加每次方法执行之间的等待时间

@retry(stop_max_attempt_number=5,wait_fixed=2000)

# 随机的等待时间

@retry(stop_max_attempt_number=5,wait_random_min=100,wait_random_max=2000)

# 每调用一次增加固定时长

@retry(stop_max_attempt_number=5,wait_incrementing_increment=1000)

# 根据异常重试，先看个简单的例子

def retry_if_io_error(exception):

return isinstance(exception, IOError)

@retry(retry_on_exception=retry_if_io_error)

def read_a_file():

with open("file", "r") as f:

return f.read()

read_a_file函数如果抛出了异常，会去retry_on_exception指向的函数去判断返回的是True还是False，如果是True则运行指定的重试次数后，抛出异常，False的话直接抛出异常。

当时自己测试的时候网上一大堆抄来抄去的，意思是retry_on_exception指定一个函数，函数返回指定异常，会重试，不是异常会退出。真坑人啊！

来看看获取代理的应用(仅仅是为了测试retrying模块)

到此这篇关于python爬虫多次请求超时的几种重试方法的文章就介绍到这了,更多相关python爬虫多次请求超时内容请搜索我们以前的文章或继续浏览下面的相关文章希望大家以后多多支持我们！

本文标题: python爬虫多次请求超时的几种重试方法(6种)

本文地址: http://www.cppcns.com/jiaoben/python/367053.html

python网页请求超时_python爬虫多次请求超时的几种重试方法(6种)相关推荐

python网页抓包_python爬虫入门01：教你在 Chrome 浏览器轻松抓包
通过我们知道了什么是爬虫也知道了爬虫的具体流程那么在我们要对某个网站进行爬取的时候要对其数据进行分析就要知道应该怎么请求就要知道获取的数据是什么样的所以我们要学会怎么抓咪咪! 哦,不对. ...
python爬虫下载重试_python爬虫多次请求超时的几种重试方法(6种)
第一种方法 headers = Dict() url = 'https://www.baidu.com' try: proxies = None response = requests.get(url ...
python网页结构分析图_Python爬虫解析网页的4种方式值得收藏
用Python写爬虫工具在现在是一种司空见惯的事情,每个人都希望能够写一段程序去互联网上扒一点资料下来,用于数据分析或者干点别的事情. 我们知道,爬虫的原理无非是把目标网址的内容下载下来存储到内存中, ...
python爬虫今日头条_python爬虫—分析Ajax请求对json文件爬取今日头条街拍美图
python爬虫-分析Ajax请求对json文件爬取今日头条街拍美图前言本次抓取目标是今日头条的街拍美图,爬取完成之后,将每组图片下载到本地并保存到不同文件夹下.下面通过抓取今日头条街拍美图讲解一 ...
python爬取网页数据软件_python爬虫入门10分钟爬取一个网站
一.基础入门 1.1什么是爬虫爬虫(spider,又网络爬虫),是指向网站/网络发起请求,获取资源后分析并提取有用数据的程序. 从技术层面来说就是通过程序模拟浏览器请求站点的行为,把站点返回的HT ...
python实现get请求模块_python爬虫基于requests模块发起ajax的get请求实现解析
基于requests模块发起ajax的get请求需求:爬取豆瓣电影分类排行榜 https://movie.douban.com/中的电影详情数据用抓包工具捉取使用ajax加载页面的请求鼠标往下 ...
python requests的作用_Python爬虫第一课：requests的使用
requests模块的入门使用注意是requests不是request. 1.为什么使用requests模块,而不是用python自带的urllib requests的底层实现就是urllib re ...
python爬取方式_Python 爬虫入门（三）—— 寻找合适的爬取策略
写爬虫之前,首先要明确爬取的数据.然后,思考从哪些地方可以获取这些数据.下面以一个实际案例来说明,怎么寻找一个好的爬虫策略.(代码仅供学习交流,切勿用作商业或其他有害行为) 1).方式一:直接爬取网站 ...
python 爬网站实例_python爬虫实战：之爬取京东商城实例教程！（含源代码）
前言: 本文主要介绍的是利用python爬取京东商城的方法,文中介绍的非常详细,下面话不多说了,来看看详细的介绍吧. 主要工具 scrapy BeautifulSoup requests 分析步骤 1 ...

python网页请求超时_python爬虫多次请求超时的几种重试方法(6种)

python网页请求超时_python爬虫多次请求超时的几种重试方法(6种)相关推荐

最新文章

热门文章