Get the App
SLTechnology News&Howtos  ›  Internet Technology  › 

How does Python crawl the joke collection

Shulou Source: shulou.com Published: 2022-06-02 02:59:11 09月28日 Update

Xiaobian to share with you Python how to crawl jokes, I believe most people do not know how, so share this article for everyone's reference, I hope you have a lot of harvest after reading this article, let's go to understand it together!

code

import requestfrom bs4 import BeautifulSoupheaders={ 'user-agent':'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.25 Safari/537.36 Core/1.70.3775.400 QQBrowser/10.6.4208.400'}#request header, crawler masquerade for i in range(0,100): url = 'http://xiaohua.zol.com.cn/detail15/{}.html'.format(i) #Crawler target website html = request.get(url, headers=headers) #Source code returned after request # html.encoding = 'utf-8' soup = BeautifulSoup(html.text, 'lxml')#parse source code if html.status_code==200: #Visit successfully title = soup.select(".article-title")[0].text.replace(' ', '') content = soup.select(".article-text")[0].text.replace(' ', '') with open('D:/xh.txt', 'a',encoding='utf-8') as f: #Save file in D:/xh.txt file f.write(title) f.write(content) f.write('\n\n') f.close() print(title, content) else: #Access failed Continue The above is "Python how to crawl jokes" all the content of this article, thank you for reading! I believe that everyone has a certain understanding, hope to share the content to help everyone, if you still want to learn more knowledge, welcome to pay attention to the industry information channel!

Tags: Articles Daquan jokes content files source code crawlers success not much code most more goals knowledge websites industry information information channels channels references Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Docker Shulou Technology Shulou Tech Info macOS Linux