Get the App
SLTechnology News&Howtos  ›  Development  › 

How to use Python to crawl web content

Shulou Source: shulou.com Published: 2022-06-03 16:28:06 10月04日 Update

This article mainly explains "how to use Python to crawl web content", the content of the article is simple and clear, easy to learn and understand, the following please follow the editor's ideas slowly in depth, together to study and learn "how to use Python to crawl web content"!

Write execution code

First, you need to install requests and BeautifulSoup4, and then execute the following code. Import requests from bs4 import BeautifulSoup iurl = 'http://news.sina.com.cn/c/nd/2017-08-03/doc-ifyitapp0128744.shtml' res = requests.get (iurl) res.encoding =' utf-8' # print (len (res.text)) soup = BeautifulSoup (res.text) 'html.parser') # title H1 = soup.select (' # artibodyTitle') [0] .text # Source time_source = soup.select ('. Time-source') [0] .text # Source origin = soup.select ('# artibody p') [0] .text.strip () # original title oriTitle = soup.select ('# artibody p') [1] .text.strip () # content raw_content = soup.select ('# artibody p') [2: 19] content = [] for paragraph in raw_content: content.append (paragraph.text.strip ())'@ .join (content) # responsible Editor ae = soup.select ('.article-editor') [0] .text Thank you for your reading The above is the content of "how to use Python to crawl web content". After the study of this article, I believe you have a deeper understanding of how to use Python to crawl web content, and the specific use needs to be verified in practice. Here is, the editor will push for you more related knowledge points of the article, welcome to follow!

Tags: Content web page learning code source title that is ideas situations articles more knowledge knowledge points articles responsibilities follow questions practice push research Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Microsoft macOS NVidia Shulou Information OPPO Reno