How to use Python to crawl web content
This article mainly explains "how to use Python to crawl web content", the content of the article is simple and clear, easy to learn and understand, the following please follow the editor's ideas slowly in depth, together to study and learn "how to use Python to crawl web content"!
Write execution code
First, you need to install requests and BeautifulSoup4, and then execute the following code. Import requests from bs4 import BeautifulSoup iurl = 'http://news.sina.com.cn/c/nd/2017-08-03/doc-ifyitapp0128744.shtml' res = requests.get (iurl) res.encoding =' utf-8' # print (len (res.text)) soup = BeautifulSoup (res.text) 'html.parser') # title H1 = soup.select (' # artibodyTitle') [0] .text # Source time_source = soup.select ('. Time-source') [0] .text # Source origin = soup.select ('# artibody p') [0] .text.strip () # original title oriTitle = soup.select ('# artibody p') [1] .text.strip () # content raw_content = soup.select ('# artibody p') [2: 19] content = [] for paragraph in raw_content: content.append (paragraph.text.strip ())'@ .join (content) # responsible Editor ae = soup.select ('.article-editor') [0] .text Thank you for your reading The above is the content of "how to use Python to crawl web content". After the study of this article, I believe you have a deeper understanding of how to use Python to crawl web content, and the specific use needs to be verified in practice. Here is, the editor will push for you more related knowledge points of the article, welcome to follow!