How to crawl various document types in Python
This article introduces you how to crawl various document types in Python, the content is very detailed, interested friends can refer to, hope to be helpful to you.
Crawl TXT documents
Under python3, the common method is to get it directly using the urllib.request.urlopen method. After that, we use regular expressions and other methods to retrieve sensitive words.
Crawl CSV documents
Grab word
Methods:
(1) grab remote word docx files using urlopen
(2) convert it to memory byte stream
(3) decompress (docx is a compressed file)
(4) read the decompressed file as xml
(5) find the tags (body content) in xml and process them.
About how to crawl a variety of document types in Python to share here, I hope the above content can be of some help to you, can learn more knowledge. If you think the article is good, you can share it for more people to see.