Novel website crawler
The first day of the novel website crawler
From today on, learn about crawlers and crawl the novel website.
Day one:
Website: http://www.bxwx9.org
Novel: the Great Master
Language: IDEA+java
Jar package: maven project, so put dependencies. Let's study the function of each jar package.
Project structure:
Requirements: get the title and URL from the chapter list of the novel
Principle:
Use Google browser F12 to view the contents of the page and find the element where the chapter list is located.
Use the tag selector to select what you want
The code is as follows:
The solution of Chinese garbled code:
The effect picture of the operation:
Continue tomorrow!