How does HTML parse the module
In this issue, the editor will bring you about how to parse the HTML module. The article is rich in content and analyzes and narrates it from a professional point of view. I hope you can get something after reading this article.
This is relatively simple, there is nothing to emphasize, if the returned json is taken directly according to the key value, if it is a web page, it uses the html of the lxml module for xpath parsing.
From lxml import html
Import json
Class GetNodeList ():
Def _ init__ (self):
Self.getdivxpath= "/ / div [@ class='demo']"
Def use_xpath (self,source):
If len (source):
Root=html.fromstring (source) # html converted to dom object
Nodelist=root.xpath (self.getdivxpath) # xpath parsing of dom objects
If len (nodelist):
Return nodelist
Return None
Def use_json (self, source,keyname):
If len (source):
Jsonstr=json.loads (source)
Value=jsonstr.get (keyname) # modify according to the specific key value
If len (value):
Return value
Return None
The above is the HTML parsing module shared by the editor. If you happen to have similar doubts, you might as well refer to the above analysis to understand. If you want to know more about it, you are welcome to follow the industry information channel.