What is the connection between hive and hadoop
This article mainly explains "what is the relationship between hive and hadoop". The content of the explanation is simple and clear, and it is easy to learn and understand. Please follow the editor's train of thought to study and learn "what is the relationship between hive and hadoop"?
Parsing:
1. Submit the sql to the driver
2. Driver compilation: parsing relevant field table information
3. Go to metastore to query relevant information and return field table information.
4. The information returned by the compiler is sent to the driver
5. The driver sends an execution plan to the execution engine
6. Execution plan (three forms: metastore, namenode, metastore+namenode+mapreduce)
Form 1 > the direct interaction between DDL and metastore on the operation of database tables. For example: create table T1 (name string)
Form 2 > dfs ops fetches data directly from namenode. For example: select * from T1
Form 3 > give job to job tracker, let task tracker execute return execution information + complete job return data information, find namenode to check data.
For example: select * from T1 where col=X
7. Return the result information set
Summary: hive runs on top of hadoop, and some operations require calls to mapreduce in hdfs. Hive metadata is stored in matestore, while non-metadata (such as data in table) is stored on hdfs.
Thank you for your reading, the above is the content of "what is the relationship between hive and hadoop". After the study of this article, I believe you have a deeper understanding of the relationship between hive and hadoop, and the specific use needs to be verified in practice. Here is, the editor will push for you more related knowledge points of the article, welcome to follow!