What are the components of the Hadoop ecosystem
This article will explain in detail what are the components of the Hadoop ecosystem. The editor thinks it is very practical, so I share it for you as a reference. I hope you can get something after reading this article.
The components of the Hadoop ecosystem include:
HDFS: distributed file system
YARN: resource management and scheduling
MapReduce: parallel computing framework
HBase: scalable distributed NoSQL database
Hive: suitable for ETL big data warehouse, support SQL query language, based on MapReduce
Impala: a new query system that provides interactive SQL query
ZooKeeper: distributed Application Coordination Service
Spark: distributed memory computing engine that supports ETL, machine learning, Streaming and graph computing
Flume: distributed log collection and aggregation system
Pig: large-scale data Analysis platform
PrestoDB: big data's distributed SQL query engine
Phoenix: is the SQL driver of Hbase
Drill: a tool for speeding up Hadoop data query
Hue is a graphical user interface for operating and developing Hadoop applications.
Divided by service system:
Computing cloud: virtual host / elastic computing / load balancing QLB
Storage cloud: GlusterFS/Swift/FastDFS/ production storage / cloud disk
Service Cloud-Database: MySQL/Couchbase/Redis/MongoDB
Service Cloud-distributed Middleware: RPC/MQ/ZooKeeper
Service Cloud-Hadoop:HDFS/MR/Hive/HBase
Service cloud-real-time computing: Spark/Storm/ real-time log collection and analysis
This is the end of this article on "what are the components of the Hadoop ecosystem?". I hope the above content can be of some help to you, so that you can learn more knowledge. if you think the article is good, please share it for more people to see.