Get the App
SLTechnology News&Howtos  ›  Internet Technology  › 

A general introduction to hadoop ecology

Shulou Source: shulou.com Published: 2022-06-03 05:35:52 10月04日 Update

Key components:

Distributed File Storage system of HDFS:Hadoop

MapReduce:Hadoop 's distributed program computing framework can also be called a programming model.

Hive: SQL-like data Warehouse tool based on Hadoop

HBase: column distributed NoSQL Database based on Hadoop

ZooKeeper: distributed coordination service component

Mahout: a machine learning algorithm library based on distributed computing frameworks such as MapReduce/Flink/Spark

Oozie/Azkaban: workflow scheduling engine

Sqoop: data move in and out tool

Flume: log collection tool

Data processing flow:

A. data acquisition: customize the development of acquisition programs, or use the open source framework Flume or LogStash

B, data preprocessing: custom development MapReduce programs run in Hadoop cluster, or special data collection tools can also preprocess data.

C, data warehouse technology: Hive based on Hadoop

D, data export: Sqoop data import and export tool based on Hadoop

E, data visualization: customize the development of web programs or use products such as Kettle

F, statistical analysis of data: MapReduce in Hadoop is either Hive based on Hadoop, or Spark,Flink

G, process scheduling of the whole process: Oozie/Azkaban tools in the Hadoop ecosystem or other similar open source products

Tags: Data tools distributed programs custom development frameworks development products warehouses processes components scheduling computing preprocessing ecology workflow engines technology databases data statistics Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Apple OPPO Reno Xiaomi NVidia Shulou Technology