Hadoop fully distributed deployment
Hadoop fully distributed deployment I. Overview of concepts:
Is a reliable, scalable, distributed computing open source software.
Is a framework that allows big data and distributed processing across computer clusters, using a simple programming model (mapreduce)
Can scale from a single server to thousands of hosts, with each node providing computing and storage capabilities.
Do not rely on hardware to process HA, and implement it at the application level
Property 4V:
Volumn has a large mass
Velocity is fast.
Variaty has many styles
Low value density of value
Module:
Hadoop common common class library, which supports other modules
HDFS hadoop distributed file system,hadoop distributed file system
Hadoop yarn Job scheduling and Resource Management Framework
Hadoop mapreduce is based on the parallel processing technology of large data sets in yarn system.
Installation and deployment 2.1host planning host name IP address installation node application hadoop-1172.20.2.203namenode/datanode/nodemanagerhadoop-2172.20.2.204secondarynode/datanode/nodemanagerhadoop-3172.20.2.205resourcemanager/datanode/nodemanager2.2 deployment 2.2.1 basic environment configuration
a. Configure the java environment
Yum install java-1.8.0-openjdk.x86_64 java-1.8.0-openjdk-devel-ycat > / etc/profile.d/java.sh/etc/hosts