Matters needing attention for setting up stand-alone Spark in Ubuntu system
For Spark, if you just want to touch and be familiar with it, you can build a stand-alone Spark. The general steps are as follows (I use Ubuntu 14.04under VMWare, which is run under root without considering security issues for the time being):
1. Install Ubuntu 14.04. after installation, you need to disable the firewall (ufw disable), install the SSH server, and enable root users.
2. Download and install JDK-1.8, scala 2.11.8 (need to cooperate with the jar version of spark, this is not necessary for practice), spark, maven (for build, the scala version here needs to be consistent with spark, otherwise a ClassNotDef exception may occur)
3. Configure environment variables in .profile, such as:
Export SPARK_HOME=/root/spark-2.2.0-bin-hadoop2.7
Export SPARK_LOCAL_HOST=192.168.162.132
Export SPARK_MASTER_HOST=192.168.162.132
4. Start spark:
$SPARK_HOME/sbin/start-master.sh
# it must also be started in the case of a stand-alone machine, otherwise there is no worker
$SPARK_HOME/sbin/start-slave.sh
5. Use maven to compile a sample program (of course, sbt can also)
6. Submit your test program as follows:
$SPARK_HOME/spark-submit-- class "class name"-- master spark://IP:Port package file name
In addition, it is important to note that hostnames need to be correctly configured in / etc/hosts and / etc/hostname, otherwise IOException may appear