Spark-yarn mode jar package optimization
In yarn mode, the jar package will be uploaded to yarn to execute the spark program. If you upload it every time, it will take a lot of time. Moreover, if it is Ali Yun's machine, the upload is very slow, and the 180m jar will take more than ten minutes to upload, so you have to upload it to hdfs in advance.
Spark supports the following parameters
Spark.yarn.jars: you can only specify a specific jar package. Before spark1.6.2 (including), you can download a large jar package from the official website and write this jar package, but after 2.0, it becomes a lot of small packages.
Spark.yarn.archive: this folder is supported, but there is one thing to note
.set ("spark.yarn.archive", "hdfs://node2:8020/user/xiaokan/assembly/target/scala-2.11/jars")
.set ("spark.yarn.archive", "hdfs://node2:8020/user/xiaokan/assembly/target/scala-2.11/jars/")
Only the first one is correct, the second is wrong, and the second does not read any jar packets.