Get the App
SLTechnology News&Howtos  ›  Internet Technology  › 

Spark-yarn mode jar package optimization

Shulou Source: shulou.com Published: 2022-06-03 06:05:08 09月24日 Update

In yarn mode, the jar package will be uploaded to yarn to execute the spark program. If you upload it every time, it will take a lot of time. Moreover, if it is Ali Yun's machine, the upload is very slow, and the 180m jar will take more than ten minutes to upload, so you have to upload it to hdfs in advance.

Spark supports the following parameters

Spark.yarn.jars: you can only specify a specific jar package. Before spark1.6.2 (including), you can download a large jar package from the official website and write this jar package, but after 2.0, it becomes a lot of small packages.

Spark.yarn.archive: this folder is supported, but there is one thing to note

.set ("spark.yarn.archive", "hdfs://node2:8020/user/xiaokan/assembly/target/scala-2.11/jars")

.set ("spark.yarn.archive", "hdfs://node2:8020/user/xiaokan/assembly/target/scala-2.11/jars/")

Only the first one is correct, the second is wrong, and the second does not read any jar packets.

Tags: Writing support mode parameters only big piles small packets files folders is in machine program after error Ali Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno NVidia Huawei Shulou Information Docker vpn