Get the App
SLTechnology News&Howtos  ›  Internet Technology  › 

Spark Overview and programming Model

Shulou Source: shulou.com Published: 2022-06-02 19:31:43 10月02日 Update

The reason why spark is fast

1. Memory computing

2.DAG

Spark shell has initialized SparkContext. You can call it directly with sc.

Lineage bloodline

RDD wide and narrow dependencies

Narrow dependencies each RDD partition is dependent on at most one child RDD partirion

/ sbin (system binary) are all commands related to system administration.

In some systems, ordinary users do not have permission to execute these commands.

In some systems, the PATH of the average user does not include / sbin

Data.cache data is put in memory

Spark-submit submit Task

Scala code

Package cn.chinahadoop.sparkimport org.apache.spark. {SparkContext, SparkConf} import scala.collection.mutable.ListBufferimport org.apache.spark.SparkContext._/** * Created by chenchao on 14-3-1. * / class Analysis {} object Analysis {def main (args: Array [String]) {if (args.length! = 2) {println ("Usage: java-jar code.jar file_location save_location") System.exit (0)} val conf = new SparkConf () conf.setSparkHome ("/ data/software/crazyjvm/spark") val sc = new SparkContext (conf) val data = sc.textFile (args (0)) Data.cache println (data.count) data.filter (_ .split (''). Length = = 3). Map (_ .split (') (1)). Map ((_ ) .reduceByKey (_ + _) .map (x = > (x.room2, x.room1)) .sortByKey (false) .map (x = > (x.room2, x.room1)) .saveAsTextFile (args (1))}}

Tags: Systems normal memory commands users size code tasks reasons data permissions lineage management models programming Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno MySQL OPPO Reno Xiaomi Huawei macOS