Spark Overview and programming Model
The reason why spark is fast
1. Memory computing
2.DAG
Spark shell has initialized SparkContext. You can call it directly with sc.
Lineage bloodline
RDD wide and narrow dependencies
Narrow dependencies each RDD partition is dependent on at most one child RDD partirion
/ sbin (system binary) are all commands related to system administration.
In some systems, ordinary users do not have permission to execute these commands.
In some systems, the PATH of the average user does not include / sbin
Data.cache data is put in memory
Spark-submit submit Task
Scala code
Package cn.chinahadoop.sparkimport org.apache.spark. {SparkContext, SparkConf} import scala.collection.mutable.ListBufferimport org.apache.spark.SparkContext._/** * Created by chenchao on 14-3-1. * / class Analysis {} object Analysis {def main (args: Array [String]) {if (args.length! = 2) {println ("Usage: java-jar code.jar file_location save_location") System.exit (0)} val conf = new SparkConf () conf.setSparkHome ("/ data/software/crazyjvm/spark") val sc = new SparkContext (conf) val data = sc.textFile (args (0)) Data.cache println (data.count) data.filter (_ .split (''). Length = = 3). Map (_ .split (') (1)). Map ((_ ) .reduceByKey (_ + _) .map (x = > (x.room2, x.room1)) .sortByKey (false) .map (x = > (x.room2, x.room1)) .saveAsTextFile (args (1))}}