Wang Jialin Daily big data Quotations Spark 0018 (November 7, 2015 in Nanning)
The process of Shuffle will be triggered during the reduceByKey operation of Spark. Before Shuffle, a local aggregation process will generate MapPartitionsRDD, then a specific Shuffle will generate ShuffledRDD, and then do a global aggregation to generate the result MapPartitionsRDD.