What's the difference between shuffle and map shuffle?
This article will explain in detail what is the difference between shuffle and map shuffle. The editor thinks it is very practical, so I share it for you as a reference. I hope you can get something after reading this article.
General shuffle structure diagram:
Different tables are completed by different map, and shuffle distributes key with equal conditions to reduce task for execution.
Join is equal to being completed in the reduce phase.
Disadvantages:
The cost is high and the efficiency is slow. All the data needs to be completed by shuffle.
Map shuffle structure diagram:
Mapjoin: join occurs in the map phase, without shuffle
Prerequisites for the use of mapjoin: large tables join small tables (small tables have size limits maximum; live metadata to determine size tables)
The local map task reads the data from the small table to generate HashTable File, and then upload into the distributed cache.
After completing the local map task small table, start the Mapjoin task job to read the large table data, and match each data with the data in the cache
This is the end of the article on "what's the difference between shuffle and map shuffle". I hope the above content can be of some help to you, so that you can learn more knowledge. if you think the article is good, please share it for more people to see.