How to use hadoop archive to merge small files and mapreduce to reduce the number of map
This article is about how to use hadoop archive to merge small files and mapreduce to reduce the number of map. The editor thinks it is very practical, so share it with you as a reference and follow the editor to have a look.
As follows: four original files
After hadoop archive:
The command executed is: hadoop archive-archiveName words.har-p / words-r 1 / wordhar
The generated file is in / wordhar/words.har
Where part-0 is a data file
In mapreduce, files at the beginning of the underscore are ignored, that is, the _ SUCCESS,_index,_masterindex in the image above will not be processed.
Then only the data file part-0 will be processed.
The input path for the job setting is
The number of map executed in running mapreduce is 1
Slice into one
The number of map is one
The courseware can also be mapreduce through hadoop archive files.
Thank you for reading! On "how to use hadoop archive to merge small files and mapreduce to reduce the number of map" this article is shared here, I hope the above content can be of some help to you, so that you can learn more knowledge, if you think the article is good, you can share it out for more people to see it!