Get the App
SLTechnology News&Howtos  ›  Internet Technology  › 

How to use hadoop archive to merge small files and mapreduce to reduce the number of map

Shulou Source: shulou.com Published: 2022-06-01 14:25:05 09月26日 Update

This article is about how to use hadoop archive to merge small files and mapreduce to reduce the number of map. The editor thinks it is very practical, so share it with you as a reference and follow the editor to have a look.

As follows: four original files

After hadoop archive:

The command executed is: hadoop archive-archiveName words.har-p / words-r 1 / wordhar

The generated file is in / wordhar/words.har

Where part-0 is a data file

In mapreduce, files at the beginning of the underscore are ignored, that is, the _ SUCCESS,_index,_masterindex in the image above will not be processed.

Then only the data file part-0 will be processed.

The input path for the job setting is

The number of map executed in running mapreduce is 1

Slice into one

The number of map is one

The courseware can also be mapreduce through hadoop archive files.

Thank you for reading! On "how to use hadoop archive to merge small files and mapreduce to reduce the number of map" this article is shared here, I hope the above content can be of some help to you, so that you can learn more knowledge, if you think the article is good, you can share it out for more people to see it!

Tags: Files quantity content data more articles processing good original practical so that first the picture that is command beginning article look knowledge. Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Shulou Tech Info vpn Shulou Information Xiaomi Apple