Get the App
SLTechnology News&Howtos  ›  Servers  › 

How to use order by,distribute by,sort by,cluster by in hive

Shulou Source: shulou.com Published: 2022-05-31 22:23:47 10月03日 Update

This article mainly introduces how to use order by,distribute by,sort by,cluster by in hive. It is very detailed and has a certain reference value. Friends who are interested must finish it!

Instructions for using order by,distribute by,sort by,cluster by query

Sort meteorological data by year and temperature to ensure that all rows with the same year end up in one reducer partition / / one reduce (massive data, very slow) select year, temperatureorder by year asc, temperature desclimit 100; / multiple reduce (massive data, fast) select year, temperature distribute by year sort by year asc, temperature desclimit 100

Order by (global sort)

Order by will globally sort the input, so there is only one reducer (multiple reducer cannot guarantee global order)

With only one reducer, it takes a long time to calculate when the input size is large.

In hive.mapred.mode=strict mode, it is mandatory to add limit restrictions to reduce the size of reducer data

For example, when limit 100 is limited, if the number of map is 50, the input size of reducer is 100 to 50

Distribute by (similar to split buckets)

The data is divided into different output reduce files according to the fields specified by distribute by.

Sort by (similar to in-bucket sorting)

Sort by is not a global sort, it sorts the data before it enters the reducer.

Therefore, if you sort with sort by and set mapred.reduce.tasks > 1, sort by only guarantees that the output of each reducer is ordered, not globally.

Cluster by

Cluster by not only has the function of distribute by but also has the function of sort by.

However, sorting can only be in reverse order, and the collation cannot be specified as asc or desc.

Therefore, it is often thought that cluster by = distribute by + sort by

The above is all the contents of the article "how to use order by,distribute by,sort by,cluster by in hive". Thank you for reading! Hope to share the content to help you, more related knowledge, welcome to follow the industry information channel!

Tags: Sort data global orderly scale guarantee input content function only multiple year mass article speed output limit different same larger Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno NVidia vpn Docker Huawei Shulou Information