Get the App
SLTechnology News&Howtos  ›  Internet Technology  › 

How to realize union Operation in Spark SQL

Shulou Source: shulou.com Published: 2022-06-02 09:07:11 10月05日 Update

Today, I will talk to you about how to implement union operation in Spark SQL. Many people may not know much about it. In order to let everyone know more, Xiaobian summarized the following contents for everyone. I hope you can gain something according to this article.

union all is a direct connection, get all values, records may have duplicates

union is a unique value, records are not duplicated

1. The syntax of UNION is as follows:

[SQL Statement 1]

UNION

[SQL Statement 2]

The syntax of UNION ALL is as follows:

[SQL Statement 1]

UNION ALL

[SQL Statement 2]

Comparative summary:

The UNION and UNION ALL keywords both combine two result sets into one, but both differ in terms of usage and efficiency.

1. Processing duplicate results: UNION filters out duplicate records after table linking, Union All does not remove duplicate records.

2, sorting processing: Union will sort according to the order of the fields;UNION ALL simply combines the two results and returns.

In terms of efficiency, UNION ALL is much faster than UNION, so if you can confirm that the two result sets merged do not contain duplicate data and do not need sorting, then use UNION ALL.

Spark SQL

In fact, the DataSet API of Spark SQL has no union all operation, only union operation, and its union operation is union all operation.

In this case, to implement the union operation, you need to add the distinct operation after union.

sales.union(sales).show()

The output is duplicate data

The action needs to be changed to:

sales.union(sales).distinct().show()

After reading the above, do you have any further understanding of how to implement union operations in Spark SQL? If you still want to know more knowledge or related content, please pay attention to the industry information channel, thank you for your support.

Tags: Results statements two content sorting efficiency data syntax processing differences keys keywords just only fields actually that is more different Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno macOS Shulou Information OPPO Reno Linux vpn