Get the App
SLTechnology News&Howtos  ›  Database  › 

Discussion on the applicability of Hadoop, Spark, HBase and Redis (2): HBase

Shulou Source: shulou.com Published: 2022-06-01 11:33:48 10月03日 Update

Let's talk about HBase. In this regard, it is often heard that HBase is only suitable for supporting offline analytical applications, especially as a background data source for MapReduce tasks. Hold this view a lot, even in a well-known telecom equipment provider in China, HBase is included in the data analysis product line, and it is clear that HBase is not recommended for online applications. But is this really the case? Let's take a look at some of its major cases: Facebook's messaging applications, including Messages, Chats, Emails and SMS systems, all use Aliwangwang, a WEB version of HBase; Taobao, while Mi chatting with HBase; Xiaomi also uses HBase. Last year, the mobile detailed order query system of a certain provincial company was changed from the original Oracle to a 32-node HBase cluster-brothers, these are the key applications of well-known companies, which is telling enough.

In fact, from the technical characteristics of HBase, it is especially suitable for simple data writing (such as "message class" application) and massive, simple structure data query (such as "detailed single class" application). Among the four HBase applications mentioned above, Facebook message, WEB version of Ali Wangwang and Mixiao all belong to the message application based on data writing, while the mobile phone detailed order query system of the mobile company belongs to the detailed application based on data query.

Another use of HBase is to serve as a background data source for MapReduce to support offline analytical applications. Of course, this is possible, but its performance is open to question. For example, after comparing "Hive over HBase" and "Hive over HDFS" through experiments, superlxw1234 students are surprised to find that except that when using rowkey filtering, the performance based on HBase is slightly better than that based on HDFS, and when using full table scanning and filtering based on value, the performance of directly based on HDFS is much better than that of HBase-this is really a fallacy! However, for this problem, I personally feel that when using rowkey filtering, the higher the degree of filtering, the better the performance of the HBase-based scheme, while the performance of the scheme directly based on HDFS has nothing to do with the degree of filtering. [to be continued]

Although Hadoop is powerful, it is not omnipotent. Http://database.51cto.com/art/201402/429789.htm

2. Performance comparison and analysis of Hiveover HBase and Hiveover HDFS. Http://superlxw1234.iteye.com/blog/2008274

Tags: Application data performance message analysis query company background solution system actual mobile phone data source query system problem Ali Wangwang mobile powerful well-known Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno OPPO Reno Apple Microsoft vpn Shulou Technology