Get the App
SLTechnology News&Howtos  ›  Database  › 

Design principles of Hbase watches

Shulou Source: shulou.com Published: 2022-06-01 14:25:55 10月03日 Update

1. The design of column clusters

The number of clusters should be as few as possible, preferably no more than 3. Because each column cluster exists in an independent HFile, flush and compaction operations are carried out for a Region. When a lot of data in one column cluster needs flush, other column clusters need flush even if they have little data, thus resulting in a large number of unnecessary io operations.

In the case of multi-column clusters, note that the order of magnitude of the data of each column cluster should be the same. If the order of magnitude difference between the two column clusters is too large, the data scanning efficiency of the column clusters with a small order of magnitude will be inefficient.

Put frequently queried and infrequently queried data into different column clusters.

Because column clusters and column names are stored in each Cell of HBase, their names should be as short as possible. For example, replace mycolumnfamily:mycolumnqualifier with FRV Q

2. The design of rowkey

Avoid using incremental numbers or times as rowkey.

If rowkey is an integer, it saves more space in binary than in string.

Reasonably control the length of the rowkey, as short as possible, because the rowkey data will also be stored in each Cell.

If you need to pre-split the table into multiple region, it is best to customize the rules for splitting.

Tags: Data quantity order of magnitude design name as far as possible best query different low consistent two binary multiple case efficiency number way time time Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno OPPO Reno Huawei Xiaomi Shulou Tech Info Docker