Performance optimization tips-cluster dimension table
When fact table and dimension table are associated, dimension table needs frequent random access, so dimension table should be stored in memory as much as possible to improve the performance of association calculation. If the dimension table is too large to fit in the memory of a single machine, it should be considered to read the dimension table into the memory of multiple machines in a cluster manner. The following example illustrates the use of cluster dimension tables.
Suppose there are two compute nodes, 127.0.0.1: 8281 and 127.0.0.1: 8282. Execute the following script to load the product table into node machine memory:
AB1=["127.0.0.1:8281","127.0.0.1:8282"]
2fork [[1,20000000],[20000000,40000000]];A1=connect("demo1").query@x("select productID,name,price from product where productID>? and productID