Efficient Aggregation Query Processing for Large-Scale Multidimensional Data by Combining RDB and KVS

Efficient Aggregation Query Processing for Large-Scale Multidimensional Data by Combining RDB and KVS
复制标题

DOI:
10.1007/978-3-319-98809-2_9
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Y. Watari;Atsushi Keyaki;Jun Miyazaki;Masahide Nakamura
Y. Watari;Atsushi Keyaki;Jun Miyazaki;Masahide Nakamura
中科院分区:
其他
文献类型:
--
作者:
Y. Watari;Atsushi Keyaki;Jun Miyazaki;Masahide Nakamura

文献摘要

相似文献

提出了一种高效的大规模多维数据聚合查询处理方法。网络技术的最新发展导致了大量多维数据的产生,例如传感器数据。聚合查询在分析此类数据时起着重要作用。尽管关系数据库(rdb)支持高效的聚合查询,其索引支持更快的查询处理,但是增加数据大小可能会导致瓶颈。另一方面,使用分布式键值存储(D-KVS)是获得数据插入吞吐量的横向扩展性能的关键。然而,由于对索引的支持不足,查询多维数据有时需要进行完整的数据扫描。本文提出的方法将RDB和D-KVS相结合,使两者的优势互补。此外,提出了一种新的技术,其中将数据划分为称为网格的几个子集,并预先计算每个网格的聚合值。该技术通过减少扫描数据量来提高查询处理性能。我们通过将所提方法的性能与当前最先进的方法进行比较来评估所提方法的效率,并表明所提方法在查询和插入方面的性能优于当前方法。
This paper presents a highly efficient aggregation query processing method for large-scale multidimensional data. Recent developments in network technologies have led to the generation of a large amount of multidimensional data, such as sensor data. Aggregation queries play an important role in analyzing such data. Although relational databases (RDBs) support efficient aggregation queries with indexes that enable faster query processing, increasing data size may lead to bottlenecks. On the other hand, the use of a distributed key-value store (D-KVS) is key to obtaining scale-out performance for data insertion throughput. However, querying multidimensional data sometimes requires a full data scan owing to its insufficient support for indexes. The proposed method combines an RDB and D-KVS to use their advantages complementarily. In addition, a novel technique is presented wherein data are divided into several subsets called grids, and the aggregated values for each grid are precomputed. This technique improves query processing performance by reducing the amount of scanned data. We evaluated the efficiency of the proposed method by comparing its performance with current state-of-the-art methods and showed that the proposed method performs better than the current ones in terms of query and insertion.