Accurate Aggregation Query-Result Estimation and Its Efficient Processing on Distributed Key-Value Store

Accurate Aggregation Query-Result Estimation and Its Efficient Processing on Distributed Key-Value Store
复制标题

DOI:
10.1007/978-3-030-27520-4_22
复制
发表时间:
2019-08
期刊:
--
影响因子:
--
通讯作者:
Kosuke Yuki;Atsushi Keyaki;Jun Miyazaki;Masahide Nakamura
Kosuke Yuki;Atsushi Keyaki;Jun Miyazaki;Masahide Nakamura
中科院分区:
其他
文献类型:
--
作者:
Kosuke Yuki;Atsushi Keyaki;Jun Miyazaki;Masahide Nakamura

文献摘要

相似文献

我们提出了四种方法来提高聚合查询结果估计的准确性,使用直方图和/或内核密度估计和查询处理的效率上的分布式键值存储(D-KVS)。最近,聚合查询在分析从传感器、物联网设备等生成的大量多维数据中发挥了关键作用。D-KVS是管理和处理此类大规模多维数据的平台。然而,在D-KVS上查询大规模多维数据有时需要昂贵的数据扫描,因为它对索引的支持不足。由于聚合查询结果并不总是需要准确的,我们的四种方法不仅用于估计准确的查询结果,而不是通过扫描所有数据来获得准确的结果,而且还提高了查询处理性能。首先,我们提出了两种基于核密度估计的方法。为了进一步提高查询结果的估计精度,我们将这两种方法与基于直方图的方案相结合,以便我们可以根据查询和数据分布之间的关系动态地选择最佳的估计方法。我们评估了所提出的方法的效率和准确性,通过比较它们与当前的方法,并表明所提出的方法执行更好。
We propose four methods for improving the accuracy of aggregation query-result estimation using histograms and/or kernel density estimation and the efficiency of query processing on a distributed key-value store (D-KVS). Recently, aggregation queries have played a key role in analyzing a large amount of multidimensional data generated from sensors, Internet-of-Things devices, etc. A D-KVS is a platform to manage and process such large-scale multidimensional data. However, querying large-scale multidimensional data on a D-KVS sometimes requires a costly data scan owing to its insufficient support for indexes. Since aggregation-query results do not always need to be accurate, our four methods are not only for estimating accurate query results rather than obtaining accurate results by scanning all data, but also improving query-processing performance. We first propose two kernel density estimation-based methods. To further improve query-result estimation accuracy, we combined each of these two methods with a histogram-based scheme so that we can dynamically select an optimal estimation method based on the relationship between a query and the data distribution. We evaluated the efficiency and accuracy of the proposed methods by comparing them with a current method and showed that the proposed methods perform better.