Massively Distributed Time Series Indexing and Querying

Massively Distributed Time Series Indexing and Querying
复制标题

DOI:
10.1109/tkde.2018.2880215
复制
发表时间:
2020-01
影响因子:
8.9
通讯作者:
D. Yagoubi;Reza Akbarinia;F. Masseglia;Themis Palpanas
D. Yagoubi;Reza Akbarinia;F. Masseglia;Themis Palpanas
中科院分区:
计算机科学2区
文献类型:
--
作者:
D. Yagoubi;Reza Akbarinia;F. Masseglia;Themis Palpanas

文献摘要

被引文献

相似文献

对于许多依赖于高效和有效的相似查询处理的数据挖掘任务,索引是至关重要的。因此,对大量时间序列进行索引以及高性能的相似查询处理成为人们非常感兴趣的话题。然而,对于跨不同域的许多应用程序来说,要处理的数据量对于一台机器来说可能很难处理,这使得现有的集中式索引解决方案效率低下。我们提出了一种优雅地扩展到数十亿时间序列的并行索引解决方案,以及一种在给定一批查询的情况下高效地利用索引的并行查询处理策略。我们在合成数据和真实世界数据上的实验表明,我们的索引创建算法在不到5小时的时间内处理40亿个时间序列,而最新的集中式算法不能扩展并限制在10亿个时间序列上,其中需要超过5天。此外,我们的分布式查询算法能够高效地处理数十亿时间序列集合上的数百万查询,这要归功于有效的负载平衡机制。
Indexing is crucial for many data mining tasks that rely on efficient and effective similarity query processing. Consequently, indexing large volumes of time series, along with high performance similarity query processing, have became topics of high interest. For many applications across diverse domains though, the amount of data to be processed might be intractable for a single machine, making existing centralized indexing solutions inefficient. We propose a parallel indexing solution that gracefully scales to billions of time series, and a parallel query processing strategy that, given a batch of queries, efficiently exploits the index. Our experiments, on both synthetic and real world data, illustrate that our index creation algorithm works on four billion time series in less than five hours, while the state of the art centralized algorithms do not scale and have their limit on 1 billion time series, where they need more than five days. Also, our distributed querying algorithm is able to efficiently process millions of queries over collections of billions of time series, thanks to an effective load balancing mechanism.