A Distributed Indexing Method for Timeline Similarity Query

A Distributed Indexing Method for Timeline Similarity Query
复制标题

一种时间线相似度查询的分布式索引方法

DOI:
10.3390/a11040041
复制
发表时间:
2018-03
期刊:
影响因子:
2.3
通讯作者:
Ma Xiaogang
Ma Xiaogang
中科院分区:
--
文献类型:
--
作者:
He Zhenwen;Ma Xiaogang

文献摘要

参考文献

被引文献

相似文献

时间轴已经使用了几个世纪,近年来随着社交媒体的发展,时间轴的使用越来越广泛。每天,各种智能手机和其他物联网设备都会产生大量与时间相关的数据。这些数据大多可以用时间轴的方式进行管理。然而,如何有效、高效地存储、查询和处理大时间轴数据,特别是基于时间轴相似度的即时推荐,仍然是一个挑战。大多数现有的研究都集中在索引空间和区间数据集,而不是时间数据集。此外,它们中的许多是为集中式系统设计的。迫切需要一种适应并行和分布式计算框架的时间轴索引结构。在本研究中,我们定义了时间轴相似度查询,并在分布式系统中开发了一种新的时间轴索引,称为分布式三角增量树(DTI-Tree)来支持相似度查询。DTI-Tree由一个T-Tree和一个或多个基于Apache Spark的三角形增量分区策略的ti - tree组成。此外,我们还提供了一个开源的时间线基准数据生成器,名为TimelineGenerator,用于为不同的条件生成各种时间线测试数据集。DTI-Tree的构建、插入、删除和相似度查询的实验在一个集群上执行,该集群有两个由TimelineGenerator生成的基准数据集。实验结果表明,dti树为大时间轴数据提供了一种高效的分布式索引解决方案。
Timelines have been used for centuries and have become more and more widely used with the development of social media in recent years. Every day, various smart phones and other instruments on the internet of things generate massive data related to time. Most of these data can be managed in the way of timelines. However, it is still a challenge to effectively and efficiently store, query, and process big timeline data, especially the instant recommendation based on timeline similarities. Most existing studies have focused on indexing spatial and interval datasets rather than the timeline dataset. In addition, many of them are designed for a centralized system. A timeline index structure adapting to parallel and distributed computation framework is in urgent need. In this research, we have defined the timeline similarity query and developed a novel timeline index in the distributed system, called the Distributed Triangle Increment Tree (DTI-Tree), to support the similarity query. The DTI-Tree consists of one T-Tree and one or more TI-Trees based on a triangle increment partition strategy with the Apache Spark. Furthermore, we have provided an open source timeline benchmark data generator, named TimelineGenerator, to generate various timeline test datasets for different conditions. The experiments for DTI-Tree’s construction, insertion, deletion, and similarity queries have been executed on a cluster with two benchmark datasets that are generated by TimelineGenerator. The experimental results show that the DTI-tree provides an effective and efficient distributed index solution to big timeline data.
DOI: 10.1016/j.cageo.2011.02.011
发表时间: 2011-10
期刊: Comput. Geosci.
影响因子: --
作者:
Xiaogang Ma;E. Carranza;Chonglong Wu;F. Meer;Gang Liu
通讯作者: Xiaogang Ma;E. Carranza;Chonglong Wu;F. Meer;Gang Liu
DOI: 10.2298/csis101020035s
发表时间: 2010
期刊: Comput. Sci. Inf. Syst.
影响因子: --
作者:
Bela Stantic;R. Topor;Justin Terry;A. Sattar
通讯作者: Bela Stantic;R. Topor;Justin Terry;A. Sattar
基于分布式索引的并行轨迹搜索
DOI: 10.1016/j.ins.2017.01.016
发表时间: 2017-05
影响因子: 8.1
作者:
Wang Hongzhi;Belhassena Amina
通讯作者: Belhassena Amina
DOI: --
发表时间: 1990-08
期刊: --
影响因子: --
作者:
R. Elmasri;G. Wuu;Y. Kim
通讯作者: R. Elmasri;G. Wuu;Y. Kim
DOI: --
发表时间: 2005-08
期刊: --
影响因子: --
作者:
Jost Enderle;Nicole L. Schneider;T. Seidl
通讯作者: Jost Enderle;Nicole L. Schneider;T. Seidl