M-Grid: a distributed framework for multidimensional indexing and querying of location based data

M-Grid: a distributed framework for multidimensional indexing and querying of location based data
复制标题

DOI:
10.1007/s10619-017-7194-0
复制
发表时间:
2017-03
影响因子:
1.2
通讯作者:
Shashank Kumar;S. Madria;M. Linderman
Shashank Kumar;S. Madria;M. Linderman
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shashank Kumar;S. Madria;M. Linderman

文献摘要

被引文献

相似文献

移动的设备的广泛使用和用户位置信息的真实的时间可用性正在促进新的个性化的、基于位置的应用和服务(LBS)的开发。这些应用程序需要多属性查询处理,可扩展性,以支持数百万用户,实时查询能力和分析大量的数据。云计算辅助了新一代分布式数据库,通常称为键值存储。键值存储旨在从大量数据中提取值,同时具有高度可用性,容错性和可扩展性,因此提供了急需的基础设施来支持LBS。然而,多维数据上的复杂查询不能被有效地处理,因为它们不提供访问多个属性的方法。在本文中,我们提出了M-Grid,一个统一的索引和数据分布框架,使键值存储,以支持多维查询。我们组织一组节点在一个修改后的P-Grid覆盖网络,提供高效的数据分布,容错和查询处理多维数据。为了索引,我们使用基于Hilbert空间填充曲线的线性化技术,该技术保留了数据的局部性,以有效地管理键值存储中的多维数据。我们提出了算法来动态处理范围和最近邻(kNN)查询线性化的值。这消除了维护单独索引表的开销。我们的方法完全独立于底层存储层,可以在任何云基础设施上实现。我们在Amazon EC2上的实验表明,与MapReduce相比,M-Grid的性能提高了三个数量级,是MD-HBase方案的四倍。
The widespread use of mobile devices and the real time availability of user-location information is facilitating the development of new personalized, location-based applications and services (LBSs). Such applications require multi-attribute query processing, scalability for supporting millions of users, real-time querying capability and analyzing large volumes of data. Cloud computing aided a new generation of distributed databases commonly known as key-value stores. Key-value stores were designed to extract values from very large volumes of data while being highly available, fault-tolerant and scalable, hence providing much needed infrastructure to support LBSs. However, complex queries over multidimensional data cannot be processed efficiently as they do not provide means to access multiple attributes. In this paper, we present M-Grid, a unifying indexing and a data distribution framework which enables key-value stores to support multidimensional queries. We organize a set of nodes in a modified P-Grid overlay network which provides efficient data distribution, fault-tolerance and query processing over multidimensional data. To index, we use Hilbert Space Filling Curve based linearization technique which preserves the data locality to efficiently manage multidimensional data in a key-value store. We propose algorithms to dynamically process range andknearest neighbor (kNN) queries on linearized values. This removes the overhead of maintaining a separate index table. Our approach is completely independent from the underlying storage layer and can be implemented on any cloud infrastructure. Our experiments on Amazon EC2 show that M-Grid achieves a performance improvement of three orders of magnitude in comparison to MapReduce and four times to that of MD-HBase scheme.