Handling distributed XML queries over large XML data based on MapReduce framework

Handling distributed XML queries over large XML data based on MapReduce framework
复制标题

DOI:
10.1016/j.ins.2018.04.028
复制
发表时间:
2018-07
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Hongjie Fan;Zhiyi Ma;Dianhui Wang;Junfei Liu
Hongjie Fan;Zhiyi Ma;Dianhui Wang;Junfei Liu
中科院分区:
其他
文献类型:
--
作者:
Hongjie Fan;Zhiyi Ma;Dianhui Wang;Junfei Liu

文献摘要

相似文献

随着可扩展标记语言(XML)文档的增加,文献中提出了许多查询方法。XML查询和Twig模式查询是XML操作的两种基本方式,直接影响XML操作的效率。分布式操作海量XML数据是一个挑战。本文的目的是开发一种高效的分布式XML查询处理方法,使用MapReduce,同时处理多个查询的大容量的XML数据。首先,我们将一个大规模的XML数据文件分割成文件分割,并将它们放在一个分布式存储系统中。然后,我们提出了一个有效的算法来计算不同的片段的文档树使用MapReduce框架并行。为了有效地处理大量的XML数据,我们构建了一个分区索引,并对特定的查询使用随机访问机制。实验结果表明,我们提出的方法是有效的,并且具有良好的可扩展性。
With the increase in available extensible markup language (XML) documents, numerous approaches to querying have been proposed in the literature. XPath queries and Twig pattern queries are the two basic approaches, directly affecting the efficiency of XML operations. Distributive manipulation of massive XML data is challenging. This paper aims to develop an efficient distributed XML query processing method using MapReduce, which simultaneously processes several queries on large volumes of XML data. First, we split up a large-scale XML data file into file-splits and put them in a distributed storage system. Then, we present an efficient algorithm to compute different fragments of the document tree using the MapReduce framework in parallel. In order to efficiently handle a large amount of XML data, we built a partition index and used a random access mechanism for specific queries. The experiment results show that our proposed approach is efficient with good scalability.