Enabling Massive XML-Based Biological Data Management in HBase

Enabling Massive XML-Based Biological Data Management in HBase
复制标题

在 HBase 中实现基于 XML 的海量生物数据管理

DOI:
10.1109/tcbb.2019.2915811
复制
发表时间:
2020-11-01
影响因子:
4.5
通讯作者:
Liu,Yongzhuang
Liu,Yongzhuang
中科院分区:
工程技术3区
文献类型:
--
作者:
Liu,Jian;Liu,Qiuru;Liu,Yongzhuang

文献摘要

被引文献

相似文献

对于希望以可扩展和机器可读的格式提供生物信息学资源的组织来说,以XML格式发布生物数据很有吸引力。在大数据时代,基于XML的海量生物数据管理成为一个具有挑战性的问题。随着基于XML的生物数据集的不断增长,使用传统的声明式查询语言在处理速度和规模方面提供高效的查询能力通常是令人沮丧的。在这项研究中,我们报告了一个新的平台来存储和查询基于XML的海量生物数据集合。首先开发了一个从基于XML的生物数据集合构建HBase表的原型工具,然后提出了一种将XML查询模型转换为MapReduce查询模型的形式化方法。最后,对提出的方法在现有的基于XML的生物数据库上的查询性能进行了评估,表明了该方法的性能优势。基于可扩展标记语言的海量生物数据管理平台的源代码可在https://github.com/lyotvincent/X2H.免费获得
Publishing biological data in XML formats is attractive for organizations who would like to provide their bioinformatics resources in an extensible and machine-readable format. In the era of big data, massive XML-based biological data management is emerged as a challengeable issue. With the continuous growth of the XML-based biological data sets, it is usually frustrating to use traditional declarative query languages to provide efficient query capabilities in terms of processing speed and scale. In this study, we report a novel platform to store and query massive XML-based biological data collections. A prototype tool for constructing HBase tables from XML-based biological data collections is first developed, and then a formal approach to transform the XML query model into the MapReduce query model is proposed. Finally, an evaluation of the query performance of the proposed approach on the existing XML-based biological databases is presented, showing that the performance advantages of the proposed solution. The source code of the massive XML-based biological data management platform is freely available at https://github.com/lyotvincent/X2H.