HAlign: Fast multiple similar DNA/RNA sequence alignment based on the centre star strategy

HAlign: Fast multiple similar DNA/RNA sequence alignment based on the centre star strategy
复制标题

HAlign:基于中心星策略的快速多个相似DNA/RNA序列比对

DOI:
10.1093/bioinformatics/btv177
复制
发表时间:
2015-08-01
期刊:
影响因子:
5.8
通讯作者:
Wang, Guohua
Wang, Guohua
中科院分区:
生物学3区
文献类型:
--
作者:
Zou, Quan;Hu, Qinghua;Wang, Guohua

文献摘要

被引文献

相似文献

动机 多序列比对(MSA)是一项重要的工作,但在大量同源DNA或基因组序列的MSA中存在瓶颈。大多数现有的最先进的软件工具不能处理大规模的数据集,或者它们运行得很慢。同源DNA序列的相似性往往被忽视。缺乏并行化仍然是MSA研究的一个挑战。 结果 我们开发了两个软件工具来解决DNA MSA问题。第一种算法使用trie树来加速中心星星MSA策略。期望的时间复杂度从平方时间降低到线性时间。为了处理大规模数据,Hadoop平台应用了并行性。实验证明了我们提出的方法的性能,包括他们的运行时间,对分数和可扩展性。此外,我们提供了两个大规模的DNA/RNA MSA数据集,以供进一步的测试和研究。
MOTIVATION Multiple sequence alignment (MSA) is important work, but bottlenecks arise in the massive MSA of homologous DNA or genome sequences. Most of the available state-of-the-art software tools cannot address large-scale datasets, or they run rather slowly. The similarity of homologous DNA sequences is often ignored. Lack of parallelization is still a challenge for MSA research. RESULTS We developed two software tools to address the DNA MSA problem. The first employed trie trees to accelerate the centre star MSA strategy. The expected time complexity was decreased to linear time from square time. To address large-scale data, parallelism was applied using the hadoop platform. Experiments demonstrated the performance of our proposed methods, including their running time, sum-of-pairs scores and scalability. Moreover, we supplied two massive DNA/RNA MSA datasets for further testing and research.