Accelerating RepeatClassifier Based on Spark and Greedy Algorithm with Dynamic Upper Boundary

Accelerating RepeatClassifier Based on Spark and Greedy Algorithm with Dynamic Upper Boundary
复制标题

基于 Spark 和动态上边界贪婪算法的加速重复分类器

DOI:
10.1101/2021.06.03.446998
复制
发表时间:
2021-06
期刊:
bioRxiv
影响因子:
--
通讯作者:
Jianxin Wang
Jianxin Wang
中科院分区:
其他
文献类型:
--
作者:
Kang Hu;Xingyu Liao;You Zou;Jianxin Wang

文献摘要

参考文献

相似文献

转座因子(Transposable elements,TE)是基因组序列的重要组成部分,占小麦基因组序列的90%,在基因组的组织和进化中起着重要作用。推广转座因子的无监督注释具有重要意义。分类是TE注释的重要步骤,它概括了原始重复序列的类型或机制。RepeatClassifier是一个基本的基于同源性的分类工具,它将TE家族与重复蛋白质数据库(DB)和RepeatMasker文库进行比较。不幸的是,RepeatClassifier效率低下,需要几天时间才能对大型基因组的重复序列进行分类。为此,本文提出了基于Spark的重复分类器(SRC),该算法利用动态上边界贪婪算法(GDUB)进行数据划分和负载均衡,并利用Spark提高了重复分类器的并行性。实验结果表明,SRC不仅可以保证与RepeatClassifier相同的精度水平,而且与RepeatClassifier相比,可以实现42-88倍的加速。同时,SRC在处理长度分布不均衡的输入数据集时表现出了良好的并行性能。SRC可在https://github.com/BioinformaticsCSU/SRC上公开获取。
Transposable elements (TEs) represent quantitatively important components of genome sequences (e.g. 90% of the wheat genome), and play important roles in genome organization and evolution. The promotion of unsupervised annotation of transposable elements is of great significance. Classification is an important step in TE annotation, which summarize the information about the type or mechanism for the raw repetitive sequences. RepeatClassifier is a basic homology-based classification tool which compares the TE families to both the Repeat Protein Database (DB) and libraries of RepeatMasker. Unfortunately, RepeatClassifier is inefficient and takes a few days to classify the repetitive sequences of large genomes. Hence, we proposed Spark-based RepeatClassifier (SRC) which uses Greedy Algorithm with Dynamic Upper Boundary (GDUB) for data division and load balancing, and Spark to improve the parallelism of RepeatClassifier. Experimental results show that SRC can not only ensure the same level of accuracy as that of RepeatClassifier, but also achieve 42-88 times of acceleration compared to RepeatClassifier. At the same time, SRC shows excellent parallel performance when dealing with input datasets with unbalanced length distribution. SRC is publicly available at https://github.com/BioinformaticsCSU/SRC.
DOI: 10.1371/journal.pcbi.0010043
发表时间: 2005-09
影响因子: 4.3
作者:
Li R;Ye J;Li S;Wang J;Han Y;Ye C;Wang J;Yang H;Yu J;Wong GK;Wang J
通讯作者: Wang J
DOI: 10.1159/000084979
发表时间: 2005-01-01
影响因子: 1.7
作者:
Jurka, J;Kapitonov, VV;Walichiewicz, J
通讯作者: Walichiewicz, J
DOI: 10.1186/s12862-018-1153-x
发表时间: 2018-05-29
影响因子: 3.4
作者:
Kaltenegger E;Leng S;Heyl A
通讯作者: Heyl A
DOI: 10.1128/mcb.13.5.2802-2814.1993
发表时间: 1993-05
影响因子: 5.3
作者:
Qin Lu;Lori L. Wallrath;H. Granok;Sarah C. R. Elgin
通讯作者: Qin Lu;Lori L. Wallrath;H. Granok;Sarah C. R. Elgin
重复的DNA和下一代测序:计算挑战和解决方案。
DOI: 10.1038/nrg3117
发表时间: 2011-11-29
期刊: Nature reviews. Genetics
影响因子: --
作者:
通讯作者: --