RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads.

RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads.
复制标题

RepAHR:一种通过组装高频读数进行从头重复识别的改进方法

DOI:
10.1186/s12859-020-03779-w
复制
发表时间:
2020-10-19
期刊:
影响因子:
3
通讯作者:
Wang J
Wang J
中科院分区:
生物学4区
文献类型:
--
作者:
Liao X;Gao X;Zhang X;Wu FX;Wang J

文献摘要

参考文献

被引文献

相似文献

背景重复序列在真核生物基因组中占有很大比例。重复序列的识别在许多应用中起着重要作用,例如结构变异检测和基因组组装。许多现有的从头重复识别管道或工具利用高频链节的组装来获得重复。然而,一定程度的序列覆盖率是汇编程序获得所需汇编所必需的。另一方面,组装器将读段切割成更短的片段进行组装,这可能会破坏重复区域的结构。由于上述原因,它是很难获得完整和准确的重复区域的基因组中,通过使用现有的tools.ResultsIn这项研究中,我们提出了一种新的方法称为RepAHR从头重复识别的组装的高频读取。首先,RepAHR扫描下一代测序(NGS)读数以找到高频k-聚体。其次,RepAHR基于高频k-mer根据一定的规则从整个NGS读段中过滤高频读段。结论在5个数据集上对RepAHR进行了测试,实验结果表明RepAHR在N50、参考比对率、参考覆盖率、Repbase屏蔽率等指标上的重复检测性能优于RepARK和REPdenovo。
BackgroundRepetitive sequences account for a large proportion of eukaryotes genomes. Identification of repetitive sequences plays a significant role in many applications, such as structural variation detection and genome assembly. Many existing de novo repeat identification pipelines or tools make use of assembly of the high-frequencyk-mersto obtain repeats. However, a certain degree of sequence coverage is required for assemblers to get the desired assemblies. On the other hand, assemblers cut the reads into shorterk-mersfor assembly, which may destroy the structure of the repetitive regions. For the above reasons, it is difficult to obtain complete and accurate repetitive regions in the genome by using existing tools.ResultsIn this study, we present a new method called RepAHR for de novo repeat identification by assembly of the high-frequency reads. Firstly, RepAHR scans next-generation sequencing (NGS) reads to find the high-frequencyk-mers. Secondly, RepAHR filters the high-frequency reads from whole NGS reads according to certain rules based on the high-frequencyk-mer. Finally, the high-frequency reads are assembled to generate repeats by using SPAdes, which is considered as an outstanding genome assembler with NGS sequences.ConlusionsWe test RepAHR on five data sets, and the experimental results show that RepAHR outperforms RepARK and REPdenovo for detecting repeats in terms of N50, reference alignment ratio, coverage ratio of reference, mask ratio of Repbase and some other metrics.
DOI: 10.1159/000084979
发表时间: 2005-01-01
影响因子: 1.7
作者:
Jurka, J;Kapitonov, VV;Walichiewicz, J
通讯作者: Walichiewicz, J
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1371/journal.pgen.1002384
发表时间: 2011-12
期刊: PLoS genetics
影响因子: 4.5
作者:
de Koning AP;Gu W;Castoe TA;Batzer MA;Pollock DD
通讯作者: Pollock DD
DOI: 10.1093/bioinformatics/btt054
发表时间: 2013-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Novak, Petr;Neumann, Pavel;Macas, Jiri
通讯作者: Macas, Jiri
重复的DNA和下一代测序:计算挑战和解决方案。
DOI: 10.1038/nrg3117
发表时间: 2011-11-29
期刊: Nature reviews. Genetics
影响因子: --
作者:
通讯作者: --