RepAHR: An improved approach for de novo repeat identication by assembly of the high-frequency reads
RepAHR: An improved approach for de novo repeat identication by assembly of the high-frequency reads
复制标题
RepAHR:通过组装高频读数进行从头重复识别的改进方法
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Jianxin Wang
中科院分区:
文献类型:
--
作者:
Xingyu Liao;Xin Gao;Xiankai Zhang;Fang-Xiang Wu;Jianxin Wang
Background: Repetitive sequences account for a large proportion of eukaryotes.genomes. Identication of repetitive sequences plays a signicant role in many.applications, such as structural variation detection and genome assembly. Many.existing de novo repeat identication pipelines or tools make use of assembly of.high-frequency k-mers to obtain repeats. However, a certain degree of sequence.coverage is required for assemblers to get the desired assemblies. On the other.hand, assemblers cut the reads into shorter k-mers for assembly, which may.destroy the structure of the repetitive regions. For the above reasons, it is.dicult to obtain complete and accurate repetitive regions in the genome by.using existing tools..Result: In this study, we present a new method for de novo repeat identication.by assembly of high-frequency reads (RepAHR). Firstly, RepAHR scans.next-generation sequencing (NGS) reads to nd high-frequency k-mers. Secondly,.RepAHR lters high-frequency reads from whole NGS reads according to certain.rules based on high-frequency k-mer. Thirdly, the high-frequency reads are.assembled to generate repeats using SPAdes, which is considered as an.outstanding genome assembler with NGS sequences..Conlusion: We test RepAHR on ve data sets, and the experimental results.show that RepAHR outperforms RepARK and REPdenovo for detecting repeats.in terms of N50, reference alignment ratio, cover ratio of reference, mask ratio of.Repbase and some other metrics.