A sensitive repeat identification framework based on short and long reads.

A sensitive repeat identification framework based on short and long reads.
复制标题

基于短读和长读的敏感重复识别框架

DOI:
10.1093/nar/gkab563
复制
发表时间:
2021-09-27
影响因子:
14.9
通讯作者:
Wang J
Wang J
中科院分区:
生物学2区
文献类型:
--
作者:
Liao X;Li M;Hu K;Wu FX;Gao X;Wang J

文献摘要

参考文献

被引文献

相似文献

大量研究表明,基因组中的重复区在生物的进化、遗传和变异中起着不可或缺的作用。然而,大多数现有的方法在识别重复序列的准确性和大小方面都不能达到令人满意的性能,因为NGS读数太短而不能识别长重复序列,而SMS(单分子测序)长读数具有高错误率。在这项研究中,我们提出了一种新的识别框架LongRepMarker,该框架基于全局从头组装和基于k-mer的多序列比对,用于精确标记基因组中的长重复序列。LongRepMarker的主要特点是:(I)通过引入条形码连锁阅读和短信长阅读来辅助所有短对端阅读片段的组装,可以在更大程度上识别重复序列;(Ii)通过发现组件或结合体之间的重叠序列,它可以更快、更准确地定位重复序列;(Iii)通过使用多比对唯一的k-MERS而不是高频k-MERS来识别重叠序列中的重复序列,它可以更全面和稳定地获得重复序列;(4)通过应用基于多比对唯一k-MERS的并行比对模型,可以大大优化数据处理的效率;(V)通过采取相应的识别策略,可以识别重复之间发生的结构变化。综合实验结果表明,LongRepMarker能够取得比现有的从头检测方法(https://github.com/BioinformaticsCSU/LongRepMarker).更令人满意的结果
Numerous studies have shown that repetitive regions in genomes play indispensable roles in the evolution, inheritance and variation of living organisms. However, most existing methods cannot achieve satisfactory performance on identifying repeats in terms of both accuracy and size, since NGS reads are too short to identify long repeats whereas SMS (Single Molecule Sequencing) long reads are with high error rates. In this study, we present a novel identification framework, LongRepMarker, based on the global de novo assembly and k-mer based multiple sequence alignment for precisely marking long repeats in genomes. The major characteristics of LongRepMarker are as follows: (i) by introducing barcode linked reads and SMS long reads to assist the assembly of all short paired-end reads, it can identify the repeats to a greater extent; (ii) by finding the overlap sequences between assemblies or chomosomes, it locates the repeats faster and more accurately; (iii) by using the multi-alignment unique k-mers rather than the high frequency k-mers to identify repeats in overlap sequences, it can obtain the repeats more comprehensively and stably; (iv) by applying the parallel alignment model based on the multi-alignment unique k-mers, the efficiency of data processing can be greatly optimized and (v) by taking the corresponding identification strategies, structural variations that occur between repeats can be identified. Comprehensive experimental results show that LongRepMarker can achieve more satisfactory results than the existing de novo detection methods (https://github.com/BioinformaticsCSU/LongRepMarker).
DOI: 10.1007/s11295-018-1257-x
发表时间: 2018-08-01
影响因子: 2.4
作者:
Du, Dongliang;Du, Xiaoyun;Gmitter, Fred G., Jr.
通讯作者: Gmitter, Fred G., Jr.
DOI: 10.1186/s12862-018-1153-x
发表时间: 2018-05-29
影响因子: 3.4
作者:
Kaltenegger E;Leng S;Heyl A
通讯作者: Heyl A
DOI: 10.1186/s12859-018-2376-y
发表时间: 2018-10-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Crescente JM;Zavallo D;Helguera M;Vanzetti LS
通讯作者: Vanzetti LS
DOI: 10.1089/cmb.2012.0021
发表时间: 2012-05-01
影响因子: 1.7
作者:
Bankevich, Anton;Nurk, Sergey;Pevzner, Pavel A.
通讯作者: Pevzner, Pavel A.
DOI: 10.1093/bioinformatics/bti1003
发表时间: 2005-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Edgar, RC;Myers, EW
通讯作者: Myers, EW