RepLong: de novo repeat identification using long read sequencing data

RepLong: de novo repeat identification using long read sequencing data
复制标题

RepLong:使用长读测序数据从头重复识别

DOI:
10.1093/bioinformatics/btx717
复制
发表时间:
2018-04-01
期刊:
影响因子:
5.8
通讯作者:
Zhu, Zexuan
Zhu, Zexuan
中科院分区:
生物学3区
文献类型:
--
作者:
Guo, Rui;Li, Yan-Ran;Zhu, Zexuan

文献摘要

被引文献

相似文献

动机 重复元件的鉴定在基因组组装和系统发育分析中具有重要意义。现有的从头重复序列鉴定方法利用短读段的使用在鉴定长重复序列方面是无效的。由于长读段更可能完全覆盖重复区域,因此使用长读段更有利于识别长重复。 结果 在这项研究中,我们提出了一种新的从头重复序列识别方法,即RepLong基于PacBio长读段。考虑到映射到重复区域的读段彼此高度重叠,重复元件的鉴定等同于读段之间的共有重叠的发现,这可以进一步转化为读段重叠网络中的社区检测问题。在RepLong中,我们首先基于读段的成对比对构建读段重叠的网络,其中每个顶点指示读段,并且边缘指示相应的两个读段之间的实质性重叠。其次,基于网络模块度优化,提取出内部连通度大于内部连通度的社区。最后,提取每个社区中的代表性读数以形成重复文库。通过对果蝇和人类长片段测序数据与基于基因组和基于短片段的方法的比较研究,证明了RepLong在识别长重复序列方面的有效性。RepLong可以处理覆盖率较低的数据,并作为现有方法的补充解决方案,以提高长读段测序数据的重复识别性能。 可用性和实施 RepLong的软件可在https://github.com/ruiguo-bio/replong上免费获得。 接触 ywsun@szu.edu.cn或zhuzx@szu.edu.cn。 补充资料 补充数据可在Bioinformatics在线获得。
Motivation The identification of repetitive elements is important in genome assembly and phylogenetic analyses. The existing de novo repeat identification methods exploiting the use of short reads are impotent in identifying long repeats. Since long reads are more likely to cover repeat regions completely, using long reads is more favorable for recognizing long repeats. Results In this study, we propose a novel de novo repeat elements identification method namely RepLong based on PacBio long reads. Given that the reads mapped to the repeat regions are highly overlapped with each other, the identification of repeat elements is equivalent to the discovery of consensus overlaps between reads, which can be further cast into a community detection problem in the network of read overlaps. In RepLong, we first construct a network of read overlaps based on pair-wise alignment of the reads, where each vertex indicates a read and an edge indicates a substantial overlap between the corresponding two reads. Secondly, the communities whose intra connectivity is greater than the inter connectivity are extracted based on network modularity optimization. Finally, representative reads in each community are extracted to form the repeat library. Comparison studies on Drosophila melanogaster and human long read sequencing data with genome-based and short-read-based methods demonstrate the efficiency of RepLong in identifying long repeats. RepLong can handle lower coverage data and serve as a complementary solution to the existing methods to promote the repeat identification performance on long-read sequencing data. Availability and implementation The software of RepLong is freely available at https://github.com/ruiguo-bio/replong. Contact ywsun@szu.edu.cn or zhuzx@szu.edu.cn. Supplementary information Supplementary data are available at Bioinformatics online.