Identifying micro-inversions using high-throughput sequencing reads.

Identifying micro-inversions using high-throughput sequencing reads.
复制标题

DOI:
10.1186/s12864-015-2305-7
复制
发表时间:
2016-01-11
期刊:
影响因子:
4.4
通讯作者:
Zhu H
Zhu H
中科院分区:
生物学2区
文献类型:
--
作者:
He F;Li Y;Tang YH;Ma J;Zhu H

文献摘要

被引文献

相似文献

短于读段长度的DNA片段的倒位的鉴定(例如,100 bp),定义为微倒位(MI),对于下一代测序读数仍然具有挑战性。MI是一种重要的基因组变异,可能在遗传疾病的发生中发挥作用。然而,当前的对准方法通常对检测MI不敏感。在这里,我们开发了一种新的工具,MID(微倒置检测器),使用下一代测序读取来识别人类基因组中的MI。基于动态规划寻路方法设计了MID算法。MID与其他变体检测工具的不同之处在于,MID可以处理未映射读段中的小MI和多个断点。此外,MID通过整合多个样本来提高低覆盖率数据的可靠性。我们的评估表明,MID优于Gustaf,Gustaf目前可以检测30 bp至500 bp的倒位。据我们所知,MID是第一种可以从未映射的短下一代测序读数中有效且可靠地鉴定MI的方法。MID在低覆盖率数据上是可靠的,适用于大规模项目,如1000个基因组计划(1 KGP)。MID从1 KGP中鉴定出了以前未知的MI,这些MI与人类基因组中的基因和调控元件重叠。我们还从癌细胞系百科全书(CCLE)中鉴定了癌细胞系中的MI。因此,我们的工具有望有助于改善MI作为人类基因组中一种遗传变异的研究。源代码可以从http://cqb.pku.edu.cn/ZhuLab/MID下载。本文的在线版本(doi:10.1186/s12864-015-2305-7)包含补充材料,可供授权用户使用。
The identification of inversions of DNA segments shorter than read length (e.g., 100 bp), defined as micro-inversions (MIs), remains challenging for next-generation sequencing reads. It is acknowledged that MIs are important genomic variation and may play roles in causing genetic disease. However, current alignment methods are generally insensitive to detect MIs. Here we develop a novel tool, MID (Micro-Inversion Detector), to identify MIs in human genomes using next-generation sequencing reads. The algorithm of MID is designed based on a dynamic programming path-finding approach. What makes MID different from other variant detection tools is that MID can handle small MIs and multiple breakpoints within an unmapped read. Moreover, MID improves reliability in low coverage data by integrating multiple samples. Our evaluation demonstrated that MID outperforms Gustaf, which can currently detect inversions from 30 bp to 500 bp. To our knowledge, MID is the first method that can efficiently and reliably identify MIs from unmapped short next-generation sequencing reads. MID is reliable on low coverage data, which is suitable for large-scale projects such as the 1000 Genomes Project (1KGP). MID identified previously unknown MIs from the 1KGP that overlap with genes and regulatory elements in the human genome. We also identified MIs in cancer cell lines from Cancer Cell Line Encyclopedia (CCLE). Therefore our tool is expected to be useful to improve the study of MIs as a type of genetic variant in the human genome. The source code can be downloaded from: http://cqb.pku.edu.cn/ZhuLab/MID. The online version of this article (doi:10.1186/s12864-015-2305-7) contains supplementary material, which is available to authorized users.