MATCHCLIP: locate precise breakpoints for copy number variation using CIGAR string by matching soft clipped reads.

MATCHCLIP: locate precise breakpoints for copy number variation using CIGAR string by matching soft clipped reads.
复制标题

DOI:
10.3389/fgene.2013.00157
复制
发表时间:
2013
影响因子:
3.7
通讯作者:
Li H
Li H
中科院分区:
生物学3区
文献类型:
--
作者:
Wu Y;Tian L;Pirastu M;Stambolian D;Li H

文献摘要

参考文献

相似文献

拷贝数变异(CNVs)与许多复杂疾病有关。下一代测序数据使人们能够确定精确的CNV断点,以更好地了解潜在的分子机制,并设计更有效的检测方法。使用CIGAR字符串的读取,我们开发了一种方法,可以确定确切的CNV断点,并在断点是在一个重复的区域,该方法报告的断点可以滑动的范围。我们的方法使用覆盖CNV的断点的读段的位置和CIGAR字符串来识别CNV的断点。在其3′(右)端具有长软剪切部分(在CIGAR中表示为S)的读段可用于鉴定断裂点的5′(左)侧,并且在5′端具有长S部分的读段可用于鉴定3′侧的断裂点。为了确保两种类型的读取覆盖相同的CNV,我们要求重叠的共同串包括两个软剪切部分。当CNV在相同的重复区域中开始和结束时,它的断点不是唯一的,在这种情况下,我们的方法报告了断点的最左边的位置和断点可以在不改变变体序列的情况下递增的范围。我们已经在C++包中实现了用于当前Illumina Miseq和Hiseq平台的方法,用于全基因组和外显子测序。我们的模拟研究表明,我们的方法相比,与其他类似的方法在真正的发现率,假阳性率和断点的准确性。我们从真实的应用中得到的结果表明,检测到的CNV与接合性和读取深度信息一致。该软件包可在http://statgene.med.upenn.edu/softprog.html上获得。
Copy number variations (CNVs) are associated with many complex diseases. Next generation sequencing data enable one to identify precise CNV breakpoints to better under the underlying molecular mechanisms and to design more efficient assays. Using the CIGAR strings of the reads, we develop a method that can identify the exact CNV breakpoints, and in cases when the breakpoints are in a repeated region, the method reports a range where the breakpoints can slide. Our method identifies the breakpoints of a CNV using both the positions and CIGAR strings of the reads that cover breakpoints of a CNV. A read with a long soft clipped part (denoted as S in CIGAR) at its 3′(right) end can be used to identify the 5′(left)-side of the breakpoints, and a read with a long S part at the 5′ end can be used to identify the breakpoint at the 3′-side. To ensure both types of reads cover the same CNV, we require the overlapped common string to include both of the soft clipped parts. When a CNV starts and ends in the same repeated regions, its breakpoints are not unique, in which case our method reports the left most positions for the breakpoints and a range within which the breakpoints can be incremented without changing the variant sequence. We have implemented the methods in a C++ package intended for the current Illumina Miseq and Hiseq platforms for both whole genome and exon-sequencing. Our simulation studies have shown that our method compares favorably with other similar methods in terms of true discovery rate, false positive rate and breakpoint accuracy. Our results from a real application have shown that the detected CNVs are consistent with zygosity and read depth information. The software package is available at http://statgene.med.upenn.edu/softprog.html.
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1038/nature08516
发表时间: 2010-04-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/nar/gkn835
发表时间: 2009-01
影响因子: 14.9
作者:
Basu SN;Kollu R;Banerjee-Basu S
通讯作者: Banerjee-Basu S
DOI: 10.1093/bioinformatics/btp208
发表时间: 2009-06-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Sindi S;Helman E;Bashir A;Raphael BJ
通讯作者: Raphael BJ