svclassify: a method to establish benchmark structural variant calls.

svclassify: a method to establish benchmark structural variant calls.
复制标题

DOI:
10.1186/s12864-016-2366-2
复制
发表时间:
2016-01-16
期刊:
影响因子:
4.4
通讯作者:
Salit M
Salit M
中科院分区:
生物学2区
文献类型:
--
作者:
Parikh H;Mohiyuddin M;Lam HY;Iyer H;Chen D;Pratt M;Bartha G;Spies N;Losert W;Zook JM;Salit M

文献摘要

被引文献

相似文献

人类基因组包含大小从小的单核苷酸多态性(SNP)到大的结构变体(SV)的变体。瓶内基因组联盟(Genome in a Bottle Consortium)已经开发了用于美国国家标准与技术研究所(NIST)试验参考材料(NA 12878)的高质量基准小变体调用,但该基因组不存在类似的高质量基准SV调用。由于SV调用者输出高度不一致的结果,我们开发了方法来联合收割机组合来自多种测序技术的多种形式的证据,以将候选SV分类为可能的真阳性或假阳性。我们的方法(svclassify)从许多高通量测序技术的一个或多个对齐的bam文件中计算注释,然后使用这些注释构建一类模型,将候选SV分类为可能的真阳性或假阳性。我们首先使用系谱分析来开发一组高置信度的断点分辨大缺失。然后,我们使用svclassify对这些缺失以及来自1000个基因组计划的一组高置信度缺失和来自Spiral Genetics的一组断点解析的复杂插入进行聚类和分类。根据我们的注释,我们发现可能的SV与可能的非SV分开聚集,并且SV聚集成不同类型的缺失。然后,我们开发了一种有监督的一类分类方法,该方法使用随机非SV区域的训练集来确定候选SV是否具有不同于大多数基因组的异常注释。为了测试这种分类方法,我们使用我们的基于谱系的断点解析SV,由1000个基因组计划验证的SV,和基于组装的断点解析插入,沿着使用svviz的半自动可视化。我们发现,来自多种技术的具有高分的候选SV与PCR验证和正交一致性方法MetaSV具有高一致性(99.7%一致性),而具有低分数的候选SV是可疑的。我们分发了一组2676高置信度删除和68高置信度插入高svclassify分数从这些调用集为基准SV调用者。我们希望这些方法对于建立高置信度SV调用的基准样品,其特征在于多种技术特别有用。本文的在线版本(doi:10.1186/s12864-016-2366-2)包含补充材料,可供授权用户使用。
The human genome contains variants ranging in size from small single nucleotide polymorphisms (SNPs) to large structural variants (SVs). High-quality benchmark small variant calls for the pilot National Institute of Standards and Technology (NIST) Reference Material (NA12878) have been developed by the Genome in a Bottle Consortium, but no similar high-quality benchmark SV calls exist for this genome. Since SV callers output highly discordant results, we developed methods to combine multiple forms of evidence from multiple sequencing technologies to classify candidate SVs into likely true or false positives. Our method (svclassify) calculates annotations from one or more aligned bam files from many high-throughput sequencing technologies, and then builds a one-class model using these annotations to classify candidate SVs as likely true or false positives. We first used pedigree analysis to develop a set of high-confidence breakpoint-resolved large deletions. We then used svclassify to cluster and classify these deletions as well as a set of high-confidence deletions from the 1000 Genomes Project and a set of breakpoint-resolved complex insertions from Spiral Genetics. We find that likely SVs cluster separately from likely non-SVs based on our annotations, and that the SVs cluster into different types of deletions. We then developed a supervised one-class classification method that uses a training set of random non-SV regions to determine whether candidate SVs have abnormal annotations different from most of the genome. To test this classification method, we use our pedigree-based breakpoint-resolved SVs, SVs validated by the 1000 Genomes Project, and assembly-based breakpoint-resolved insertions, along with semi-automated visualization using svviz. We find that candidate SVs with high scores from multiple technologies have high concordance with PCR validation and an orthogonal consensus method MetaSV (99.7 % concordant), and candidate SVs with low scores are questionable. We distribute a set of 2676 high-confidence deletions and 68 high-confidence insertions with high svclassify scores from these call sets for benchmarking SV callers. We expect these methods to be particularly useful for establishing high-confidence SV calls for benchmark samples that have been characterized by multiple technologies. The online version of this article (doi:10.1186/s12864-016-2366-2) contains supplementary material, which is available to authorized users.