Using genotype array data to compare multi- and single-sample variant calls and improve variant call sets from deep coverage whole-genome sequencing data.

Using genotype array data to compare multi- and single-sample variant calls and improve variant call sets from deep coverage whole-genome sequencing data.
复制标题

DOI:
10.1093/bioinformatics/btw786
复制
发表时间:
2017-04-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
CAAPA Consortium
CAAPA Consortium
中科院分区:
其他
文献类型:
--
作者:
Shringarpure SS;Mathias RA;Hernandez RD;O'Connor TD;Szpiech ZA;Torres R;De La Vega FM;Bustamante CD;Barnes KC;Taub MA;CAAPA Consortium

文献摘要

被引文献

相似文献

基于下一代测序(NGS)数据的变异调用由于测序、定位和其他错误而容易出现假阳性调用。为了更好地区分真假阳性呼叫,我们提出了一种方法,该方法使用来自测序样本的基因型阵列数据,而不是公共数据,如HapMap或dbSNP,来训练使用随机森林的准确分类器。我们在一组来自美洲非洲裔人群哮喘联盟(CAAPA)的642个非洲裔基因组的变异呼叫上展示了我们的方法,这些基因组进行了高深度(30倍)测序。我们已经应用我们的分类器来比较不同调用方法生成的调用集,包括单样本和多样本调用者。在假阳性率为5%的情况下,我们的方法分别确定了使用Illuminas单样本调用者CASAVA、Real Time Genomics多样本变体调用者和GATK UnifiedGenotyper获得的变体调用的真阳性率为97.5%、95%和99%。由于NGS测序数据可能伴随着相同样本的基因型数据,无论是与测序同时收集的还是来自先前研究的,我们的方法可以在每个数据集上进行训练,以提供比通用方法更准确的位点调用计算验证。此外,我们的方法允许基于等位基因频率进行调整(例如,一套不同的标准来确定罕见变异和常见变异的质量),从而提供了对指示不同频率变异的呼叫质量的测序特征的洞察。代码可在Github上获得:https://github.com/suyashss/variant_validation补充数据可在Bioinformatics在线获得。
Variant calling from next-generation sequencing (NGS) data is susceptible to false positive calls due to sequencing, mapping and other errors. To better distinguish true from false positive calls, we present a method that uses genotype array data from the sequenced samples, rather than public data such as HapMap or dbSNP, to train an accurate classifier using Random Forests. We demonstrate our method on a set of variant calls obtained from 642 African-ancestry genomes from the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA), sequenced to high depth (30X). We have applied our classifier to compare call sets generated with different calling methods, including both single-sample and multi-sample callers. At a False Positive Rate of 5%, our method determines true positive rates of 97.5%, 95% and 99% on variant calls obtained using Illuminas single-sample caller CASAVA, Real Time Genomics multisample variant caller, and the GATK UnifiedGenotyper, respectively. Since NGS sequencing data may be accompanied by genotype data for the same samples, either collected concurrent to sequencing or from a previous study, our method can be trained on each dataset to provide a more accurate computational validation of site calls compared to generic methods. Moreover, our method allows for adjustment based on allele frequency (e.g. a different set of criteria to determine quality for rare versus common variants) and thereby provides insight into sequencing characteristics that indicate call quality for variants of different frequencies. Code is available on Github at: https://github.com/suyashss/variant_validation Supplementary data are available at Bioinformatics online.