Comparison and integration of deleteriousness prediction methods for nonsynonymous SNVs in whole exome sequencing studies

Comparison and integration of deleteriousness prediction methods for nonsynonymous SNVs in whole exome sequencing studies
复制标题

DOI:
10.1093/hmg/ddu733
复制
发表时间:
2015-04-15
影响因子:
3.5
通讯作者:
Liu, Xiaoming
Liu, Xiaoming
中科院分区:
生物学2区
文献类型:
--
作者:
Dong, Chengliang;Wei, Peng;Liu, Xiaoming

文献摘要

被引文献

相似文献

在全外显子组测序(WES)研究中,准确预测非同义变异的多态性对于区分致病性突变和背景多态性至关重要。虽然目前已经发展了许多预测方法,但在实际应用中,它们的预测结果有时并不一致,其相对优劣也不清楚。为了解决这些问题,我们全面评估了18种当前的准确性评分方法的预测性能,包括11种功能预测评分(PolyPhen-2,SIFT,MutationTaster,Mutation Assessor,FATHMM,LRT,PANTHER,PhD-SNP,SNAP,SNPs&GO和MutPred),3种保守评分(GERP++,SiPhy和PhyloP)和4种集成评分(CADD,PON-P,KGGSeq和CONDEL)。我们发现FATHMM和KGGSeq分别在独立分数和集成分数中具有最高的区分力。此外,为了确保对这些预测分数的无偏性能评估,我们手动收集了三个不同的测试数据集,其中没有调整当前的预测分数。此外,我们开发了两个新的集成分数,集成九个独立的分数和等位基因频率。我们的分数达到了最高的判别力相比,所有的恶性预测分数测试,并显示低的假阳性预测率良性但罕见的非同义变异,这证明了从多个orthopathy的方法相结合的信息的价值。最后,为了促进WES研究中的变体优先级排序,我们预先计算了全外显子组中87347044个可能变体的总体得分,并通过ANNOVAR软件和dbNSFP数据库公开提供。
Accurate deleteriousness prediction for nonsynonymous variants is crucial for distinguishing pathogenic mutations from background polymorphisms in whole exome sequencing (WES) studies. Although many deleteriousness prediction methods have been developed, their prediction results are sometimes inconsistent with each other and their relative merits are still unclear in practical applications. To address these issues, we comprehensively evaluated the predictive performance of 18 current deleteriousness-scoring methods, including 11 function prediction scores (PolyPhen-2, SIFT, MutationTaster, Mutation Assessor, FATHMM, LRT, PANTHER, PhD-SNP, SNAP, SNPs&GO and MutPred), 3 conservation scores (GERP++, SiPhy and PhyloP) and 4 ensemble scores (CADD, PON-P, KGGSeq and CONDEL). We found that FATHMM and KGGSeq had the highest discriminative power among independent scores and ensemble scores, respectively. Moreover, to ensure unbiased performance evaluation of these prediction scores, we manually collected three distinct testing datasets, on which no current prediction scores were tuned. In addition, we developed two new ensemble scores that integrate nine independent scores and allele frequency. Our scores achieved the highest discriminative power compared with all the deleteriousness prediction scores tested and showed low false-positive prediction rate for benign yet rare nonsynonymous variants, which demonstrated the value of combining information from multiple orthologous approaches. Finally, to facilitate variant prioritization in WES studies, we have pre-computed our ensemble scores for 87 347 044 possible variants in the whole-exome and made them publicly available through the ANNOVAR software and the dbNSFP database.