Assessment of computational methods for predicting the effects of missense mutations in human cancers.

Assessment of computational methods for predicting the effects of missense mutations in human cancers.
复制标题

DOI:
10.1186/1471-2164-14-s3-s7
复制
发表时间:
2013
期刊:
影响因子:
4.4
通讯作者:
Zhang Z
Zhang Z
中科院分区:
生物学2区
文献类型:
--
作者:
Gnad F;Baucom A;Mukhyala K;Manning G;Zhang Z

文献摘要

被引文献

相似文献

测序技术的最新进展大大增加了对癌症基因组中突变的鉴定。然而,识别癌症驱动突变仍然是一个重大挑战,因为大多数观察到的错义变化是中性乘客突变。已经开发了各种计算方法来预测氨基酸取代对蛋白质功能的影响,并将突变分类为有害或良性。这些方法包括依赖于进化保守性、结构限制或氨基酸取代的物理化学属性的方法。在这里,我们回顾了现有的方法,并进一步研究了八个工具:SIFT,PolyPhen2,Condel,CHASM,mCluster,logRE,SNAP和MutationAssessor,关于它们的覆盖范围,准确性,可用性和对其他工具的依赖性。使用具有高次要等位基因频率的单核苷酸多态性作为阴性(中性)组进行测试,使用来自COSMIC数据库的复发性突变以及在最近的癌症研究中鉴定的新复发性体细胞突变作为阳性(非中性)组。基于保守性的方法通常在区分中性突变和有害突变方面具有中等高的准确性,而具有综合特征空间的基于机器学习的预测器的性能在使用不同阳性集的评估之间变化。MutationAssessor始终提供最高的准确性。对于某些组合,元预测器稍微提高了所包含的单个方法的性能,但作为独立工具,它的性能并没有超过MutationAssessor。我们对现有工具的独立评估揭示了各种性能差异。癌症训练的方法并没有改善更一般的预测因素。没有一种方法或方法组合的准确率超过81%,这表明驱动突变预测仍有很大的改进空间,也许需要更复杂的特征集成来开发更强大的工具。
Recent advances in sequencing technologies have greatly increased the identification of mutations in cancer genomes. However, it remains a significant challenge to identify cancer-driving mutations, since most observed missense changes are neutral passenger mutations. Various computational methods have been developed to predict the effects of amino acid substitutions on protein function and classify mutations as deleterious or benign. These include approaches that rely on evolutionary conservation, structural constraints, or physicochemical attributes of amino acid substitutions. Here we review existing methods and further examine eight tools: SIFT, PolyPhen2, Condel, CHASM, mCluster, logRE, SNAP, and MutationAssessor, with respect to their coverage, accuracy, availability and dependence on other tools. Single nucleotide polymorphisms with high minor allele frequencies were used as a negative (neutral) set for testing, and recurrent mutations from the COSMIC database as well as novel recurrent somatic mutations identified in very recent cancer studies were used as positive (non-neutral) sets. Conservation-based methods generally had moderately high accuracy in distinguishing neutral from deleterious mutations, whereas the performance of machine learning based predictors with comprehensive feature spaces varied between assessments using different positive sets. MutationAssessor consistently provided the highest accuracies. For certain combinations metapredictors slightly improved the performance of included individual methods, but did not outperform MutationAssessor as stand-alone tool. Our independent assessment of existing tools reveals various performance disparities. Cancer-trained methods did not improve upon more general predictors. No method or combination of methods exceeds 81% accuracy, indicating there is still significant room for improvement for driver mutation prediction, and perhaps more sophisticated feature integration is needed to develop a more robust tool.