Statistical geometry based prediction of nonsynonymous SNP functional effects using random forest and neuro-fuzzy classifiers

Statistical geometry based prediction of nonsynonymous SNP functional effects using random forest and neuro-fuzzy classifiers
复制标题

DOI:
10.1002/prot.21838
复制
发表时间:
2008-06-01
影响因子:
2.9
通讯作者:
Jamison, D. Curtis
Jamison, D. Curtis
中科院分区:
生物学4区
文献类型:
--
作者:
Barenboim, Maxim;Masso, Majid;Jamison, D. Curtis

文献摘要

被引文献

相似文献

鉴于非同义单核苷酸多态(NsSNPs)与遗传病的潜在关系,人们对预测其对蛋白质功能的影响的方法非常感兴趣。目前最先进的监督机器学习算法,如随机森林(RF),训练模型,将蛋白质中的单一氨基酸突变归类为中性或对功能有害的。然而,多态对蛋白质的功能影响通常存在于这两个极端之间。结合模糊逻辑的分类器的使用提供了自然的延伸,以便考虑可能的功能后果的频谱。我们生成了一个人类蛋白质中具有已知三维结构的单一氨基酸替换的数据集。每个变量被唯一地表示为一个特征向量,该特征向量包括通过应用蛋白质结构的Delaunay镶嵌获得的基于计算几何和基于知识的统计势能预测因子。其他属性包括天然氨基酸和替代氨基酸的物理化学性质以及突变残基在已解决结构中的位置的拓扑位置。在由疾病相关和中性nsSNP组成的训练集上对RF算法的分类性能进行了评估,并根据它们的相对重要性对属性进行了排序。类似地,我们评估了自适应神经模糊推理系统(ANFIS)的性能。统计几何预报器的效用与其他研究人员使用的传统结构和进化属性的效用进行了比较,揭示了一种同样有效但互补的方法。在我们的特征集中的所有属性中,统计几何预测器被发现是排名最高的。基于性能的AUC(ROC曲线下面积)测量,当仅使用统计几何特征时,ANFIS和RF模型同样有效。评估AUC、平衡错误率(BER)和马修相关系数(MCC)的十次交叉验证研究表明,我们的RF模型至少可以与SIFT和PolyPhen的成熟方法相媲美。训练好的RF和ANFIS模型随后分别用于预测我们数据集中目前未分类的人类nsSNP的疾病潜力(http://rna.gmu.edu/FuzzySnps/)。
There is substantial interest in methods designed to predict the effect of nonsynonymous single nucleotide polymorphisms (nsSNPs) on protein function, given their potential relationship to heritable diseases. Current state-of-the-art supervised machine learning algorithms, such as random forest (RF), train models that classify single amino acid mutations in proteins as either neutral or deleterious to function. However, it is frequently the case that the functional effect of a polymorphism on a protein resides between these two extremes. The utilization of classifiers that incorporate fuzzy logic provides a natural extension in order to account for the spectrum of possible functional consequences. We generated a dataset of single amino acid substitutions in human proteins having known three-dimensional structures. Each variant was uniquely represented as a feature vector that included computational geometry and knowledge-based statistical potential predictors obtained though application of Delaunay tessellation of protein structures. Additional attributes consisted of physicochemical properties of the native and replacement amino acids as well as topological location of the mutated residue position in the solved structure. Classification performance of the RF algorithm was evaluated on a training set consisting of the disease-associated and neutral nsSNPs taken from our dataset, and attributes were ranked according to their relative importance. Similarly, we evaluated the performance of adaptive neuro-fuzzy inference system (ANFIS). The utility of statistical geometry predictors was compared with that of traditional structural and evolutionary attributes employed by other researchers, revealing an equally effective yet complementary methodology. Among all attributes in our feature set, the statistical geometry predictors were found to be the most highly ranked. On the basis of the AUC (area under the ROC curve) measure of performance, the ANFIS and RF models were equally effective when only statistical geometry features were utilized. Tenfold cross-validation studies evaluating AUC, balanced error rate (BER), and Matthew's correlation coefficient (MCC) showed that our RF model was at least comparable with the well-established methods of SIFT and PolyPhen. The trained RF and ANFIS models were each subsequently used to predict the disease potential of human nsSNPs in our dataset that are currently unclassified (http:// rna.gmu.edu/FuzzySnps/).