A three-state prediction of single point mutations on protein stability changes.

A three-state prediction of single point mutations on protein stability changes.
复制标题

单点突变对蛋白质稳定性的三态预测发生了变化。

DOI:
10.1186/1471-2105-9-s2-s6
复制
发表时间:
2008-03-26
期刊:
影响因子:
3
通讯作者:
Casadio, Rita
Casadio, Rita
中科院分区:
生物学4区
文献类型:
--
作者:
Capriotti, Emidio;Fariselli, Piero;Rossi, Ivan;Casadio, Rita

文献摘要

被引文献

相似文献

蛋白质结构研究的一个基本问题是突变在多大程度上影响稳定性。这个问题可以从序列和/或结构着手解决。在蛋白质组学和基因组学研究中,单点突变时蛋白质稳定性自由能变化(ΔΔG)的预测可能也有助于注释过程。实验测得的ΔΔG值受标准差所衡量的不确定性影响。大多数ΔΔG值近乎为零(约32%的ΔΔG数据集范围在−0.5到0.5千卡/摩尔之间),而且对于同一突变,ΔΔG的值和符号可能为正也可能为负,这模糊了突变与预期ΔΔG值之间的关系。为了克服这个问题,我们描述了一种新的预测器,它能区分3种突变类别:不稳定突变(ΔΔG < −1.0千卡/摩尔)、稳定突变(ΔΔG > 1.0千卡/摩尔)和中性突变(−1.0 ≤ ΔΔG ≤ 1.0千卡/摩尔)。 在本文中,一种基于蛋白质序列或结构的支持向量机能区分稳定、不稳定和中性突变。我们根据一种三态分类系统对所有可能的替换进行排序,并表明当基于序列信息进行预测时,我们的预测器总体准确率高达56%,当有蛋白质结构可用时为61%,平均相关系数分别为0.27和0.35。这些值比随机预测器的值高约20个百分点。 我们的方法通过采用现有实验数据的热力学可逆性假设,提高了单点蛋白质突变引起的自由能变化的预测质量。通过这种方式,我们既重塑了问题的热力学对称性,又平衡了可用的自由能变化实验测量值的分布。这消除了先前描述的基于不平衡数据集(包含的不稳定突变数量多于稳定突变)训练的方法可能出现的高估情况。
A basic question of protein structural studies is to which extent mutations affect the stability. This question may be addressed starting from sequence and/or from structure. In proteomics and genomics studies prediction of protein stability free energy change (ΔΔG) upon single point mutation may also help the annotation process. The experimental ΔΔG values are affected by uncertainty as measured by standard deviations. Most of the ΔΔG values are nearly zero (about 32% of the ΔΔG data set ranges from −0.5 to 0.5 kcal/mole) and both the value and sign of ΔΔG may be either positive or negative for the same mutation blurring the relationship among mutations and expected ΔΔG value. In order to overcome this problem we describe a new predictor that discriminates between 3 mutation classes: destabilizing mutations (ΔΔG<−1.0 kcal/mol), stabilizing mutations (ΔΔG>1.0 kcal/mole) and neutral mutations (−1.0≤ΔΔG≤1.0 kcal/mole). In this paper a support vector machine starting from the protein sequence or structure discriminates between stabilizing, destabilizing and neutral mutations. We rank all the possible substitutions according to a three state classification system and show that the overall accuracy of our predictor is as high as 56% when performed starting from sequence information and 61% when the protein structure is available, with a mean value correlation coefficient of 0.27 and 0.35, respectively. These values are about 20 points per cent higher than those of a random predictor. Our method improves the quality of the prediction of the free energy change due to single point protein mutations by adopting a hypothesis of thermodynamic reversibility of the existing experimental data. By this we both recast the thermodynamic symmetry of the problem and balance the distribution of the available experimental measurements of free energy changes. This eliminates possible overestimations of the previously described methods trained on an unbalanced data set comprising a number of destabilizing mutations higher than stabilizing ones.