Quantification of biases in predictions of protein stability changes upon mutations

Quantification of biases in predictions of protein stability changes upon mutations
复制标题

DOI:
10.1093/bioinformatics/bty348
复制
发表时间:
2018-11-01
期刊:
影响因子:
5.8
通讯作者:
Rooman, Marianne
Rooman, Marianne
中科院分区:
生物学3区
文献类型:
--
作者:
Pucci, Fabrizio;Bernaerts, Katrien, V;Rooman, Marianne

文献摘要

被引文献

相似文献

动机:预测点突变时蛋白质稳定性变化的生物信息学工具在过去几十年中取得了很大进展,并且已经变得足够准确和快速,使得计算诱变实验变得可行,甚至在蛋白质组规模上也是如此。尽管取得了这些成就,但它们仍然面临必须解决的重要问题,以进一步提高其性能并利用它们加深我们对蛋白质折叠和稳定性机制的了解。这些问题之一是他们对学习数据集的偏见,这些数据集以不稳定突变为主,导致不稳定突变的预测比稳定突变更好。结果:我们彻底分析了点突变(Delta Delta G(0))折叠自由能变化预测的偏差,并提出了一些无偏差的解决方案。我们首先通过收集野生型和突变蛋白结构均可用的突变,构建实验测量的 Delta Delta G(0) 数据集 S-sym,该数据集具有相同数量的稳定和不稳定突变。在这个平衡数据集上,我们评估了 15 个广泛使用的 Delta Delta G(0) 预测器的性能。在令人惊讶地观察到几乎所有这些方法,尤其是那些使用黑盒机器学习的方法,都强烈偏向于不稳定的突变之后,我们提出了一种优雅的方法来解决偏差问题,即在模型结构上强加逆突变下的物理对称性,并在 PoPMuSiC(sym) 中实现。这个新的预测器在准确性和无偏差之间进行了有效的权衡。讨论了进一步改进预测器的一些最终考虑因素和建议。
Motivation: Bioinformatics tools that predict protein stability changes upon point mutations have made a lot of progress in the last decades and have become accurate and fast enough to make computational mutagenesis experiments feasible, even on a proteome scale. Despite these achievements, they still suffer from important issues that must be solved to allow further improving their performances and utilizing them to deepen our insights into protein folding and stability mechanisms. One of these problems is their bias toward the learning datasets which, being dominated by destabilizing mutations, causes predictions to be better for destabilizing than for stabilizing mutations.Results: We thoroughly analyzed the biases in the prediction of folding free energy changes upon point mutations (Delta Delta G(0)) and proposed some unbiased solutions. We started by constructing a dataset S-sym of experimentally measured Delta Delta G(0)s with an equal number of stabilizing and destabilizing mutations, by collecting mutations for which the structure of both the wild-type and mutant protein is available. On this balanced dataset, we assessed the performances of 15 widely used Delta Delta G(0) predictors. After the astonishing observation that almost all these methods are strongly biased toward destabilizing mutations, especially those that use black-box machine learning, we proposed an elegant way to solve the bias issue by imposing physical symmetries under inverse mutations on the model structure, which we implemented in PoPMuSiC(sym). This new predictor constitutes an efficient trade-off between accuracy and absence of biases. Some final considerations and suggestions for further improvement of the predictors are discussed.