Quantitative Missense Variant Effect Prediction Using Large-Scale Mutagenesis Data

Quantitative Missense Variant Effect Prediction Using Large-Scale Mutagenesis Data
复制标题

DOI:
10.1016/j.cels.2017.11.003
复制
发表时间:
2018-01-24
期刊:
影响因子:
9.3
通讯作者:
Fowler, Douglas M.
Fowler, Douglas M.
中科院分区:
生物学1区
文献类型:
--
作者:
Gray, Vanessa E.;Hause, Ronald J.;Fowler, Douglas M.

文献摘要

被引文献

相似文献

描述突变对蛋白质功能的定量影响的大型数据集变得越来越可用。在这里,我们利用这些数据集来开发Envision,它预测错义变体的分子效应的大小。Envision结合了来自9个大规模实验诱变数据集的21,026个变异效应测量值,这是迄今为止尚未开发的训练资源,具有监督的随机梯度提升学习算法。Envision在大规模诱变数据和包含2,312个TP 53变体的独立测试数据集上均优于其他错义变体效应预测因子,这些变体的效应使用低通量方法进行测量。该数据集从未用于超参数调整或模型训练,因此作为独立的验证集。与其他预测因子相比,Envision预测准确性在氨基酸之间也更一致。最后,我们证明了Envision的性能提高了更多的大规模诱变数据。我们预先计算了人类、小鼠、青蛙、斑马鱼、果蝇、蠕虫和酵母蛋白质组中每一个可能的单个氨基酸变体的Envision预测(https://envision.gs.washington.edu/)。
Large datasets describing the quantitative effects of mutations on protein function are becoming increasingly available. Here, we leverage these datasets to develop Envision, which predicts the magnitude of a missense variant's molecular effect. Envision combines 21,026 variant effect measurements from nine large-scale experimental mutagenesis datasets, a hitherto untapped training resource, with a supervised, stochastic gradient boosting learning algorithm. Envision outperforms other missense variant effect predictors both on large-scale mutagenesis data and on an independent test dataset comprising 2,312 TP53 variants whose effects were measured using a low-throughput approach. This dataset was never used for hyperparameter tuning or model training and thus serves as an independent validation set. Envision prediction accuracy is also more consistent across amino acids than other predictors. Finally, we demonstrate that Envision's performance improves as more large-scale mutagenesis data are incorporated. We precompute Envision predictions for every possible single amino acid variant in human, mouse, frog, zebrafish, fruit fly, worm, and yeast proteomes (https://envision.gs.washington.edu/).