Updated benchmarking of variant effect predictors using deep mutational scanning.

Updated benchmarking of variant effect predictors using deep mutational scanning.
复制标题

DOI:
10.15252/msb.202211474
复制
发表时间:
2023-08-08
影响因子:
9.9
通讯作者:
Marsh, Joseph A.
Marsh, Joseph A.
中科院分区:
生物学1区
文献类型:
--
作者:
Livesey, Benjamin J.;Marsh, Joseph A.

文献摘要

参考文献

被引文献

相似文献

变异效应预测因子(VEP)性能的评估充满了通过对临床观察进行基准测试而引入的偏差。在这项研究中,在我们以前的工作的基础上,我们使用从26种人类蛋白质的深度突变扫描(DMS)实验中独立生成的蛋白质功能测量值来基准55种不同的VEP,同时引入最小的数据循环。许多表现最好的VEP是无监督方法,包括EVE,DeepSequence和ESM‐1v,这是一种蛋白质语言模型,总体排名第一。然而,最近监督VEP的强劲表现,特别是VARITY,表明开发人员正在认真对待数据循环和偏差问题。我们还评估了DMS和无监督VEP用于区分已知致病性和无害性错义变体的性能。我们的研究结果是混合的,表明一些DMS数据集在变体分类方面表现异常,而其他数据集则很差。值得注意的是,我们观察到VEP与DMS数据的一致性与识别临床相关变体的性能之间存在显著相关性,强烈支持我们的排名的有效性和DMS用于独立基准测试的实用性。使用来自深度突变扫描实验的数据评估变异效应预测基准中的常见偏倚来源。ESM‐1v、EVE和DeepSequence在功能验证和临床观察到的变体方面都表现最佳。
The assessment of variant effect predictor (VEP) performance is fraught with biases introduced by benchmarking against clinical observations. In this study, building on our previous work, we use independently generated measurements of protein function from deep mutational scanning (DMS) experiments for 26 human proteins to benchmark 55 different VEPs, while introducing minimal data circularity. Many top‐performing VEPs are unsupervised methods including EVE, DeepSequence and ESM‐1v, a protein language model that ranked first overall. However, the strong performance of recent supervised VEPs, in particular VARITY, shows that developers are taking data circularity and bias issues seriously. We also assess the performance of DMS and unsupervised VEPs for discriminating between known pathogenic and putatively benign missense variants. Our findings are mixed, demonstrating that some DMS datasets perform exceptionally at variant classification, while others are poor. Notably, we observe a striking correlation between VEP agreement with DMS data and performance in identifying clinically relevant variants, strongly supporting the validity of our rankings and the utility of DMS for independent benchmarking. Common sources of bias in variant effect predictor benchmarking are assessed using data from deep mutational scanning experiments. ESM‐1v, EVE and DeepSequence are among the top performers on both functionally validated and clinically observed variants.
DOI: 10.1038/nmeth.3027
发表时间: 2014-08
期刊: NATURE METHODS
影响因子: 48
作者:
Fowler, Douglas M.;Fields, Stanley
通讯作者: Fields, Stanley
DOI: 10.4049/jimmunol.1800343
发表时间: 2018-06-01
期刊: Journal of immunology (Baltimore, Md. : 1950)
影响因子: --
作者:
Heredia JD;Park J;Brubaker RJ;Szymanski SK;Gill KS;Procko E
通讯作者: Procko E
DOI: 10.1016/j.cels.2017.11.003
发表时间: 2018-01-24
期刊: CELL SYSTEMS
影响因子: 9.3
作者:
Gray, Vanessa E.;Hause, Ronald J.;Fowler, Douglas M.
通讯作者: Fowler, Douglas M.
DOI: 10.1074/jbc.c500288200
发表时间: 2005-09-02
影响因子: 4.8
作者:
Bertoncini, CW;Fernandez, CO;Zweckstetter, M
通讯作者: Zweckstetter, M
DOI: 10.1016/j.ajhg.2021.07.001
发表时间: 2021-09-02
影响因子: 9.8
作者:
Amorosi, Clara J.;Chiasson, Melissa A.;Dunham, Maitreya J.
通讯作者: Dunham, Maitreya J.