Genome-wide prediction of disease variant effects with a deep protein language model.

Genome-wide prediction of disease variant effects with a deep protein language model.
复制标题

DOI:
10.1038/s41588-023-01465-0
复制
发表时间:
2023-09
期刊:
影响因子:
30.8
通讯作者:
Ntranos, Vasilis
Ntranos, Vasilis
中科院分区:
生物学1区
文献类型:
--
作者:
Brandes, Nadav;Goldman, Grant;Wang, Charlotte H. H.;Ye, Chun Jimmie;Ntranos, Vasilis

文献摘要

参考文献

被引文献

相似文献

预测编码变体的影响是一个重大挑战。虽然最近的深度学习模型提高了变体效应预测的准确性,但由于依赖于相近的同源物或软件限制,它们无法分析所有编码变体。在这里,我们开发了一个工作流程,使用6.5亿个参数的蛋白质语言模型ESM 1b来预测人类基因组中所有~ 4.5亿个可能的错义变体效应,并在门户网站上提供所有预测。ESM 1b在将约150,000个ClinVar/HGMD错义变异分类为致病性或良性以及预测28个深度突变扫描数据集的测量结果方面优于现有方法。我们进一步注释了约200万个变体仅在特定蛋白质亚型中具有破坏性,证明了在预测变体效应时考虑所有亚型的重要性。我们的方法还推广到更复杂的编码变体,如帧内插入缺失和停止增益。这些结果共同建立了蛋白质语言模型,作为预测变体效应的有效、准确和通用方法。利用蛋白质语言模型(ESM 1b)的修改后的框架被用于预测人类基因组中所有可能的4.5亿个错义变异效应,并显示出推广到更复杂的遗传变异(如indels和stop-gains)的潜力。
Predicting the effects of coding variants is a major challenge. While recent deep-learning models have improved variant effect prediction accuracy, they cannot analyze all coding variants due to dependency on close homologs or software limitations. Here we developed a workflow using ESM1b, a 650-million-parameter protein language model, to predict all ~450 million possible missense variant effects in the human genome, and made all predictions available on a web portal. ESM1b outperformed existing methods in classifying ~150,000 ClinVar/HGMD missense variants as pathogenic or benign and predicting measurements across 28 deep mutational scan datasets. We further annotated ~2 million variants as damaging only in specific protein isoforms, demonstrating the importance of considering all isoforms when predicting variant effects. Our approach also generalizes to more complex coding variants such as in-frame indels and stop-gains. Together, these results establish protein language models as an effective, accurate and general approach to predicting variant effects. A modified framework leveraging a protein language model (ESM1b) is used to predict all possible 450 million missense variant effects in the human genome and shows potential for generalizing to more complex genetic variations such as indels and stop-gains.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1038/ng.3404
发表时间: 2015-11
期刊: Nature genetics
影响因子: 30.8
作者:
Finucane HK;Bulik-Sullivan B;Gusev A;Trynka G;Reshef Y;Loh PR;Anttila V;Xu H;Zang C;Farh K;Ripke S;Day FR;ReproGen Consortium;Schizophrenia Working Group of the Psychiatric Genomics Consortium;RACI Consortium;Purcell S;Stahl E;Lindstrom S;Perry JR;Okada Y;Raychaudhuri S;Daly MJ;Patterson N;Neale BM;Price AL
通讯作者: Price AL
Biopython:用于计算分子生物学和生物信息学的免费 Python 工具。
DOI: 10.1093/bioinformatics/btp163
发表时间: 2009-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cock PJ;Antao T;Chang JT;Chapman BA;Cox CJ;Dalke A;Friedberg I;Hamelryck T;Kauff F;Wilczynski B;de Hoon MJ
通讯作者: de Hoon MJ
DOI: 10.1093/nar/gky1120
发表时间: 2019-01-08
影响因子: 14.9
作者:
Buniello, Annalisa;MacArthur, Jacqueline A. L.;Parkinson, Helen
通讯作者: Parkinson, Helen
DOI: 10.1038/s41586-020-2329-2
发表时间: 2020-05-28
期刊: NATURE
影响因子: 64.8
作者:
Cummings, Beryl B.;Karczewski, Konrad J.;MacArthur, Daniel G.
通讯作者: MacArthur, Daniel G.