Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval

Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval
复制标题

Tranception:利用自回归变压器和推理时间检索进行蛋白质适应性预测

DOI:
--
复制
发表时间:
2022
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Y. Gal
Y. Gal
中科院分区:
--
文献类型:
--
作者:
Pascal Notin;M. Dias;J. Frazer;Javier Marchena;Aidan N. Gomez;D. Marks;Y. Gal

文献摘要

参考文献

被引文献

相似文献

准确模拟蛋白质序列的适应性景观的能力对于广泛的应用至关重要,从量化人类变异对疾病可能性的影响,到预测病毒的免疫逃逸突变和设计新型生物治疗蛋白质。在多个序列比对上训练的蛋白质序列深度生成模型是迄今为止解决这些任务的最成功的方法。然而,这些方法的性能取决于是否有足够深度和多样化的比对来进行可靠的训练。因此,它们的潜在范围受到以下事实的限制:许多蛋白质家族即使不是不可能,也很难对齐。对来自不同家族的大量非比对蛋白质序列进行训练的大型语言模型可以解决这些问题,并显示出最终弥合性能差距的潜力。我们引入了 Tranception,一种新颖的 Transformer 架构,利用自回归预测和推理时同源序列的检索来实现最先进的适应度预测性能。鉴于其在多个突变体上显着更高的性能、对浅层比对的鲁棒性以及对插入缺失进行评分的能力,我们的方法比现有方法提供了显着的范围增益。为了在更广泛的蛋白质家族中进行更严格的模型测试,我们开发了 ProteinGym——一套广泛的变异效应多重检测方法,与现有基准相比,大大增加了检测的数量和多样性。
The ability to accurately model the fitness landscape of protein sequences is critical to a wide range of applications, from quantifying the effects of human variants on disease likelihood, to predicting immune-escape mutations in viruses and designing novel biotherapeutic proteins. Deep generative models of protein sequences trained on multiple sequence alignments have been the most successful approaches so far to address these tasks. The performance of these methods is however contingent on the availability of sufficiently deep and diverse alignments for reliable training. Their potential scope is thus limited by the fact many protein families are hard, if not impossible, to align. Large language models trained on massive quantities of non-aligned protein sequences from diverse families address these problems and show potential to eventually bridge the performance gap. We introduce Tranception, a novel transformer architecture leveraging autoregressive predictions and retrieval of homologous sequences at inference to achieve state-of-the-art fitness prediction performance. Given its markedly higher performance on multiple mutants, robustness to shallow alignments and ability to score indels, our approach offers significant gain of scope over existing approaches. To enable more rigorous model testing across a broader range of protein families, we develop ProteinGym -- an extensive set of multiplexed assays of variant effects, substantially increasing both the number and diversity of assays compared to existing benchmarks.
DOI: 10.1016/j.hrthm.2020.05.041
发表时间: 2020-12
期刊: Heart rhythm
影响因子: 5.5
作者:
Kozek KA;Glazer AM;Ng CA;Blackwell D;Egly CL;Vanags LR;Blair M;Mitchell D;Matreyek KA;Fowler DM;Knollmann BC;Vandenberg JI;Roden DM;Kroncke BM
通讯作者: Kroncke BM
DOI: 10.1016/j.celrep.2016.09.061
发表时间: 2016-10-18
期刊: Cell reports
影响因子: 8.8
作者:
Brenan L;Andreev A;Cohen O;Pantel S;Kamburov A;Cacchiarelli D;Persky NS;Zhu C;Bagul M;Goetz EM;Burgin AB;Garraway LA;Getz G;Mikkelsen TS;Piccioni F;Root DE;Johannessen CM
通讯作者: Johannessen CM
DOI: 10.1016/j.cels.2018.01.015
发表时间: 2018-04-25
期刊: Cell systems
影响因子: 9.3
作者:
Staller MV;Holehouse AS;Swain-Lenz D;Das RK;Pappu RV;Cohen BA
通讯作者: Cohen BA
DOI: 10.1016/j.jmb.2013.01.032
发表时间: 2013-04-26
影响因子: 5.6
作者:
Roscoe, Benjamin P.;Thayer, Kelly M.;Zeldovich, Konstantin B.;Fushman, David;Bolon, Daniel N. A.
通讯作者: Bolon, Daniel N. A.
DOI: 10.1016/j.cels.2016.11.004
发表时间: 2016-12-21
期刊: Cell systems
影响因子: 9.3
作者:
Kelsic ED;Chung H;Cohen N;Park J;Wang HH;Kishony R
通讯作者: Kishony R