In silico proof of principle of machine learning-based antibody design at unconstrained scale.

In silico proof of principle of machine learning-based antibody design at unconstrained scale.
复制标题

DOI:
10.1080/19420862.2022.2031482
复制
发表时间:
2022-01
期刊:
影响因子:
5.3
通讯作者:
Greiff V
Greiff V
中科院分区:
医学2区
文献类型:
--
作者:
Akbar R;Robert PA;Weber CR;Widrich M;Frank R;Pavlović M;Scheffer L;Chernigovskaya M;Snapkov I;Slabodkin A;Mehta BB;Miho E;Lund-Johansen F;Andersen JT;Hochreiter S;Hobæk Haff I;Klambauer G;Sandve GK;Greiff V

文献摘要

参考文献

被引文献

相似文献

生成式机器学习(ML)已被假定成为抗原特异性单克隆抗体(mAb)计算设计的主要驱动力。然而,确认这一假设的努力受到了阻碍,因为无法测试任意大量的抗体序列的最关键的设计参数:互补位、表位、亲和力和可开发性。为了应对这一挑战,我们利用了基于晶格的抗体-抗原结合模拟框架,该框架结合了广泛的生理抗体结合参数。模拟框架使得能够计算合成的抗体-抗原3D结构,并且它用作ML生成的抗体序列的抗体设计参数的不受限制的前瞻性评估和基准测试的预言机。我们发现,专门在抗体序列(一维:1D)数据上训练的深度生成模型可用于设计构象(三维:3D)表位特异性抗体,在亲和力和可开发性参数值变化方面匹配或超过训练数据集。此外,我们建立了高准确性生成抗体ML所需的序列多样性的较低阈值,并证明该较低阈值也适用于实验真实世界数据。最后,我们证明了迁移学习可以从低N训练数据中生成高亲和力的抗体序列。我们的工作建立了先验的可行性和高通量ML为基础的单克隆抗体设计的理论基础。
Generative machine learning (ML) has been postulated to become a major driver in the computational design of antigen-specific monoclonal antibodies (mAb). However, efforts to confirm this hypothesis have been hindered by the infeasibility of testing arbitrarily large numbers of antibody sequences for their most critical design parameters: paratope, epitope, affinity, and developability. To address this challenge, we leveraged a lattice-based antibody-antigen binding simulation framework, which incorporates a wide range of physiological antibody-binding parameters. The simulation framework enables the computation of synthetic antibody-antigen 3D-structures, and it functions as an oracle for unrestricted prospective evaluation and benchmarking of antibody design parameters of ML-generated antibody sequences. We found that a deep generative model, trained exclusively on antibody sequence (one dimensional: 1D) data can be used to design conformational (three dimensional: 3D) epitope-specific antibodies, matching, or exceeding the training dataset in affinity and developability parameter value variety. Furthermore, we established a lower threshold of sequence diversity necessary for high-accuracy generative antibody ML and demonstrated that this lower threshold also holds on experimental real-world data. Finally, we show that transfer learning enables the generation of high-affinity antibody sequences from low-N training data. Our work establishes a priori feasibility and the theoretical foundation of high-throughput ML-based mAb design.
Biopython:用于计算分子生物学和生物信息学的免费 Python 工具。
DOI: 10.1093/bioinformatics/btp163
发表时间: 2009-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cock PJ;Antao T;Chang JT;Chapman BA;Cox CJ;Dalke A;Friedberg I;Hamelryck T;Kauff F;Wilczynski B;de Hoon MJ
通讯作者: de Hoon MJ
DOI: 10.1016/j.celrep.2017.04.054
发表时间: 2017-05-16
期刊: CELL REPORTS
影响因子: 8.8
作者:
Greiff, Victor;Menzel, Ulrike;Reddy, Sai T.
通讯作者: Reddy, Sai T.
DOI: 10.7554/elife.46935
发表时间: 2019-09-05
期刊: ELIFE
影响因子: 7.7
作者:
Davidsen, Kristian;Olson, Branden J.;Matsen, Frederick A.
通讯作者: Matsen, Frederick A.
DOI: 10.1080/19420862.2020.1743053
发表时间: 2020-01-01
期刊: MABS
影响因子: 5.3
作者:
Bailly, Marc;Mieczkowski, Carl;Fayadat-Dilman, Laurence
通讯作者: Fayadat-Dilman, Laurence
DOI: 10.1038/s41592-021-01283-4
发表时间: 2021-10
期刊: Nature methods
影响因子: 48
作者:
AlQuraishi M;Sorger PK
通讯作者: Sorger PK