IgLM: Infilling language modeling for antibody sequence design.

IgLM: Infilling language modeling for antibody sequence design.
复制标题

IgLM:抗体序列设计的填充语言模型。

DOI:
10.1016/j.cels.2023.10.001
复制
发表时间:
2023
期刊:
影响因子:
9.3
通讯作者:
Gray,JeffreyJ
Gray,JeffreyJ
中科院分区:
生物学1区
文献类型:
--
作者:
Shuai,RichardW;Ruffolo,JeffreyA;Gray,JeffreyJ

文献摘要

被引文献

相似文献

用于治疗应用的单抗的发现和优化依赖于大的序列文库,但受到诸如低溶解度、高聚集性和高免疫原性等可开发性问题的阻碍。生成性语言模型对数百万个蛋白质序列进行了训练,是按需生成现实的、多样化的序列的强大工具。我们提出了免疫球蛋白语言模型(IgLM),这是一个用于创建合成抗体库的深层生成性语言模型。与以前利用单向上下文生成序列的方法不同,IgLM基于自然语言中的文本填充来制定抗体设计,允许它使用双向上下文重新设计抗体序列中的可变长度跨度。我们对5.58亿(M)抗体重链和轻链可变序列进行了IgLM训练,条件是每个序列的链类型和起源种类。我们证明了IgLM可以从各种物种中生成全长抗体序列,其填充配方允许它生成填充的互补决定区(CDR)环库,并改善了电子发育特性。补充资料中包括了本文件透明的同行审查过程的记录。
Discovery and optimization of monoclonal antibodies for therapeutic applications relies on large sequence libraries but is hindered by developability issues such as low solubility, high aggregation, and high immunogenicity. Generative language models, trained on millions of protein sequences, are a powerful tool for the on-demand generation of realistic, diverse sequences. We present the Immunoglobulin Language Model (IgLM), a deep generative language model for creating synthetic antibody libraries. Compared with prior methods that leverage unidirectional context for sequence generation, IgLM formulates antibody design based on text-infilling in natural language, allowing it to re-design variable-length spans within antibody sequences using bidirectional context. We trained IgLM on 558 million (M) antibody heavy- and light-chain variable sequences, conditioning on each sequence's chain type and species of origin. We demonstrate that IgLM can generate full-length antibody sequences from a variety of species and its infilling formulation allows it to generate infilled complementarity-determining region (CDR) loop libraries with improvedin silicodevelopability profiles. A record of this paper's transparent peer review process is included in the supplemental information.