Rapid response for pandemics: single cell sequencing and deep learning to predict antibody sequences against an emerging antigen
Rapid response for pandemics: single cell sequencing and deep learning to predict antibody sequences against an emerging antigen
批准号:
10274223
负责人:
Jeniffer Bertha Hernandez
金额:
$185.16万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-16 至 2024-08-31
关键词:
AffinityAmino Acid SequenceAntibodiesAntibody FormationAntibody SpecificityAntibody TherapyAntigen-Antibody ComplexAntigensArchitectureB-Cell Antigen ReceptorB-Cell Receptor BindingB-LymphocytesBase SequenceBindingBiologicalBiologyCellsChronic DiseaseCodeComputer ModelsComputersComputing MethodologiesCoupledDataData SetDatabasesDegenerative DisorderDevelopmentDiagnosisEconomicsElectronsEngineeringEnzymesEpitopesEquilibriumFoundationsFutureGenesGenomicsGoalsHourImmune systemImmunizeImmunoassayImmunoglobulinsImmunologistImmunologyIndustrializationLigandsLightLinkMachine LearningMalignant NeoplasmsMeasurableMechanicsMethodsMicroscopicModelingMolecularMolecular BiologyMolecular ComputationsMusNatureNetwork-basedNeural Network SimulationOutputPassive ImmunotherapyPhage DisplayPhasePlayProblem SolvingProcessProductionProteinsReadinessReagentResearch Project GrantsSARS-CoV-2 antigenSARS-CoV-2 spike proteinSavingsScientistSeriesSpecificityStructural ChemistryStructural ModelsStructureSurface Plasmon ResonanceSystemTestingTherapeuticTherapeutic antibodiesThermodynamicsTimeTrainingVaccinesValidationVariantViralViral AntigensViral ProteinsWorkbasecombatcomputer sciencedata streamsdatabase structuredeep learningdeep neural networkdeep sequencingdensitydesignexperimental studyhigh dimensionalityin silicoinnovationinsightlarge datasetsmachine learning methodmolecular modelingmouse modelneutralizing antibodynovelnovel viruspandemic diseasepandemic preparednesspathogenphysical propertyprotein structurequantumresponsescaffoldsimulationsingle cell sequencingsynthetic antibodiestherapeutic evaluationthree dimensional structure
中文摘要
摘要
免疫学的“圣杯”之一是能够直接预测紧密结合的可变链抗体。
电子计算机中针对外源或非自身‘抗原性’蛋白的序列。免疫球蛋白链重排可以
可能编码大约1016种不同的抗体重链和轻链序列变体。然而,
通常只有一小部分序列空间被用于进化针对外源蛋白的抗体。
计算上的挑战是从抗原结构的模型到预测一组抗体
能与抗原紧密结合的链序列。如果得到解决,在不到24小时内搬家是可能的。
从首次提出的一种新型病毒蛋白的冷冻电子显微镜结构出发,提出了一套有效的抗体样蛋白
待测试的分子候选者。为了解决这一问题,本项目旨在开发一种深度学习
将热力学、量子力学(密度泛函)和局域结构作为输入的体系结构-
基于抗原及其同源抗体的网络拓扑特征,并将其输出
各自的结合亲和力常数。
我们将设计一个生成性对抗网络(GAN),我们认为它是唯一适合于基于回归的网络
免疫系统的ML方法,以发现表位和可变链之间的关联
功能。这种方法需要大量的抗原和同源抗体序列的数据流,直到
最近很难买到。一种新近报道的单一B细胞受体(BCR)特异性标记方法
单细胞深度测序(通过测序将B细胞受体与抗原特异性联系起来)或Libra-
SEQ)可以快速分离和测序能够高选择性结合的BCR可变链编码区
到抗原表位。
对于具体的项目目标,在任务1中,将使用Libra-seq来快速识别和生成候选人
免疫球蛋白编码序列响应特定的线性和非线性表位(对照),
通过计算/分子模拟选择,并优先考虑SARS-CoV-2尖峰蛋白表位(但
不限于这些),注入到老鼠模型中以生成大的训练集;在任务2中,这些训练
集合以及公共数据库中已有的其他数据集将生成一系列结构特征
(如上所述),这将用于训练GaN;在任务3中,预测的表位-抗体相互作用
将通过合成抗体和噬菌体展示系统的直接实验进行验证。因此,拟议的
战略结合了进化生物学、基因组学、结构化学和计算机的基本原理
科学来解决一个普遍的生物工程问题。
该项目的结果有望为经过严格测试的全自动机器奠定基础-
一种学习系统,可以从一种新病毒的结构中快速产生合成抗体候选
蛋白质,可以增强应对未来大流行的快速反应能力。有针对性地发展的能力
针对非传染性疾病或慢性病的抗体治疗,以及基于抗体的工业生产
如果这个项目取得成功,酶也将得到极大的增强。
团队:这个多机构研究项目的团队负责人包括一名计算机科学家,一名蛋白质
结晶学家、免疫学家和分子生物学家。
1
英文摘要
ABSTRACT
One of the “holy grails” in immunology is to be able to directly predict tight-binding variable chain antibody
sequences in silico against foreign or non-self `antigenic' proteins. Immunoglobulin chain rearrangement can
potentially encode approximately 1016 different variants of antibody heavy and light chain sequences. However,
only a small fraction of the sequence space is generally accessed for evolving antibodies against foreign proteins.
The computational challenge is to go from a model of the structure of an antigen to predicting a set of antibody
chain sequences that can bind tightly to the antigen. If solved, it might be possible to move in less than 24 hours
from the first cryo-electron-microscopic structure of a novel viral protein to advance a set of potent antibody-like
molecular candidates for testing. Towards solving this problem, this project aims to develop a deep learning
architecture that will take as input thermodynamic, quantum mechanical (density functional), and local structure-
based network topographical features of the antigens and their cognate antibodies, and will output their
respective binding affinity constants.
We will design a generative adversarial network (GAN), which we think is uniquely suited for regression-based
ML approaches for the immune system, to discover associations between the epitope and the variable chain
features. This approach requires a large data stream of antigen and cognate antibody sequences, which until
recently was difficult to obtain. A recently described single B-cell receptor (BCR) specific tagging method coupled
with single cell deep sequencing (“linking B cell receptor to antigen specificity through sequencing” or LIBRA-
seq) can rapidly isolate and sequence the BCR variable chain coding regions that can bind with high selectivity
to antigenic epitopes.
Towards the specific project goals, in Task 1, LIBRA-seq will be used to rapidly identify and generate candidate
immunoglobulin coding sequences in response to specific linear and nonlinear epitopes (against controls),
chosen through computational/molecular modeling and prioritized with SARS-CoV-2 Spike protein epitopes (but
not restricted to these), injected into a mouse model, to generate large training sets; in Task 2, these training
sets, along with other data sets already available in public databases, will generate a series of structural features
(described above), which will be used to train the GAN; in Task 3, the predicted epitope-antibody interactions
will be validated by direct experiments with synthetic antibody and phage-display systems. Thus, the proposed
strategy combines foundational principles in evolutionary biology, genomics, structural chemistry, and computer
science to the solution of a general biological engineering problem.
Results from this project are expected to lay the foundations for a rigorously tested and fully automated machine-
learning system that could rapidly generate synthetic antibody candidates from the structure of a novel virus
protein, which can enhance the rapid response ability against a future pandemic. The ability to develop targeted
antibody therapy against non-infectious or chronic diseases, and on the production of antibody-based industrial
enzymes, will also be dramatically enhanced if this project were to be successful.
The team: The team-leads of this multi-institutional research project comprise a computer scientist, a protein
crystallographer, an immunologist, and a molecular biologist.
1
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Rapid response for pandemics: single cell sequencing and deep learning to predict antibody sequences against an emerging antigen
-
批准号:10845715
-
项目类别:
-
资助金额:$121.99万
-
财政年份:2021
-
负责人:Jeniffer Bertha Hernandez
-
依托单位:
海外基金