Investigating B-cell repertoire data using deep learning approaches to aid in the development of antibody therapeutics
Investigating B-cell repertoire data using deep learning approaches to aid in the development of antibody therapeutics
批准号:
2271214
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
抗体是免疫系统的重要蛋白质。它们识别潜在的有害分子,与它们结合,并开始将它们从体内清除。到目前为止,大约有100种抗体被批准用于临床,它们已经成为一种重要的且不断增长的药物类别。然而,治疗性抗体的开发因许多要求而变得复杂,包括化学稳定性、溶解性、低粘度、生物利用度、长的血清半衰期、无免疫原性、抗碎裂、聚集、翻译后修饰和蛋白水解性切割,同时还保持其所需的功能(结合亲和力、特异性和功能活性)。尽管治疗性抗体的发现方法不断发展,但它仍然是一个昂贵和繁琐的过程。因此,计算机方法以及机器学习算法引入的见解,可以从以前的实验中提取和利用信息来预测新抗体的性质和功能,因此受到高度关注。最近,UniRep等方法显示了通过在蛋白质数据上应用自然语言处理(NLP)启发的方法,特别是迁移学习来改进蛋白质预测的可能性。转移学习是指将来自一个领域的信息用于另一个相关领域,当后一个领域只有很小的数据集可用时,这可能特别强大,抗体数据通常会出现这种情况。因此,开发基于最先进的NLP技术的新的抗体ML工具可以对治疗性抗体的发现产生巨大的影响。本项目的目的是探索抗体数据并开发新的机器学习工具,以提高抗体性质和功能的预测。这包括:调查观察到的抗体空间数据库中可获得的大量序列数据。开发新的ML技术,并采用新的NLP技术来处理生物数据。探索这些新技术在抗体性质和功能预测中的应用。这个DPhil项目是牛津大学牛津蛋白质信息学小组(OPIG)的夏洛特·迪恩教授和伦敦葛兰素史克(GSK)的Iain H.Moal博士合作的。该项目与EPSRC的几个战略和研究领域相一致。它主要属于EPSRC生物信息学研究领域,因为它开发了新的计算技术来模拟和分析生物数据(用于抗体预测的机器学习工具)。此外,该项目还属于EPSRC分析科学和人工智能技术研究领域,我们使用新颖的ML技术从大数据集中提取信息,用于分析和预测抗体的特性。
英文摘要
Antibodies are important proteins of the immune system. They recognize potentially harmful molecules, binding to them and initiating their removal from the body. With approximately 100 antibodies approved for clinical use to date, they have become an important and growing class of pharmaceuticals. However, therapeutic antibody development is complicated by numerous requirements, including chemical stability, solubility, low viscosity, bioavailability, long serum half-life, non-immunogenicity, and resistance to fragmentation, aggregation, post-translational modification and proteolytic cleavage, while also retaining their desired functions (binding affinity, specificity and functional activity).Whilst methodologies for therapeutic antibody discovery are constantly evolving, it remains an expensive and cumbersome process. Insights introduced by in silico approaches, along with machine learning algorithms, which can extract and utilize information from previous experiments to predict the properties and functions of new antibodies, are therefore highly sought. Recently, methods such as UniRep have shown the possibilities of improving protein predictions by applying Natural Language Processing (NLP) inspired methods, notably transfer learning, on protein data. Transfer learning is when information from one domain is used in another related domain, which can be particularly powerful when only small data sets are available for the latter domain, a common occurrence with antibody data. Developing new ML tools for antibodies built on state-of-the-art NLP techniques can therefore have a large impact on the therapeutic antibody discovery.The aim of this project is to explore antibody data and develop novel machine learning tools for improving the predictions of antibody properties and functions. This includes; Investigating the large amount of sequence data available in the Observed Antibody Space database. Develop novel ML techniques, and adaptation of novel NLP techniques to work on biological data. Explore the use of these new techniques in antibody property and function prediction.This DPhil project is a collaboration between Prof. Charlotte Deane at the Oxford Protein Informatics Group (OPIG), University of Oxford and Dr. Iain H. Moal at GlaxoSmithKline (GSK), London. This project aligns with several of EPSRC's strategies and research areas. It mainly falls within the EPSRC Biological Informatics research area for its development of novel computational techniques to model and analyse biological data (machine learning tools for antibody predictions). Additionally, the project also falls within the EPSRC Analytical Science and Artificial Intelligence Technologies research areas, for our use of novel ML techniques to extract information from large datasets for analyzing and predicting properties of antibodies.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1002/pro.4205
发表时间:
2022-01
期刊:
Protein science : a publication of the Protein Society
影响因子:
--
作者:
[Olsen TH, Boyles F, Deane CM]
通讯作者:
Deane CM
AbLang: An antibody language model for completing antibody sequences
AbLang:用于完成抗体序列的抗体语言模型
DOI:
10.1101/2022.01.20.477061
发表时间:
2022
期刊:
影响因子:
--
作者:
[Olsen T]
通讯作者:
Olsen T
国内基金
海外基金
登录
查看更多内容
全细胞疫苗Cell@MnO2的乳腺癌术后免疫响应监测与放射免疫治疗研究
-
批准号:QN25H220002
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:顾媛
-
依托单位:
染色体外环状DNA以cell-in-cell途径促进基因横向传递和扩增的研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:15.0万元
-
批准年份:2024
-
负责人:王锐智
-
依托单位:
GMFG/F-actin/cell adhesion 轴驱动 EHT 在造
血干细胞生成中的作用及机制研究
-
批准号:TGY24H080011
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:李鸿鹄
-
依托单位:
基于In-cell NMR策略对“舟楫之剂”桔梗中引经药效物质的快速发现研究
-
批准号:82305053
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:王丽明
-
依托单位:
配子生成素GGN不同位点突变损伤分子伴侣BIP及HSP90B1功能导致精子形成障碍的发病机理
-
批准号:82371616
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:姚晨成
-
依托单位:
糖尿病ED中成纤维细胞衰老调控内皮细胞线粒体稳态失衡的机制研究
-
批准号:82371634
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:赵福军
-
依托单位:
骨髓ISG+NAMPT+中性粒细胞介导抗磷脂综合征B细胞异常活化的机制研究
-
批准号:82371799
-
项目类别:面上项目
-
资助金额:47.00万元
-
批准年份:2023
-
负责人:杨程德
-
依托单位:
利用CRISPR内源性激活Atoh1转录促进前庭毛细胞再生和功能重建
-
批准号:82371145
-
项目类别:面上项目
-
资助金额:46.00万元
-
批准年份:2023
-
负责人:陶永
-
依托单位:
IL-4协同精氨酸优化种植初期巨噬细胞胞葬作用和成骨微环境的作用及机制研究
-
批准号:82370923
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:张文杰
-
依托单位:
胆固醇合成蛋白CYP51介导线粒体通透性转换诱发Th17/Treg细胞稳态失衡在舍格伦综合征中的作用机制研究
-
批准号:82370976
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:郑凌艳
-
依托单位: