Navigating Chemical Space with Natural Language Processing and Deep Learning
Navigating Chemical Space with Natural Language Processing and Deep Learning
批准号:
EP/Y004167/1
负责人:
Jiayun Pang
金额:
$11.41万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --
中文摘要
自然语言处理(NLP)是语言学和计算机科学的交汇点,其目的是处理和分析人类语言,通常以书面文本的形式提供。NLP现在非常关注使用机器学习来完成具有挑战性的任务,在过去的几年里已经开发了一些革命性的算法。它们现在支撑了一系列现实生活中的应用,如ChatGPT、虚拟助手和我们写电子邮件时的自动文本补全。创新的研究思路往往来自于跨学科的技术和概念的整合。对于这一学科跳跃拨款,我们希望探索Transformer Models如何才能适应化学领域的研究挑战。Transformer Models是谷歌在2017年开发的一种开创性的深度学习算法,为目前NLP中的大部分尖端研究提供了动力。化学结构通常是三维的。然而,它们也经常被转换成序列,称为微笑。微笑有一个简单的化学元素和键符号的词汇表,以及一些关于化学元素位置的语法规则。由于这种对文本序列的直接类比,通过微笑,可以使用NLP算法以类似于分析文本的方式来分析化学结构。在拟议的研究中,化学家彭博士将与自然语言处理和机器学习专家Vulic博士合作,以加快了解自然语言处理领域的最新发展,并研究它们在她的专业领域中的进一步适用性。我们将探索和利用一个现在机器学习和NLP中普遍存在的概念,称为转移学习,它1)预训练大型通用模型,2)针对特定任务和应用微调(即专门化)这些通用模型,其中标记的数据创建成本很高(因为它们需要专业知识和复杂的注释协议),因此固有的稀缺性。具体地说,我们将预先训练变形金刚模型,以学习由数千万微笑定义的化学空间的潜在表示。然后,在微调过程中,这种学习的潜在表示可以用于预测给定化学结构的分子性质。这种方法的优点是,得到的机器学习模型对所谓的标记数据(具有实验确定的性质的分子)的依赖较少,考虑到相关的成本和实验挑战,在化学中生成这些数据是耗时的,甚至是不可能的。我们的目标是使用两种最新的机器学习技术,即句子编码和对比学习,使Transformer模型在计算上更加高效和准确。我们希望这种新的分子表示方法可以补充现有的分子表示方法,并提供一种替代方法来评估分子结构与其性质,这是化学和制药行业许多研究和开发任务的基础。
英文摘要
Natural language processing (NLP) lies at the intersection between linguistics and computer science which aims to process and analyse human language, typically provided as written text. NLP is now strongly focused on the use of machine learning for challenging tasks with some revolutionary algorithms having been developed in the last few years. They now underpin a wide range of real-life applications, such as ChatGPT, virtual assistants and automatic text completion when we write emails. Innovative research ideas often come from integrating techniques and concepts across disciplines. For this discipline-hopping grant, we would like to explore how Transformer models, a ground-breaking deep learning algorithm developed by Google in 2017 which fuels majority of the current cutting-edge research in NLP, can be adapted to solve research challenges in chemistry. Chemical structures are usually three dimensional. However, they are also often converted into sequences, called SMILES. SMILES has a simple vocabulary of chemical elements and bond symbols and a few grammatical rules of how the chemical elements are positioned. Owing to this direct analogy to text sequences, through SMILES it is possible to use NLP algorithms to analyse chemical structures in a similar fashion as they are used to analyse text. For the proposed research, Dr Pang, a chemist will work with Dr Vulic, an NLP and machine learning expert in order to get up to speed with the latest developments in the field of NLP and to examine their further applicability in her domain of expertise. We will explore and utilise a concept which is now pervasive in machine learning and NLP, termed transfer learning, which 1) pretrains large general-purpose models, and 2) fine-tunes (i.e., specialises) those general models for specific tasks and applications, where labelled data are expensive to create (as they require expert knowledge and complex annotation protocols) and thus inherently scarce. Specifically, we will pretrain Transformer models to learn a latent representation of the chemical space defined by tens of millions of SMILES. This learned latent representation can then be used to predict molecular properties for a given chemical structure during fine-tuning. The advantage of this type of approach is that the resulting machine learning models rely less on the so-called labelled data (molecules with experimentally determined properties), which are time-consuming or even impossible to generate in chemistry considering the associated cost and experimental challenges. We will aim to make the Transformer models more computationally efficient and accurate using two latest machine learning techniques, termed sentence encoding and contrastive learning. We hope that this new molecular representation can complement existing molecular representation methods and provide an alternative approach to evaluate molecular structures against their properties, which underpins many research and development tasks in the chemical and pharmaceutical industries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Chinese Journal of Chemical Engineering
-
批准号:21224004
-
项目类别:专项基金项目
-
资助金额:20.0万元
-
批准年份:2012
-
负责人:廖叶华
-
依托单位:
Chinese Journal of Chemical Engineering
-
批准号:21024805
-
项目类别:专项基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:廖叶华
-
依托单位: