课题基金 / 基金详情

Machine Learning Approaches to Predict Enzyme Function

Machine Learning Approaches to Predict Enzyme Function
预测酶功能的机器学习方法
批准号:
BB/I00596X/1
负责人:
John Mitchell
金额:
$33.78万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --

项目摘要

项目成果

John Mitchell的其他基金

相似基金

相关文献

中文摘要
翻译
蛋白质是生物系统中最重要的分子之一。它们对生物体至关重要,生物体利用它们来执行各种各样的基本功能:催化、运输、储存、运动功能、信号传导、陪伴折叠、调节、分子识别、结构作用和DNA修复。由于蛋白质在生物学中无处不在,如果我们想了解生物过程,了解它们的特性是必不可少的。这个项目的重点是所有蛋白质功能中最重要的一个:酶催化。酶催化或促进生物体内发生的化学反应。了解它们的工作原理本身就很有趣,而且在药物设计、诊断、生物燃料、食品科学和洗衣等多种领域都很有用。这个项目是关于蛋白质的结构和它所执行的酶功能之间的关系。我们的目标是从蛋白质结构的知识来预测催化功能。为了实现这一点,我们将使用机器学习方法,特别是一种称为随机森林的技术。森林由几百棵“决策树”组成,每棵树基本上都是一个流程图。我们将训练他们学习现有酶结构的已知性质和它们催化的反应的化学步骤的模式。然而,我们生成树的方法涉及到计算机模拟的掷骰子。这将确保它们都是不同的,尽管它们基于相同的底层信息。然后每个决策树对未知的可能的催化功能进行预测。这些预测被视为对蛋白质功能的投票。这个投票过程产生了许多决策树的共识,并最大限度地利用了底层数据中包含的信息,产生的结果比任何一个决策树的结果都要准确得多。由于许多原因,酶功能的预测是非常重要的。首先,能够更准确地预测酶的功能将提高基因组的功能注释,降低目前通过生物信息学数据库传播错误注释的风险。结构基因组学的快速发展,对来自各种生物体的各种蛋白质的高通量结构测定,意味着许多功能尚不清楚的酶的结构是可用的。其次,这个项目将使我们认识到进化上不相关的酶之间的化学相似性,这些酶催化相似的步骤,尽管不一定是相似的总体反应。第三,这项工作将帮助我们理解蛋白质结构、功能和进化之间复杂关系的关键决定因素,特别是在反应步骤的催化方面。第四,该项目将有助于设计具有新功能或对现有功能进行精心修改的新酶。这个项目位于学科之间的界面,结合了化学、生物学和计算机科学。广泛的技能和专业知识是增加我们对催化的理解所必需的,这一直是一个重要的学术目标。在商业上,这项工作为制药和生物技术行业奠定了直接有用的基础,其中酶被用作诊断和治疗;农化工业,其产品经常以酶为目标;在生物燃料的开发中,需要强大的酶来提高生产力和降低成本;在洗衣业,酶已经被用于日常用品;在营养和食品行业。特别是这个项目将有助于设计新的和重新利用的酶。
英文摘要
Proteins are amongst the most important of all molecules in biological systems. They are crucial to organisms which use them to carry out a huge variety of essential functions: catalysis, transport, storage, motor functions, signalling, chaperoning folding, regulation, molecular recognition, structural roles, and DNA Repair. As proteins are so ubiquitous in biology, understanding their properties is essential if we want to know about biological processes. This project is focused on one of the most significant of all protein functions: enzyme catalysis. Enzymes catalyse, or facilitate, the chemical reactions that occur in living organisms. Understanding how they work is both interesting in itself and useful in areas as diverse as drug design, diagnostics, biofuels, food science and laundry. This project is about the relationship between the structure of a protein and the enzyme function it carries out. We aim to predict the catalytic functionality from a knowledge of the protein structure. In order to achieve this, we will use machine learning methods, and in particular a technique called Random Forest. The forest consists of several hundred 'decision trees', each of which is basically a flow diagram. We will train them to learn patterns in the known properties of existing enzyme structures and the chemistry of the steps comprising the reactions they catalyse. However, the way in which we will generate the trees involves computer-simulated dice-rolling. This will ensure that they are all different, though based on the same underlying information. The decision trees then each make a prediction of the unknown possible catalytic functions. These predictions are treated as votes as to the function of the protein. This voting process produces a consensus of many decision trees and maximises the use of the information contained in the underlying data, generating results which are much more accurate than those of any one decision tree. The prediction of enzyme function is immensely important for a number of reasons. Firstly, being able to predict enzyme function more accurately will improve the functional annotation of genomes and reduce the current risk of misannotations being propagated through bioinformatics databases. Rapid developments in structural genomics, high throughput structure determination of diverse proteins from a wide variety of organisms, mean that many structures are available for enzymes whose functions are not yet known. Secondly, this project will allow us to recognise chemical similarities between evolutionarily unrelated enzymes that catalyse similar steps, though not necessarily similar overall reactions. Thirdly, this work will help us to understand the key determinants of the complex relationship between protein structure, function and evolution, particularly in terms of catalysis of reaction steps. Fourthly, the project will facilitate the design of new enzymes with either novel functions or carefully modified versions of existing functions. This project sits at an interface between disciplines, combining chemistry, biology and computer science. A wide range of skills and expertise is necessary to increase our understanding of catalysis, which has long been an important academic goal. Commercially, this work lays a foundation which is directly useful to the pharmaceutical and biotechnology industries, where enzymes are used both as diagnostics and therapeutics; the agrochemical industry, whose products often target enzymes; in the development of biofuels, which need robust enzymes to improve productivity and reduce costs; in laundry, where enzymes are already used in everyday products; and in the nutrition and food industries. In particular this project will aid in the design of new and repurposed enzymes.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1016/j.patrec.2015.06.002
发表时间: 2015-10-01
期刊: Pattern recognition letters
影响因子: 5.1
作者: [Mussa HY, Mitchell JB, Afzal AM]
通讯作者: Afzal AM
DOI: 10.1186/1471-2105-15-150
发表时间: 2014-05-19
期刊: BMC bioinformatics
影响因子: 3
作者: [De Ferrari L, Mitchell JB]
通讯作者: Mitchell JB
DOI: 10.1186/s13321-015-0105-3
发表时间: 2015
期刊: Journal of cheminformatics
影响因子: 8.6
作者: [Mussa HY, Mitchell JB, Glen RC]
通讯作者: Glen RC
DOI: 10.1186/s13104-015-1730-7
发表时间: 2015-12-03
期刊: BMC research notes
影响因子: 1.8
作者: [Mussa HY, De Ferrari L, Mitchell JB]
通讯作者: Mitchell JB
共 9 条
    AMPS: Mathematical Foundations of Market Operations with Renewable Bidders
    • 批准号:
      2229335
    • 项目类别:
      Standard Grant
    • 资助金额:
      $30.0万
    • 财政年份:
      2023
    • 负责人:
      John Mitchell
    • 依托单位:
    AMPS: Rank Minimization Algorithms for Wide-Area Phasor Measurement Data Processing
    • 批准号:
      1736326
    • 项目类别:
      Standard Grant
    • 资助金额:
      $24.0万
    • 财政年份:
      2017
    • 负责人:
      John Mitchell
    • 依托单位:
    SaTC-EDU: EAGER: Cybersecurity education for public policy
    • 批准号:
      1500089
    • 项目类别:
      Standard Grant
    • 资助金额:
      $30.0万
    • 财政年份:
      2015
    • 负责人:
      John Mitchell
    • 依托单位:
    Collaborative Research: Binary Constrained Convex Quadratic Programs with Complementarity Constraints and Extensions
    • 批准号:
      1334327
    • 项目类别:
      Standard Grant
    • 资助金额:
      $15.0万
    • 财政年份:
      2013
    • 负责人:
      John Mitchell
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
    Understanding structural evolution of galaxies with machine learning
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      Nicola Rosario Napolitano
    • 依托单位:
    煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
    • 批准号:
      --
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2022
    • 负责人:
      吉建娇
    • 依托单位:
    基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
    • 批准号:
      62003314
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      24.0万元
    • 批准年份:
      2020
    • 负责人:
      沈剑
    • 依托单位: