CATH-FunL: Improving Gene Target Selection by Predicting Functional Modules in Biological Systems
CATH-FunL: Improving Gene Target Selection by Predicting Functional Modules in Biological Systems
批准号:
BB/M020088/1
负责人:
Christine Orengo
金额:
$14.42万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
在过去的几十年里,数据可用性的显著增加已经彻底改变了生物学的研究。实验技术的进步意味着我们现在对细胞中的基因和蛋白质及其相互作用有了丰富的了解。这种前所未有的数据量给生物学家提出了一个挑战:如何最好地结合和利用不同的数据源来获得有意义的生物学见解。ath - funl就是为解决这个问题而设计的工具。FunL将允许用户预测可能与他们感兴趣的一组蛋白质相关的新蛋白质(“目标”)——例如,蛋白质信号通路中的已知成分。CATH-FunL还将允许用户通过组织和注释这个预测基因列表来进一步了解这些预测目标。最后,cats - funl将提供预测目标及其之间函数关系的直观可视化。CATH-FunL的预测方法是基于有充分证据的联想内疚概念。现代实验技术产生的许多数据可以用来推断蛋白质是否参与相同的生物过程——也就是说,它们是否在功能上相关。功能关联的证据包括蛋白质之间的物理结合,表达模式的相关性和许多其他更间接的指标。联想犯罪方法将这些信息表示为蛋白质之间的功能关联网络,并试图使用该网络的结构来预测新的关联。最简单的方法只是根据蛋白质的直接网络邻居进行预测。然而,这忽略了网络整体拓扑结构中存在的丰富信息:例如,已知具有相同功能的蛋白质组在网络中形成紧密连接的簇,与其他蛋白质的连接较少。FunL的目标是利用这种类型的结构,使用一种被称为图核(graph kernel)的强大且经过充分研究的方法。cats - funl将整合大量的蛋白质相互作用/关联信息,这些信息来自几个公共存储库和我们自己的内部蛋白质关联预测工具。这些数据将被表示为网络,结合起来,然后根据用户提供的一组查询和已知蛋白质,使用基于核的方法转换成潜在目标的排名列表。查询蛋白质将根据其与已知蛋白质的关联强度进行排序。FunL将通过提供有关其功能的信息,进一步深入了解目标蛋白。功能注释通常使用基因本体(GO)中的术语来执行。然而,平均而言,生物体中小于10%的基因已经被实验表征-因此GO注释对于许多蛋白质来说可能是稀疏的或不可靠的。因此,我们将使用最先进的、内部的、基于序列的预测方法来补充实验GO注释。一旦计算出目标列表,cat - funl将把列表组织成功能一致的子组。这将允许用户检测预测目标中的潜在模式,并将重点放在他们感兴趣的特定生物过程上。由于这种聚类所涉及的许多计算工作已经由FunL在查询阶段完成,因此这提供了一种非常有效的对目标列表蛋白质进行分类的方法。最后,FunL将以一种直观的方式将结果可视化。我们将使用基于网络的可视化,并探索与基于内核的方法相关的更多创新方法。总之,cat - funl将允许用户将他们自己的实验分析基因数据集与来自异构公共可用存储库和我们内部功能注释数据集的信息相结合,以获得他们感兴趣的生物过程的有价值的功能见解。
英文摘要
In the past decades, a marked increase in data availability has revolutionized the study of biology. Advances in experimental techniques mean that we now have an abundance of information about the genes and proteins in our cells and their interactions. This unprecedented volume of data presents a challenge for biologists: how to best combine and exploit different data sources to gain meaningful biological insights.CATH-FunL is a tool designed to address this problem. FunL will allow users to predict novel proteins ('targets') likely to be associated with a set of proteins they are interested in - for example, known components in a protein signalling pathway. CATH-FunL will also allow users to gain further insight into these predicted targets by organizing and annotating this list of predicted genes. Finally, CATH-FunL will provide intuitive visualizations of the predicted targets and the functional relations between them.CATH-FunL's prediction methods are based on the well-documented concept of guilt-by-association. Much of the data produced by modern experimental techniques can be used to infer whether proteins participate in the same biological process - that is, whether they are functionally associated. Evidence for functional association comprises physical binding between proteins, correlation in expression patterns and numerous other, more indirect indicators. Guilty-by-association methods represent this information as a network of functional associations between proteins and attempts to use the structure of the network to predict new associations.The simplest methods simply make predictions based on the direct network neighbours of a protein. This, however, ignores the rich information present in the overall topology of the network: for example, groups of proteins relating to the same function are known to form densely connected clusters within the network, with fewer connections to other proteins. FunL aims to exploit this type of structure using a powerful and well-studied approach known as graph kernels.CATH-FunL will integrate a large volume of protein interaction/association information, from several public repositories and our own in-house tools for protein association prediction. These data will be represented as networks, combined and then transformed into a ranked list of potential targets using kernel-based methods, based on a set of query and known proteins provided by the user. Query proteins will be ranked by the strength of their association to known proteins.FunL will provide further insight into the target proteins by providing information about their function. Functional annotation is often performed using terms from the Gene Ontology (GO). However, on average, <10% of genes in an organism have been experimentally characterised - GO annotations can therefore be sparse or unreliable for many proteins. Therefore, we will supplement experimental GO annotations with predicted annotations using state-of-the-art, in-house, sequence based prediction methods. Once the target list has been computed, CATH-FunL will organise the list into functionally coherent sub-groups. This will allow users to detect potential patterns in the predicted targets and to focus on particular biological processes of interest to them. Because much of the computational work involved in this clustering will already be done by FunL at the query stage, this provides a very efficient way of classifying the target list proteins.Finally, FunL will visualise the results in an intuitive way. We will use both network based visualisations and explore more innovative approaches related to the kernel-based methods.In summary, CATH-FunL will allow users to combine their own datasets of experimentally analysed genes with information from heterogeneous publicly available repositories and our in-house functional annotation datasets to gain valuable functional insights into biological processes they are interested in.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Novel Computational Protocols for Functionally Classifying and Characterising Serine Beta-Lactamases.
用于在功能上分类和表征丝氨酸β-内酰胺酶的新型计算方案。
DOI:
10.1371/journal.pcbi.1004926
发表时间:
2016-06
期刊:
PLoS computational biology
影响因子:
4.3
作者:
[Lee D, Das S, Dawson NL, Dobrijevic D, Ward J, Orengo C]
通讯作者:
Orengo C
DOI:
10.1038/ncomms13542
发表时间:
2016-12-06
期刊:
NATURE COMMUNICATIONS
影响因子:
16.6
作者:
[Erasmus, J. C., Bruche, S., Pizarro, L., Maimari, N., Pogglioli, T., Tomlinson, C., Lees, J., Zalivina, I., Wheeler, A., Alberts, A., Russo, A., Braga, V. M. M.]
通讯作者:
Braga, V. M. M.
DOI:
10.1038/s41598-017-05780-5
发表时间:
2017-07-17
期刊:
Scientific reports
影响因子:
4.6
作者:
[Yu-Wai-Man C, Owen N, Lees J, Tagalakis AD, Hart SL, Webster AR, Orengo CA, Khaw PT]
通讯作者:
Khaw PT
DOI:
10.1371/journal.pcbi.1005791
发表时间:
2017-10
期刊:
PLoS computational biology
影响因子:
4.3
作者:
[Wan C, Lees JG, Minneci F, Orengo CA, Jones DT]
通讯作者:
Jones DT
BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
-
批准号:BB/Y001117/1
-
项目类别:Research Grant
-
资助金额:$34.21万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
ProtFunAI: AI based methods for functional annotation of proteins in crop genomes
-
批准号:BB/Y514044/1
-
项目类别:Research Grant
-
资助金额:$32.43万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Improving accuracy, coverage, and sustainability of functional protein annotation in InterPro, Pfam and FunFam using Deep Learning methods PID 7012435
-
批准号:BB/X018563/1
-
项目类别:Research Grant
-
资助金额:$16.68万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Transforming the Structural Landscape of CATH to Aid Variant Analyses in Human and Agricultural Organisms and their Pathogens
-
批准号:BB/W018802/1
-
项目类别:Research Grant
-
资助金额:$111.5万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
Unlocking the chemical potential of plants: Predicting function from DNA sequence for complex enzyme superfamilies
-
批准号:BB/V014722/1
-
项目类别:Research Grant
-
资助金额:$39.23万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
CATH-FunVar - Predicting Viral and Human Variants Affecting COVID-19 Susceptibility and Severity and Repurposing Therapeutics
-
批准号:BB/W003368/1
-
项目类别:Research Grant
-
资助金额:$14.89万
-
财政年份:2021
-
负责人:Christine Orengo
-
依托单位:
3D-Gateway - Gateway to protein structure and function
-
批准号:BB/S020144/1
-
项目类别:Research Grant
-
资助金额:$37.37万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
-
批准号:BB/S020039/1
-
项目类别:Research Grant
-
资助金额:$3.42万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
SENSE - Screening of ENvironmental SEquences to discover novel protein functions, using informatics target selection and high-throughput validation
-
批准号:BB/T002735/1
-
项目类别:Research Grant
-
资助金额:$29.22万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
BBSRC-NSF/BIO Expanding the fold library in the twilight zone to facilitate structure determination of macromolecular machines
-
批准号:BB/S016007/1
-
项目类别:Research Grant
-
资助金额:$43.85万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Increasing the Coverage and Accuracy of CATH for Comparative Genomics and Variant Interpretation
-
批准号:BB/R014892/1
-
项目类别:Research Grant
-
资助金额:$79.16万
-
财政年份:2018
-
负责人:Christine Orengo
-
依托单位:
FunPDBe - Community driven enrichment of PDB data with structural and functional annotations
-
批准号:BB/P023940/1
-
项目类别:Research Grant
-
资助金额:$13.34万
-
财政年份:2017
-
负责人:Christine Orengo
-
依托单位:
Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
-
批准号:BB/N019253/1
-
项目类别:Research Grant
-
资助金额:$49.25万
-
财政年份:2016
-
负责人:Christine Orengo
-
依托单位:
An Greatly Expanded CATH-Gene3D with Functional Fingerprints to Characterise Proteins
-
批准号:BB/K020013/1
-
项目类别:Research Grant
-
资助金额:$78.03万
-
财政年份:2014
-
负责人:Christine Orengo
-
依托单位:
GENOME-3D: a UK network providing structure-based annotations for genotype to phenotype studies
-
批准号:BB/I025050/1
-
项目类别:Research Grant
-
资助金额:$37.5万
-
财政年份:2012
-
负责人:Christine Orengo
-
依托单位:
Exploiting High Performance Computing to Provide Functional Annotations via CATH-Gene3D
-
批准号:BB/H02364X/1
-
项目类别:Research Grant
-
资助金额:$13.88万
-
财政年份:2010
-
负责人:Christine Orengo
-
依托单位:
An Integrated CATH Resource for the Postgenomic Era
-
批准号:BB/F010451/1
-
项目类别:Research Grant
-
资助金额:$104.01万
-
财政年份:2008
-
负责人:Christine Orengo
-
依托单位: