Development of a graph-theoretic approach to predict protein function by integrating large scale heterogeneous data
Development of a graph-theoretic approach to predict protein function by integrating large scale heterogeneous data
批准号:
BB/F00964X/1
负责人:
Alberto Paccanaro
金额:
$53.49万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --
中文摘要
具有完整基因组序列的生物体名单不断增长,这导致鉴定出数千个功能未知的基因。这些基因可能潜在地参与重要的生物细胞功能,可能代表诊断和药物基因组学研究的重要目标,并具有工业和农学意义。因此,生物学的一项主要任务是在基因组尺度上识别这些未表征基因的功能。生物信息学面临的挑战是设计算法方法,给定一个基因,可以预测其功能的假设,然后通过湿实验室分析进行验证。幸运的是,新的实验技术已经出现,产生的数据提供了蛋白质功能的线索,因此可以用于功能预测,例如蛋白质相互作用数据,基因表达数据。一些实验和计算数据具有网络的自然表示(例如蛋白质相互作用数据),其他数据本质上是“一维的”(例如序列模式)。最近有三个事实变得清晰:虽然每种数据类型都包含有助于确定蛋白质功能的重要信息,但没有一种数据类型本身就足够了;通过整合不同来源的证据,大大提高了大规模功能推理;对于那些可以表示为网络的数据类型,利用网络拓扑结构的算法可以获得最佳结果。到目前为止,在网络上进行功能推断的方法在它们可以集成的数据类型方面非常有限,而可以集成更多种数据的方法没有利用网络的拓扑结构。我打算研究一种通用的方法,这种方法可以考虑到其内在结构,基本上可以集成当前可用的任何数据类型:它利用网络数据的图拓扑,并且可以将这些证据与一维信息集成在一起。我将开发图论方法,利用图上的信息扩散,从网络数据中生成功能证据。然后使用机器学习技术将这些证据与其他一维信息结合起来。该方法的优势在于它能够使用不同的噪声数据集,并将它们组合起来以获得可靠的统计推断;每个数据集中包含的弱信号通过数据整合得到增强。该方法将首先在酵母上发展,然后我将把这种方法转移到更高级的生物,如秀丽隐杆线虫、黑腹线虫、拟南芥和智人。对于所有这些生物,算法的性能将通过测试集“在计算机中”进行评估;也就是说,我将验证预测已知注释的基因功能的方法的准确性。然后,这种方法将在形成信号通路(MAPK信号)的基因子网络上“在体内”进行测试,并发挥将信息从受体传递到基因表达的功能。MAPK通路组分在模式植物拟南芥中高度多样化,有123个组分。对于其中的许多,我们不知道它们是如何连接起来的,也不知道它们的生物学功能是什么。这些将由算法预测,然后通过使用RNA干扰和在突变系中沉默它们的表达来进行功能测试。我还将设计和实施独立的和基于网络的软件工具,结合所开发的算法。应用程序将使生物学家能够通过用户友好的界面轻松应用算法;可视化相关的生物网络,从而使推理过程透明化,并为系统预测的功能注释提供解释。还将创建一个web工具。所有这些工具都将免费提供给科学界。
英文摘要
The list of organisms with completed genome sequence is continuously growing and this has led to the identification of thousands of genes whose function is still unknown. These genes could potentially be involved in important biological cell functions and could represent important targets for diagnostic and pharmacogenomics studies and be of industrial and agronomical importance. A major undertaking for biology is therefore that of identifying the function of these uncharacterized genes on a genomic scale. The challenge for bioinformatics is then to devise algorithmic methods that, given a gene, can predict a hypothesis for its function that can then be validated by wet-lab assays. Luckily, new experimental techniques have become available, producing data which offer clues about protein function and can therefore be employed for function prediction, e.g. protein interaction data, gene expression data. Some experimental and computational data have a natural representation as networks (e.g. protein interaction data), others are inherently 'one-dimensional' (e.g. sequence patterns). Three facts have recently become clear: while each data type contains important information that can help in determining the function of a protein, no single data type by itself suffices; large-scale functional inference greatly improves by integrating evidence from different sources; for those data types which can be represented as networks, the best results are obtained by algorithms that take advantage of the networks' topologies. So far, methods that make functional inferences on networks are very limited in the type of data they can integrate, while methods that can integrate a greater variety of data do not take advantage of the networks' topologies. I intend to investigate a general method that can integrate essentially any data type currently available taking into account its intrinsic structure: it takes advantage of the graph topology for network data, and it can integrate this evidence together with one-dimensional information. I shall develop graph-theoretical methods that use the diffusion of information over graphs to generate functional evidence from network data. This evidence is then combined with other one-dimensional information using machine learning techniques. The strength of the methodology lies in its ability to use diverse sets of noisy data, and to combine them to obtain sound statistical inferences; the weak signals contained in each dataset is enhanced by integrating the data. The methodology will be first developed on Yeast, and I shall then transfer this approach to higher organisms such as C. elegans, D. melanogaster, A. thaliana, and H. sapiens. For all these organisms the performance of the algorithms will then be evaluated 'in silico' by means of test sets; that is I shall verify the accuracy of the methods at predicting the function for genes whose annotation is known. The approach will then be tested 'in vivo' on a sub-network of genes that form signalling pathways (MAPK signalling) and function to transmit information from receptors to gene expression. MAPK pathway components are highly diversified in the model plant, Arabidopsis thaliana, with 123 components. For many of these we do not know how they connect up and what their biological functions are. These will be predicted by the algorithms and then functionally tested by silencing their expression using RNA interference and in mutant lines. I shall also design and implement stand-alone and web-based software tools incorporating the algorithms developed. The applications will enable the biologist to easily apply the algorithms through a user-friendly interface; to visualize the relevant biological networks thus making the inference process transparent and providing an explanation for the functional annotation predicted by the system. A web tool will also be created. All these tools will be made freely available to the scientific community.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1016/j.jbi.2023.104295
发表时间:
2023-03
期刊:
JOURNAL OF BIOMEDICAL INFORMATICS
影响因子:
4.5
作者:
[Casiraghi, Elena, Wong, Rachel, Hall, Margaret, Coleman, Ben, Notaro, Marco, Evans, Michael D., Tronieri, Jena S., Blau, Hannah, Laraway, Bryan, Callahan, Tiffany J., Chan, Lauren E., Bramante, Carolyn T., Buse, John B., Moffitt, Richard A., Sturmer, Til, Johnson, Steven G., Shao, Yu Raymond, Reese, Justin, Robinson, Peter N., Paccanaro, Alberto, Valentini, Giorgio, Huling, Jared D., Wilkins, Kenneth J.]
通讯作者:
Wilkins, Kenneth J.
Combining interactomes from multiple organisms: A case study on human-mouse
结合多种生物体的相互作用组:人鼠案例研究
DOI:
10.1109/clei.2016.7833324
发表时间:
2016
期刊:
影响因子:
--
作者:
[Caceres J]
通讯作者:
Caceres J
Additional file 1 of LUMI-PCR: an Illumina platform ligation-mediated PCR protocol for integration site cloning, provides molecular quantitation of integration sites
LUMI-PCR 的附加文件 1:用于整合位点克隆的 Illumina 平台连接介导的 PCR 方案,提供整合位点的分子定量
DOI:
10.6084/m9.figshare.11805027
发表时间:
2020
期刊:
影响因子:
--
作者:
[Dawes J]
通讯作者:
Dawes J
DOI:
10.1038/s41431-023-01511-9
发表时间:
2024-01-10
期刊:
EUROPEAN JOURNAL OF HUMAN GENETICS
影响因子:
5.2
作者:
[Caniza,Horacio, Caceres,Juan J., Paccanaro,Alberto]
通讯作者:
Paccanaro,Alberto
DOI:
10.1038/srep17658
发表时间:
2015-12-03
期刊:
Scientific reports
影响因子:
4.6
作者:
[Caniza H, Romero AE, Paccanaro A]
通讯作者:
Paccanaro A
共 8 条
A GPU-based high performance system for discovering consensus domain architecture and functional annotation of protein families
-
批准号:BB/K004131/1
-
项目类别:Research Grant
-
资助金额:$14.61万
-
财政年份:2012
-
负责人:Alberto Paccanaro
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于Graph-PINN的层结稳定度参数化建模与沙尘跨介质耦合传输模拟研
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:梅奥
-
依托单位:
平面三角剖分flip graph的强凸性研究
-
批准号:12301432
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:王子丽
-
依托单位:
基于graph的多对比度磁共振图像重建方法
-
批准号:61901188
-
项目类别:青年科学基金项目
-
资助金额:24.5万元
-
批准年份:2019
-
负责人:赖宗英
-
依托单位:
基于de bruijn graph梳理的宏基因组拼接算法开发
-
批准号:61771009
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2017
-
负责人:李国君
-
依托单位:
基于Graph和ISA的红外目标分割与识别方法研究
-
批准号:61101246
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2011
-
负责人:刘靳
-
依托单位:
固定参数可解算法在平面图问题的应用以及和整数线性规划的关系
-
批准号:60973026
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2009
-
负责人:鲁道夫
-
依托单位:
图的一般染色数与博弈染色数
-
批准号:10771035
-
项目类别:面上项目
-
资助金额:18.0万元
-
批准年份:2007
-
负责人:杨大庆
-
依托单位:
中国Web Graph的挖掘与应用研究
-
批准号:60473122
-
项目类别:面上项目
-
资助金额:23.0万元
-
批准年份:2004
-
负责人:俞勇
-
依托单位: