Systematic discovery of functional modules and context-specific functional annotation of human genome

Systematic discovery of functional modules and context-specific functional annotation of human genome
复制标题

DOI:
10.1093/bioinformatics/btm222
复制
发表时间:
2007-07-01
期刊:
影响因子:
5.8
通讯作者:
Zhou, Xianghong Jasmine
Zhou, Xianghong Jasmine
中科院分区:
生物学3区
文献类型:
--
作者:
Huang, Yu;Li, Haifeng;Zhou, Xianghong Jasmine

文献摘要

被引文献

相似文献

动机:微阵列数据集的快速积累为人类基因组的系统功能表征提供了独特的机会。我们设计了一种基于图的方法来整合跨平台的微阵列数据,并提取重复表达模式。一系列微阵列数据集可以建模为一系列共表达网络,我们在其中搜索频繁出现的网络模式。整合的方法提供了三个主要的优势,通常使用的微阵列分析方法:(1)增强信号噪声分离(2)确定功能相关的基因没有共表达和(3)提供了一种方法来预测基因功能的上下文特定的方式。我们开发了一个基于频繁项集挖掘和双聚类的数据挖掘程序,系统地发现至少在五个数据集中重复出现的网络模式。由此产生了143 401个潜在的功能模块。随后,我们设计了一个网络拓扑统计的基础上图随机游走,有效地捕捉基因的局部功能环境的特点。然后,基于此统计数据的功能注释将使用随机森林方法进行评估,并结合网络模块的其他六个属性。我们将1126个功能分配给895个基因,其中779个已知,116个未知,验证准确率为70%。在我们的分配中,20%的基因被分配了基于不同网络环境的多种功能。
Motivation: The rapid accumulation of microarray datasets provides unique opportunities to perform systematic functional characterization of the human genome. We designed a graph-based approach to integrate cross-platform microarray data, and extract recurrent expression patterns. A series of microarray datasets can be modeled as a series of co-expression networks, in which we search for frequently occurring network patterns. The integrative approach provides three major advantages over the commonly used microarray analysis methods: (1) enhance signal to noise separation (2) identify functionally related genes without co-expression and (3) provide a way to predict gene functions in a context-specific way.Results: We integrate 65 human microarray datasets, comprising 1105 experiments and over 11 million expression measurements. We develop a data mining procedure based on frequent itemset mining and biclustering to systematically discover network patterns that recur in at least five datasets. This resulted in 143 401 potential functional modules. Subsequently, we design a network topology statistic based on graph random walk that effectively captures characteristics of a gene's local functional environment. Function annotations based on this statistic are then subject to the assessment using the random forest method, combining six other attributes of the network modules. We assign 1126 functions to 895 genes, 779 known and 116 unknown, with a validation accuracy of 70%. Among our assignments, 20% genes are assigned with multiple functions based on different network environments.