Extracting expression modules from perturbational gene expression compendia.

Extracting expression modules from perturbational gene expression compendia.
复制标题

DOI:
10.1186/1752-0509-2-33
复制
发表时间:
2008-04-10
影响因子:
--
通讯作者:
Kuiper, Martin
Kuiper, Martin
中科院分区:
生物2区
文献类型:
--
作者:
Maere, Steven;Van Dijck, Patrick;Kuiper, Martin

文献摘要

参考文献

被引文献

相似文献

从系统生物学的角度来看,化学和遗传扰动下的基因表达谱纲要构成了一个宝贵的资源。然而,这些数据的扰动性质对用于分析它们的计算方法提出了具体的挑战。特别是,传统的聚类算法有困难的扰动纲要,即部分基因之间的共表达关系的一个突出特点。另一方面,双聚类方法是专门设计来捕获这样的部分共表达模式,但它们显示出各种其他缺点。例如,一些双聚类方法不太适合识别重叠的双聚类,而其他方法则生成高度冗余的双聚类。此外,没有一个现有的双聚类工具利用的主食扰动表达数据分析:差异表达基因的识别。我们介绍了一种新的方法,称为ENIGMA,解决了其中的一些问题。ENIGMA利用差异表达分析结果从扰动基因表达数据中提取表达模块。ENIGMA聚类过程的核心参数自动优化,以减少模块之间的冗余。与大多数其他方法产生的双簇相反,ENIGMA模块可能显示内部子结构,即具有不同但显著相关的表达模式的基因子集。将这些(通常是功能上)相关的模式组合在一个模块中,极大地帮助了对数据的生物学解释。我们表明,ENIGMA优于人工数据集上的其他方法,使用的质量标准,与其他标准不同,可以用于生成重叠聚类的算法,并且可以修改以考虑聚类之间的冗余。最后,我们将ENIGMA应用于酿酒酵母表达谱的Rosetta纲要,并更详细地分析了一个信息素响应相关模块,展示了ENIGMA产生详细预测的潜力。人们越来越认识到,微扰表达纲要是必不可少的,以确定细胞功能的基因网络,并努力建立这些不同的生物体目前正在进行中。我们表明,ENIGMA构成了一个有价值的除了剧目的方法来分析这些数据。
Compendia of gene expression profiles under chemical and genetic perturbations constitute an invaluable resource from a systems biology perspective. However, the perturbational nature of such data imposes specific challenges on the computational methods used to analyze them. In particular, traditional clustering algorithms have difficulties in handling one of the prominent features of perturbational compendia, namely partial coexpression relationships between genes. Biclustering methods on the other hand are specifically designed to capture such partial coexpression patterns, but they show a variety of other drawbacks. For instance, some biclustering methods are less suited to identify overlapping biclusters, while others generate highly redundant biclusters. Also, none of the existing biclustering tools takes advantage of the staple of perturbational expression data analysis: the identification of differentially expressed genes. We introduce a novel method, called ENIGMA, that addresses some of these issues. ENIGMA leverages differential expression analysis results to extract expression modules from perturbational gene expression data. The core parameters of the ENIGMA clustering procedure are automatically optimized to reduce the redundancy between modules. In contrast to the biclusters produced by most other methods, ENIGMA modules may show internal substructure, i.e. subsets of genes with distinct but significantly related expression patterns. The grouping of these (often functionally) related patterns in one module greatly aids in the biological interpretation of the data. We show that ENIGMA outperforms other methods on artificial datasets, using a quality criterion that, unlike other criteria, can be used for algorithms that generate overlapping clusters and that can be modified to take redundancy between clusters into account. Finally, we apply ENIGMA to the Rosetta compendium of expression profiles for Saccharomyces cerevisiae and we analyze one pheromone response-related module in more detail, demonstrating the potential of ENIGMA to generate detailed predictions. It is increasingly recognized that perturbational expression compendia are essential to identify the gene networks underlying cellular function, and efforts to build these for different organisms are currently underway. We show that ENIGMA constitutes a valuable addition to the repertoire of methods to analyze such data.
DOI: 10.1016/s0165-1684(02)00475-9
发表时间: 2003-04-01
期刊: SIGNAL PROCESSING
影响因子: 4.4
作者:
Bolshakova, N;Azuaje, F
通讯作者: Azuaje, F
DOI: 10.1038/35011540
发表时间: 1999-12-02
期刊: NATURE
影响因子: 64.8
作者:
Hartwell, LH;Hopfield, JJ;Murray, AW
通讯作者: Murray, AW
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1186/1471-2105-4-2
发表时间: 2003-01-13
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Bader, GD;Hogue, CW
通讯作者: Hogue, CW
DOI: 10.1103/physreve.67.031902
发表时间: 2003-03-01
期刊: PHYSICAL REVIEW E
影响因子: 2.4
作者:
Bergmann, S;Ihmels, J;Barkai, N
通讯作者: Barkai, N