Gene expression module discovery using gibbs sampling.

Gene expression module discovery using gibbs sampling.
复制标题

DOI:
10.11234/gi1990.15.239
复制
发表时间:
2004
期刊:
Genome informatics. International Conference on Genome Informatics
影响因子:
--
通讯作者:
Chang-Jiun Wu;Yutao Fu;T. Murali;S. Kasif
Chang-Jiun Wu;Yutao Fu;T. Murali;S. Kasif
中科院分区:
其他
文献类型:
--
作者:
Chang-Jiun Wu;Yutao Fu;T. Murali;S. Kasif

文献摘要

相似文献

高通量基因表达谱的最新进展催化了功能基因组学的爆炸式增长,旨在阐明在各种实验条件下不同组织或细胞类型中差异表达的基因。这些研究可以鉴定诊断基因,将基因分类为功能类别,将基因与调节途径相关联,并将基因聚类成可能由一组转录因子共同调节的模块。传统的聚类方法,如层次聚类或主成分分析,很难有效地用于这些任务,因为基因很少在广泛的条件下表现出相似的表达模式。基因表达数据的双聚类是一种很有前途的方法,用于鉴定基因组,这些基因组在一组条件下显示出一致的表达谱。这种方法可以成为发现共调控和共表达基因或模块的第一步。虽然双聚类(也称为块聚类)在1974年被引入统计学,但很少有鲁棒和有效的解决方案用于提取微阵列数据中的基因表达模块。在本文中,我们提出了一种简单但有前途的基于吉布斯抽样范式的双聚类新方法。我们的算法在GEMS (Gene Expression Module Sampler)程序中实现。GEMS已经在合成数据上进行了测试,以评估噪声对算法性能的影响,并在已发表的白血病数据集上进行了测试。通过将GEMS与其他双聚类软件进行比较,我们发现GEMS是一种可靠、灵活且计算效率高的双聚类基因表达数据处理方法。
Recent advances in high throughput profiling of gene expression have catalyzed an explosive growth in functional genomics aimed at the elucidation of genes that are differentially expressed in various tissue or cell types across a range of experimental conditions. These studies can lead to the identification of diagnostic genes, classification of genes into functional categories, association of genes with regulatory pathways, and clustering of genes into modules that are potentially co-regulated by a group of transcription factors. Traditional clustering methods such as hierarchical clustering or principal component analysis are difficult to deploy effectively for several of these tasks since genes rarely exhibit similar expression pattern across a wide range of conditions. Bi-clustering of gene expression data is a promising methodology for identification of gene groups that show a coherent expression profile across a subset of conditions. This methodology can be a first step towards the discovery of co-regulated and co-expressed genes or modules. Although bi-clustering (also called block clustering) was introduced in statistics in 1974 few robust and efficient solutions exist for extracting gene expression modules in microarray data. In this paper, we propose a simple but promising new approach for bi-clustering based on a Gibbs sampling paradigm. Our algorithm is implemented in the program GEMS (Gene Expression Module Sampler). GEMS has been tested on synthetic data generated to evaluate the effect of noise on the performance of the algorithm as well as on published leukemia datasets. In our preliminary studies comparing GEMS with other bi-clustering software we show that GEMS is a reliable, flexible and computationally efficient approach for bi-clustering gene expression data.