Clustering microarray gene expression data using weighted Chinese restaurant process

Clustering microarray gene expression data using weighted Chinese restaurant process
复制标题

DOI:
10.1093/bioinformatics/btl284
复制
发表时间:
2006-08-15
期刊:
影响因子:
5.8
通讯作者:
Qin, Zhaohui S.
Qin, Zhaohui S.
中科院分区:
生物学3区
文献类型:
--
作者:
Qin, Zhaohui S.

文献摘要

被引文献

相似文献

动机:聚类微阵列基因表达数据是阐明基因间共调控关系的有力工具。许多不同的聚类技术已经成功地应用,结果是有前途的。然而,大量的波动包含在微阵列数据,缺乏知识的数量集群和复杂的调控机制的生物systems.Results:我们设计了一种改进的基于模型的贝叶斯方法聚类微阵列基因表达数据的聚类问题具有巨大的挑战性。聚类分配是通过一个迭代加权中餐馆座位计划,使最佳的集群数量可以同时确定与集群分配。预测更新技术的应用,以提高吉布斯采样器的效率。在重新分配过程中增加了一个额外的步骤,以允许显示复杂相关关系(如时移和/或反转)的基因聚集在一起。在真实的数据集上进行的分析表明,多达30%的显著基因聚集在同一组中,显示出与簇的一致模式的复杂关系。其他值得注意的功能包括自动处理缺失数据、集群强度和分配置信度的定量测量。合成和真实的微阵列基因表达数据集进行了分析,以证明其性能。
Motivation: Clustering microarray gene expression data is a powerful tool for elucidating co-regulatory relationships among genes. Many different clustering techniques have been successfully applied and the results are promising. However, substantial fluctuation contained in microarray data, lack of knowledge on the number of clusters and complex regulatory mechanisms underlying biological systems make the clustering problems tremendously challenging.Results: We devised an improved model-based Bayesian approach to cluster microarray gene expression data. Cluster assignment is carried out by an iterative weighted Chinese restaurant seating scheme such that the optimal number of clusters can be determined simultaneously with cluster assignment. The predictive updating technique was applied to improve the efficiency of the Gibbs sampler. An additional step is added during reassignment to allow genes that display complex correlation relationships such as time-shifted and/or inverted to be clustered together. Analysis done on a real dataset showed that as much as 30% of significant genes clustered in the same group display complex relationships with the consensus pattern of the cluster. Other notable features including automatic handling of missing data, quantitative measures of cluster strength and assignment confidence. Synthetic and real microarray gene expression datasets were analyzed to demonstrate its performance.