Cluster analysis of gene expression dynamics

Cluster analysis of gene expression dynamics
复制标题

DOI:
10.1073/pnas.132656399
复制
发表时间:
2002-07-09
影响因子:
11.1
通讯作者:
Kohane, IS
Kohane, IS
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Ramoni, MF;Sebatiani, P;Kohane, IS

文献摘要

被引文献

相似文献

本文提出了一种基于模型的基因表达动态聚类的贝叶斯方法。该方法将基因表达动力学表示为自回归方程,并使用聚集程序在给定可用数据的情况下搜索最可能的聚类集。这种方法的主要贡献是能够在聚类过程中考虑到基因表达时间序列的动态性质,以及确定不同聚类数量的原则方法。由于可能的聚类模型的数量随着观测时间序列的数量呈指数增长,我们设计了一个基于距离的启发式搜索过程,能够使搜索过程变得可行。这样,该方法保留了传统的基于距离的聚类的重要可视化能力,并获得了一个独立的、有原则的度量来确定两个序列的差异是否足以属于不同的聚类。这种方法依赖于基因表达动态的显式统计表示,因此可以使用标准的统计技术来评估所得模型的拟合优度并验证潜在的假设。为了研究人成纤维细胞对血清的反应,收集了一组基因表达时间序列,用于鉴定该方法的特性。
This article presents a Bayesian method for model-based clustering of gene expression dynamics. The method represents gene-expression dynamics as autoregressive equations and uses an agglomerative procedure to search for the most probable set of clusters given the available data. The main contributions of this approach are the ability to take into account the dynamic nature of gene expression time series during clustering and a principled way to identify the number of distinct clusters. As the number of possible clustering models grows exponentially with the number of observed time series, we have devised a distance-based heuristic search procedure able to render the search process feasible. In this way, the method retains the important visualization capability of traditional distance-based clustering and acquires an independent, principled measure to decide when two series are different enough to belong to different clusters. The reliance of this method on an explicit statistical representation of gene expression dynamics makes it possible to use standard statistical techniques to assess the goodness of fit of the resulting model and validate the underlying assumptions. A set of gene-expression time series, collected to study the response of human fibroblasts to serum, is used to identify the properties of the method.