Hierarchical clustering of high-throughput expression data based on general dependences.

Hierarchical clustering of high-throughput expression data based on general dependences.
复制标题

DOI:
10.1109/tcbb.2013.99
复制
发表时间:
2013-07
期刊:
IEEE/ACM transactions on computational biology and bioinformatics
影响因子:
--
通讯作者:
Peng H
Peng H
中科院分区:
其他
文献类型:
--
作者:
Yu T;Peng H

文献摘要

被引文献

相似文献

高通量表达技术,包括基因表达阵列和液相色谱-质谱(LC-MS)等,可连续测量数千个特征,即基因或代谢物。在此类数据中,特征之间既存在线性关系,也存在非线性关系。非线性关系可以反映生物系统中的关键调节模式。然而,基于线性关联的传统聚类方法并未识别和利用它们。基于一般依赖性(即线性和非线性关系)的聚类受到数据的高维度和高噪声水平的阻碍。我们开发了一种敏感的非参数测量方法,用于测量高维随机变量(组)之间的一般依赖性。基于这种依赖性度量,我们开发了一种层次聚类方法。在模拟研究中,该方法在对具有非线性依赖性的特征进行聚类方面优于基于相关性和互信息(MI)的层次聚类方法。我们将该方法应用于测量细胞周期时间序列中基因表达的微阵列数据集,以表明它产生了生物学相关的结果。 R 代码可从 http://userwww.service.emory.edu/~tyu8/GDHC 获取。
High-throughput expression technologies, including gene expression array and liquid chromatography – mass spectrometry (LC-MS) etc., measure thousands of features, i.e. genes or metabolites, on a continuous scale. In such data, both linear and nonlinear relations exist between features. Nonlinear relations can reflect critical regulation patterns in the biological system. However they are not identified and utilized by traditional clustering methods based on linear associations. Clustering based on general dependencies, i.e. both linear and nonlinear relations, is hampered by the high dimensionality and high noise level of the data. We developed a sensitive nonparametric measure of general dependency between (groups of) random variables in high-dimensions. Based on this dependency measure, we developed a hierarchical clustering method. In simulation studies, the method outperformed correlation- and mutual information (MI) – based hierarchical clustering methods in clustering features with nonlinear dependencies. We applied the method to a microarray dataset measuring the gene expression in cell-cycle time series to show it generates biologically relevant results. The R code is available at http://userwww.service.emory.edu/~tyu8/GDHC.