Measuring similarities between gene expression profiles through new data transformations.

Measuring similarities between gene expression profiles through new data transformations.
复制标题

DOI:
10.1186/1471-2105-8-29
复制
发表时间:
2007-01-27
期刊:
影响因子:
3
通讯作者:
Huang H
Huang H
中科院分区:
生物学4区
文献类型:
--
作者:
Kim K;Zhang S;Jiang K;Cai L;Lee IB;Feldman LJ;Huang H

文献摘要

参考文献

被引文献

相似文献

聚类方法广泛应用于基因表达数据,以对具有相似表达谱的基因进行分类。找到适当的(不)相似性度量对于分析至关重要。在我们的研究中,当关键因素是轮廓的形状,并且在确定基因关系时还应考虑表达量时,我们开发了一种新的基因聚类方法。这是通过在基因表达谱中分别对形状和幅度参数进行建模,然后使用估计的形状和幅度参数在新的特征空间中定义测量来实现的。我们探索了几种不同的变换方案来构造特征空间,包括由原始表达分量的相互差异确定特征的空间、从参数协方差矩阵导出的空间以及传统PCA分析中的主成分空间。前两者是新提出的,后者是出于比较目的而探索的。我们在这些特征空间中定义的新度量被用于 K 均值聚类过程来执行分析。将这些算法应用到模拟数据集、正在发育的小鼠视网膜 SAGE 数据集、小型酵母孢子形成 cDNA 数据集和玉米根 affymetrix 微阵列数据集,我们从结果中发现,与第一个特征空间相关的算法(名为 TransChisq)比其他方法表现出明显的优势。所提出的 TransChisq 在捕获有意义的基因表达簇方面非常有前途。这项研究还证明了数据转换在定义有效距离度量方面的重要性。我们的方法应该为分析基因表达数据提供新的见解。聚类算法可根据要求提供。
Clustering methods are widely used on gene expression data to categorize genes with similar expression profiles. Finding an appropriate (dis)similarity measure is critical to the analysis. In our study, we developed a new measure for clustering the genes when the key factor is the shape of the profile, and when the expression magnitude should also be accounted for in determining the gene relationship. This is achieved by modeling the shape and magnitude parameters separately in a gene expression profile, and then using the estimated shape and magnitude parameters to define a measure in a new feature space. We explored several different transformation schemes to construct the feature spaces that include a space whose features are determined by the mutual differences of the original expression components, a space derived from a parametric covariance matrix, and the principal component space in traditional PCA analysis. The former two are the newly proposed and the latter is explored for comparison purposes. The new measures we defined in these feature spaces were employed in a K-means clustering procedure to perform analyses. Applying these algorithms to a simulation dataset, a developing mouse retina SAGE dataset, a small yeast sporulation cDNA dataset, and a maize root affymetrix microarray dataset, we found from the results that the algorithm associated with the first feature space, named TransChisq, showed clear advantages over other methods. The proposed TransChisq is very promising in capturing meaningful gene expression clusters. This study also demonstrates the importance of data transformations in defining an efficient distance measure. Our method should provide new insights in analyzing gene expression data. The clustering algorithms are available upon request.
DOI: 10.1073/pnas.150242097
发表时间: 2000-07-18
影响因子: 11.1
作者:
Holter, NS;Mitra, M;Fedoroff, NV
通讯作者: Fedoroff, NV
DOI: 10.1073/pnas.97.18.10101
发表时间: 2000-08-29
影响因子: 11.1
作者:
Alter, O;Brown, PO;Botstein, D
通讯作者: Botstein, D
DOI: 10.1093/bioinformatics/16.11.953
发表时间: 2000-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Man, MZ;Wang, XN;Wang, YX
通讯作者: Wang, YX
DOI: 10.2307/2532201
发表时间: 1993-09-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
BANFIELD, JD;RAFTERY, AE
通讯作者: RAFTERY, AE
DOI: 10.1089/10665270252935485
发表时间: 2002-01-01
影响因子: 1.7
作者:
Filkov, V;Skiena, S;Zhi, JZ
通讯作者: Zhi, JZ