Reducing microarray data via nonnegative matrix factorization for visualization and clustering analysis

Reducing microarray data via nonnegative matrix factorization for visualization and clustering analysis
复制标题

DOI:
10.1016/j.jbi.2007.12.003
复制
发表时间:
2008-08-01
影响因子:
4.5
通讯作者:
Ye, Datian
Ye, Datian
中科院分区:
医学3区
文献类型:
--
作者:
Liu, Weixiang;Yuan, Kehong;Ye, Datian

文献摘要

被引文献

相似文献

在微阵列数据分析中,每个基因表达样本具有数千个基因,并且降低如此高的维度对于样本的可视化和进一步聚类都是有用的。传统的主成分分析(PCA)是一种常用的方法,但存在问题。非负矩阵分解是一种新的降维方法。在本文中,我们比较了NMF和PCA的降维。减少的数据用于可视化,并通过11个真实的基因表达数据集上的k-means聚类分析。在聚类分析之前,我们应用NMF和PCA进行可视化约简。在一个白血病数据集上的结果表明,NMF可以发现自然聚类,并清楚地检测到一个错误标记的样本,而PCA不能。对于通过k-means的聚类分析,NMF通常优于PCA。我们的研究结果表明,NMF优于PCA在减少微阵列数据。(C)2007年爱思唯尔公司All rights reserved.
In microarray data analysis, each gene expression sample has thousands of genes and reducing such high dimensionality is useful for both visualization and further clustering of samples. Traditional principal component analysis (PCA) is a commonly used method which has problems. Nonnegative Matrix Factorization (NMF) is a new dimension reduction method. In this paper we compare NMF and PCA for dimension reduction. The reduced data is used for visualization, and clustering analysis via k-means on 11 real gene expression datasets. Before the clustering analysis, we apply NMF and PCA for reduction in visualization. The results on one leukemia dataset show that NMF can discover natural clusters and clearly detect one mislabeled sample while PCA cannot. For clustering analysis via k-means, NMF most typically outperforms PCA. Our results demonstrate the superiority of NMF over PCA in reducing microarray data. (C) 2007 Elsevier Inc. All rights reserved.