A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis

A penalized matrix decomposition, with applications to sparse principal components and canonical correlation analysis
复制标题

DOI:
10.1093/biostatistics/kxp008
复制
发表时间:
2009-07-01
期刊:
影响因子:
2.1
通讯作者:
Hastie, Trevor
Hastie, Trevor
中科院分区:
数学2区
文献类型:
--
作者:
Witten, Daniela M.;Tibshirani, Robert;Hastie, Trevor

文献摘要

被引文献

相似文献

我们提出了一个惩罚矩阵分解(PMD),一个新的框架计算秩K近似矩阵。我们将矩阵X近似为(X)over cap = Sigma(K)(k=1)d(k)u(k)V(k)(Ti),其中d(k),u(k)和v(k)最小化X -(X)over cap的平方Frobenius范数,受到u(k)和v(k)的惩罚。这导致奇异值分解的正则化版本。特别感兴趣的是在u(k)和v(k)上使用L-1-罚分,这产生了使用稀疏向量的X的分解。我们表明,当PMD是适用于使用L-1-惩罚v(k),而不是u(k),稀疏的主成分的结果的方法。事实上,这为“SCoTLASS”提案(Jolliffe等人,2003年)提供了一种有效的算法,用于获得稀疏主成分。该方法在公开可用的基因表达数据集上得到证明。我们还建立了稀疏主成分分析的SCoTLASS方法与Zou等人(2006)的方法之间的联系。此外,我们表明,当PMD被施加到一个叉积矩阵,它的结果在一个惩罚典型相关分析(CCA)的方法。我们将这种惩罚CCA方法应用于模拟数据和由同一组样本上的基因表达和DNA拷贝数测量组成的基因组数据集。
We present a penalized matrix decomposition (PMD), a new framework for computing a rank-K approximation for a matrix. We approximate the matrix X as (X) over cap = Sigma(K)(k=1) d(k)u(k)V(k)(T,) where d(k), u(k), and v(k) minimize the squared Frobenius norm of X - (X) over cap, subject to penalties on u(k) and v(k). This results in a regularized version of the singular value decomposition. Of particular interest is the use of L-1-penalties on u(k) and v(k), which yields a decomposition of X using sparse vectors. We show that when the PMD is applied using an L-1-penalty on v(k) but not on u(k), a method for sparse principal components results. In fact, this yields an efficient algorithm for the "SCoTLASS" proposal (Jolliffe and others 2003) for obtaining sparse principal components. This method is demonstrated on a publicly available gene expression data set. We also establish connections between the SCoTLASS method for sparse principal component analysis and the method of Zou and others (2006). In addition, we show that when the PMD is applied to a cross-products matrix, it results in a method for penalized canonical correlation analysis (CCA). We apply this penalized CCA method to simulated data and to a genomic data set consisting of gene expression and DNA copy number measurements on the same set of samples.