Correlated variables in regression: Clustering and sparse estimation

Correlated variables in regression: Clustering and sparse estimation
复制标题

DOI:
10.1016/j.jspi.2013.05.019
复制
发表时间:
2013-11-01
影响因子:
0.9
通讯作者:
Zhang, Cun-Hui
Zhang, Cun-Hui
中科院分区:
数学3区
文献类型:
--
作者:
Buehlmann, Peter;Ruetimann, Philipp;Zhang, Cun-Hui

文献摘要

被引文献

相似文献

我们考虑一个具有强相关变量的高维线性模型的估计。我们建议首先对变量进行聚类,然后进行后续的稀疏估计,例如针对聚类代表的Lasso或基于聚类结构的组Lasso。对于第一步,我们提出了一种基于典型关联的自底向上的聚类算法,并证明了它找到了一个最优解,并且在统计上是一致的。我们还提出了一些理论论点,即基于典型相关的聚类导致设计矩阵具有更好的相容性常数,以确保可识别性,并为群Lasso提供了一个oracle不等式。此外,我们讨论了集群代表和使用Lasso作为后续估计器导致预测和检测变量的改进结果的情况。我们用各种实证结果来补充理论分析。(C) 2013 Elsevier B.V.版权所有
We consider estimation in a high-dimensional linear model with strongly correlated variables. We propose to cluster the variables first and do subsequent sparse estimation such as the Lasso for cluster-representatives or the group Lasso based on the structure from the clusters. Regarding the first step, we present a novel and bottom-up agglomerative clustering algorithm based on canonical correlations, and we show that it finds an optimal solution and is statistically consistent. We also present some theoretical arguments that canonical correlation based clustering leads to a better-posed compatibility constant for the design matrix which ensures identifiability and an oracle inequality for the group Lasso. Furthermore, we discuss circumstances where cluster-representatives and using the Lasso as subsequent estimator leads to improved results for prediction and detection of variables. We complement the theoretical analysis with various empirical results. (C) 2013 Elsevier B.V. All rights reserved.