THE COMPOSITE ABSOLUTE PENALTIES FAMILY FOR GROUPED AND HIERARCHICAL VARIABLE SELECTION

THE COMPOSITE ABSOLUTE PENALTIES FAMILY FOR GROUPED AND HIERARCHICAL VARIABLE SELECTION
复制标题

DOI:
10.1214/07-aos584
复制
发表时间:
2009-12-01
影响因子:
4.5
通讯作者:
Yu, Bin
Yu, Bin
中科院分区:
数学1区
文献类型:
--
作者:
Zhao, Peng;Rocha, Guilherme;Yu, Bin

文献摘要

被引文献

相似文献

从高维数据中提取有用信息是当今统计研究和实践的一个重要焦点。惩罚损失函数最小化在理论上和经验上都被证明是有效的。由于正则化和稀疏性的优点,l -1惩罚的平方误差最小化方法Lasso在回归模型及其他领域很受欢迎。在本文中,我们将包括L-1在内的不同规范组合成一个智能惩罚,以便在回归或分类模型的拟合中添加侧信息,以获得合理的估计。特别地,我们引入了复合绝对惩罚(CAP)家族,它允许在预测器之间表示给定的分组和层次关系。CAP惩罚是通过定义组并在组间和组内级别上组合规范惩罚的属性来构建的。分组选择发生在不重叠的组中。通过定义具有特定重叠模式的组来实现分层变量选择。我们建议一般使用BLASSO和交叉验证来计算CAP估计。对于仅涉及L-1和l -与规范成比例的CAP估计的子族,我们引入了iCAP算法来跟踪分组选择问题的整个正则化路径。在这个亚族中,导出了自由度(df)的无偏估计,以便在没有交叉验证的情况下选择正则化参数。在一系列模拟实验中,CAP被证明可以提高LASSO的预测性能,包括p >> n和可能错误指定分组的情况。当一个模型的复杂性被正确计算时,iCAP在实验中被认为是简约的。
Extracting useful information from high-dimensional data is an important focus of today's statistical research and practice. Penalized loss function minimization has been shown to be effective for this task both theoretically and empirically. With the virtues of both regularization and sparsity, the L-1-penalized squared error minimization method Lasso has been popular in regression models and beyond.In this paper, we combine different norms including L-1 to form an intelligent penalty in order to add side information to the fitting of a regression or classification model to obtain reasonable estimates. Specifically, we introduce the Composite Absolute Penalties (CAP) family, which allows given grouping and hierarchical relationships between the predictors to be expressed. CAP penalties are built by defining groups and combining the properties of norm penalties at the across-group and within-group levels. Grouped selection occurs for nonoverlapping groups. Hierarchical variable selection is reached by defining groups with particular overlapping patterns. We propose using the BLASSO and cross-validation to compute CAP estimates in general. For a subfamily of CAP estimates involving only the L-1 and L-proportional to norms, we introduce the iCAP algorithm to trace the entire regularization path for the grouped selection problem. Within this subfamily, unbiased estimates of the degrees of freedom (df) are derived so that the regularization parameter is selected without cross-validation. CAP is shown to improve on the predictive performance of the LASSO in a series of simulated experiments, including cases with p >> n and possibly mis-specified groupings. When the complexity of a model is properly calculated, iCAP is seen to be parsimonious in the experiments.