Selection and Fusion of Categorical Predictors with L 0 -Type Penalties

Selection and Fusion of Categorical Predictors with L 0 -Type Penalties
复制标题

DOI:
10.1177/1471082x14553366
复制
发表时间:
2015-10
影响因子:
1
通讯作者:
Margret-Ruth Oelker;Wolfgang Pößnecker;G. Tutz
Margret-Ruth Oelker;Wolfgang Pößnecker;G. Tutz
中科院分区:
数学4区
文献类型:
--
作者:
Margret-Ruth Oelker;Wolfgang Pößnecker;G. Tutz

文献摘要

相似文献

在回归建模中,必须对分类协变量进行编码。根据分类协变量的数量及其级别的数量,系数的数量可能会变得巨大。为了降低模型复杂度,应融合相似类别的系数,并将无影响类别的系数设置为零。为此,对系数差异进行套索式惩罚是一种标准方法。然而,这种方法的聚类/选择性能有时很差——尤其是当自适应权重条件较差或不存在时。在某些情况下,没有动力对相似的类别进行聚类。为了克服这个问题,提出了对系数差异的 L 0 惩罚,其中 L 0 “范数”被定义为向量中非零条目的数量。所提出的惩罚有利于找到对响应变量具有相同影响的类别簇,而估计精度与套索型惩罚相当。广义线性模型框架内的数值实验很有希望。为了说明这一点,我们分析了德国的失业率数据。
In regression modelling, categorical covariates have to be coded. Depending on the number of categorical covariates and on the number of levels they have, the number of coefficients can become huge. To reduce the model complexity, coefficients of similar categories should be fused and coefficients of non-influential categories should be set to zero. To this end, Lasso-type penalties on the differences of coefficients are a standard approach. However, the clustering/selection performance of this approach is sometimes poor–especially when the adaptive weights are badly conditioned or not existing. In some situations, there is no incentive to cluster similar categories. To overcome this, a L 0 penalty on the differences of coefficients is proposed, whereby the L 0 ‘norm’ is defined as the number of non-zero entries in a vector. The proposed penalty favours to find clusters of categories that share the same effect on the response variable while the estimation accuracy is comparable to Lasso-type penalties. Numerical experiments within the framework of generalized linear models are promising. For illustration, data on the unemployment rates in Germany is analyzed.