Sparsity information and regularization in the horseshoe and other shrinkage priors

Sparsity information and regularization in the horseshoe and other shrinkage priors
复制标题

DOI:
10.1214/17-ejs1337si
复制
发表时间:
2017-01-01
影响因子:
1.1
通讯作者:
Vehtari, Aki
Vehtari, Aki
中科院分区:
数学3区
文献类型:
--
作者:
Piironen, Juho;Vehtari, Aki

文献摘要

被引文献

相似文献

马蹄形先验已被证明是稀疏贝叶斯估计的一种值得注意的替代方案,但之前遇到了两个问题。首先,还没有基于关于参数向量中的稀疏度的先验信息来指定全局收缩超参数的先验的系统方法。其次,马蹄形先验具有不受欢迎的特性,即不可能分别指定关于最大系数的稀疏性和正则化量的信息,这对于弱识别的参数可能是有问题的,例如在数据分离的情况下的Logistic回归系数。本文针对这两个问题提出了相应的解决方案。我们引入了非零参数有效个数的概念,给出了基于稀疏性假设的全局超参数先验公式的直观方法,并论证了以前的缺省选择是可疑的,因为它们倾向于倾向于具有比我们通常预期的先验更多的未压缩参数的解。此外,我们引入了马蹄形先验的推广,称为正则化马蹄形,它允许我们将正则化的最小水平指定为最大值。我们证明了新的先验可以看作是有限板宽的钉板先验的连续对应,而原来的马蹄形类似于无限宽板的钉板。对合成数据和真实世界数据的数值实验说明了这两个理论进步的好处。
The horseshoe prior has proven to be a noteworthy alternative for sparse Bayesian estimation, but has previously suffered from two problems. First, there has been no systematic way of specifying a prior for the global shrinkage hyperparameter based on the prior information about the degree of sparsity in the parameter vector. Second, the horseshoe prior has the undesired property that there is no possibility of specifying separately information about sparsity and the amount of regularization for the largest coefficients, which can be problematic with weakly identified parameters, such as the logistic regression coefficients in the case of data separation. This paper proposes solutions to both of these problems. We introduce a concept of effective number of nonzero parameters, show an intuitive way of formulating the prior for the global hyperparameter based on the sparsity assumptions, and argue that the previous default choices are dubious based on their tendency to favor solutions with more unshrunk parameters than we typically expect a priori. Moreover, we introduce a generalization to the horseshoe prior, called the regularized horseshoe, that allows us to specify a minimum level of regularization to the largest values. We show that the new prior can be considered as the continuous counterpart of the spike-and-slab prior with a finite slab width, whereas the original horseshoe resembles the spike-and-slab with an infinitely wide slab. Numerical experiments on synthetic and real world data illustrate the benefit of both of these theoretical advances.