Sparse Bayesian linear regression with latent masking variables.

Sparse Bayesian linear regression with latent masking variables.
复制标题

具有潜在掩蔽变量的稀疏贝叶斯线性回归。

DOI:
10.1016/j.neucom.2016.12.080
复制
发表时间:
2017
期刊:
影响因子:
6
通讯作者:
and Shin-ichi Maeda
and Shin-ichi Maeda
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yohei Kondo;Kohei Hayashi;and Shin-ichi Maeda

文献摘要

相似文献

提取任务的少量相关特征,即特征选择,通常是监督学习问题的关键步骤。稀疏线性回归为特征选择提供了快速便捷的选择,其中正则化有助于减少不相关特征的权重参数。然而,正则化也会导致相关特征的权重出现不良收缩。在这里,我们提出贝叶斯掩蔽(BM)以解决稀疏性和收缩之间的权衡问题。我们的策略是不直接对权重进行任何正则化;相反,BM 将二元潜在变量(称为掩蔽变量)引入回归模型中以保持稀疏性;每个特征和样本都有一个二元变量,其值决定该特征是否在样本中被屏蔽。我们基于因式分解信息准则(FIC)(最近提出的边际对数似然的渐近近似)推导了增强模型的变分贝叶斯推理算法。我们分析了 Lasso、自动相关性确定(ARD)和 BM 的一维估计量,从而表明 BM 在稀疏性-收缩权衡方面的优越性。最后,我们通过实验证实了我们的理论分析,并证明与 Lasso 和 ARD 相比,BM 实现了更高的特征选择精度。
Extracting a small number of relevant features for the task, i.e., feature selection, is often a crucial step in supervised learning problems. Sparse linear regression provides a fast and convenient option for feature selection, where regularization facilitates reducing the weight parameters of irrelevant features. However, the regularization also induces undesirable shrinkage in the weights of relevant features.Here, we propose Bayesian masking (BM) in order to resolve the trade-off problem between sparsity and shrinkage. Our strategy is not to directly impose any regularization on the weights; instead, BM introduces binary latent variables, called masking variables, into a regression model to keep the sparsity; each feature and sample has a binary variable whose value determines if the feature is masked or not at the sample. We derive a variational Bayesian inference algorithm for the augmented model based on the factorized information criterion (FIC), a recently-proposed asymptotic approximation of the marginal log-likelihood. We analyze the one-dimensional estimators of Lasso, automatic relevance determination (ARD), and BM, and thus show the superiority of BM in terms of the sparsity-shrinkage trade-off. Finally, we confirm our theoretical analyses through experiments and, demonstrate that BM achieves higher feature selection accuracy compared with Lasso and ARD.