Sparse Bayesian Learning With Weakly Informative Hyperprior and Extended Predictive Information Criterion

Sparse Bayesian Learning With Weakly Informative Hyperprior and Extended Predictive Information Criterion
复制标题

DOI:
10.1109/tnnls.2021.3131357
复制
发表时间:
2021-12-09
影响因子:
10.4
通讯作者:
Kawano, Shuichi
Kawano, Shuichi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Murayama, Kazuaki;Kawano, Shuichi

文献摘要

被引文献

相似文献

本文考虑了当权重P大于数据大小N时,稀疏贝叶斯学习(SBL)的回归问题,即,P(sic)N.这种情况会导致过拟合,并使回归任务,如预测和基础选择,具有挑战性。我们展示了解决这个问题的策略。我们的战略包括两个步骤。第一种方法是在自动相关性确定(ARD)先验的噪声精度上应用形状参数接近于零的逆伽马超先验。这种超先验与增强稀疏性方面的弱信息先验的概念相关联。通过调整逆伽玛超先验的尺度参数可以控制模型的稀疏性,从而防止过拟合。二是选择最优尺度参数。我们开发了一个扩展的预测信息标准(EPIC)的最优选择。我们通过相关向量机(RVM)与多核方案处理高度非线性数据,包括光滑和不太光滑的区域的战略。这种设置是在P(sic)N情况下SBL回归任务的一种形式。作为实证评估,对四个人工数据集和八个真实的数据集进行回归分析。我们看到,过拟合被防止,而预测性能可能不会大大优于比较方法上级。我们的方法允许我们选择少量的非零权重,同时保持模型稀疏。因此,该方法有望成为有用的基础和变量的选择。
This article considers the regression problem with sparse Bayesian learning (SBL) when the number of weights P is larger than the data size N, i.e., P(sic) N. The situation induces overfitting and makes regression tasks, such as prediction and basis selection, challenging. We show a strategy to address this problem. Our strategy consists of two steps. The first is to apply an inverse gamma hyperprior with a shape parameter close to zero over the noise precision of automatic relevance determination (ARD) prior. This hyperprior is associated with the concept of a weakly informative prior in terms of enhancing sparsity. The model sparsity can be controlled by adjusting a scale parameter of inverse gamma hyperprior, leading to the prevention of overfitting. The second is to select an optimal scale parameter. We develop an extended predictive information criterion (EPIC) for optimal selection. We investigate the strategy through relevance vector machine (RVM) with a multiple-kernel scheme dealing with highly nonlinear data, including smooth and less smooth regions. This setting is one form of the regression task with SBL in the P(sic) N situation. As an empirical evaluation, regression analyses on four artificial datasets and eight real datasets are performed. We see that the overfitting is prevented, while predictive performance may be not drastically superior to comparative methods. Our methods allow us to select a small number of nonzero weights while keeping the model sparse. Thus, the methods are expected to be useful for basis and variable selection.