Latent Variable Bayesian Models for Promoting Sparsity

Latent Variable Bayesian Models for Promoting Sparsity
复制标题

DOI:
10.1109/tit.2011.2162174
复制
发表时间:
2011-09-01
影响因子:
2.5
通讯作者:
Nagarajan, Srikantan
Nagarajan, Srikantan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wipf, David P.;Rao, Bhaskar D.;Nagarajan, Srikantan

文献摘要

被引文献

相似文献

许多寻找最大稀疏系数展开式的实用方法涉及使用特定类别的凹罚函数来解决回归问题。从贝叶斯的角度来看,这个过程相当于使用稀疏诱导先验分布的最大后验概率(MAP)估计(I型估计)。使用变分技术,这个分布总是可以方便地表示为一组潜变量调制的缩放高斯分布的最大化。替代贝叶斯算法,它在潜在变量空间中操作,使这种变分表示老化,导致稀疏估计反映超出模式的后验信息(II型估计)。目前,还不清楚I型和II型的基本成本函数是如何联系的,也不清楚存在什么相关的理论属性,特别是关于II型。在此,一组公共的辅助函数用于方便地在系数或潜变量空间中表达I型和II型成本函数,以便于直接比较。在系数空间中,分析表明,II型完全等同于使用特定类别的字典和噪声相关的非阶乘系数先验进行标准MAP估计。一个先验(至少)从这个类保持几个理想的优势,所有可能的类型I的方法,并利用一种新的,非凸近似的l(0)范数与大多数,并在某些可量化的条件下,局部极小平滑。重要的是,全局最小值总是保持不变,不像标准的l(1)-范数松弛。这确保了任何适当的下降方法都能保证找到最大稀疏解。
Many practical methods for finding maximally sparse coefficient expansions involve solving a regression problem using a particular class of concave penalty functions. From a Bayesian perspective, this process is equivalent to maximum a posteriori (MAP) estimation using a sparsity-inducing prior distribution (Type I estimation). Using variational techniques, this distribution can always be conveniently expressed as a maximization over scaled Gaussian distributions modulated by a set of latent variables. Alternative Bayesian algorithms, which operate in latent variable space lever-aging this variational representation, lead to sparse estimators reflecting posterior information beyond the mode (Type II estimation). Currently, it is unclear how the underlying cost functions of Type I and Type II relate, nor what relevant theoretical properties exist, especially with regard to Type II. Herein a common set of auxiliary functions is used to conveniently express both Type I and Type II cost functions in either coefficient or latent variable space facilitating direct comparisons. In coefficient space, the analysis reveals that Type II is exactly equivalent to performing standard MAP estimation using a particular class of dictionary- and noise-dependent, nonfactorial coefficient priors. One prior (at least) from this class maintains several desirable advantages over all possible Type I methods and utilizes a novel, nonconvex approximation to the l(0) norm with most, and in certain quantifiable conditions all, local minima smoothed away. Importantly, the global minimum is always left unaltered unlike standard l(1)-norm relaxations. This ensures that any appropriate descent method is guaranteed to locate the maximally sparse solution.