Nonlinear spike-and-slab sparse coding for interpretable image encoding.

Nonlinear spike-and-slab sparse coding for interpretable image encoding.
复制标题

DOI:
10.1371/journal.pone.0124088
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Lücke J
Lücke J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Shelton JA;Sheikh AS;Bornschein J;Sterne P;Lücke J

文献摘要

参考文献

被引文献

相似文献

稀疏编码是一种流行的自然图像建模方法,但面临着两个主要挑战:对低层图像分量(如边缘状结构及其遮挡)建模和对不同像素强度的建模。传统上,图像被建模为字典元素的稀疏线性叠加,其中该问题的概率观点是系数服从拉普拉斯或柯西先验分布。我们提出了一种新的模型,而不是使用尖峰和板条先验和组件的非线性组合。有了先验知识,我们的模型可以很容易地表示准确的零点,例如没有图像分量,例如边缘,以及非零像素强度上的分布。对于非线性(非线性最大组合规则),其思想是以遮挡为目标;字典元素对应于可以彼此遮挡的图像分量。这两种(非线性)方法所作的模型假设都有主要的后果,因此本文的主要目的是隔离和突出它们之间的差异。在我们的模型中,参数优化在分析和计算上都是困难的,因此作为主要贡献,我们设计了一个精确的Gibbs采样器来进行有效的推断,我们可以使用潜变量预选来应用于高维数据。在稀疏结构形式可控的自然和人工遮挡数据上的实验结果表明,该模型可以提取出与生成过程紧密匹配的稀疏类边缘成分,我们称之为可解释成分。此外,解的稀疏性与图像中组件/边缘的基本真实数量密切相关。线性模型没有学习到这种具有任何稀疏程度的边缘状分量。这表明我们的模型可以自适应地很好地逼近和表征有意义的生成过程。
Sparse coding is a popular approach to model natural images but has faced two main challenges: modelling low-level image components (such as edge-like structures and their occlusions) and modelling varying pixel intensities. Traditionally, images are modelled as a sparse linear superposition of dictionary elements, where the probabilistic view of this problem is that the coefficients follow a Laplace or Cauchy prior distribution. We propose a novel model that instead uses a spike-and-slab prior and nonlinear combination of components. With the prior, our model can easily represent exact zeros for e.g. the absence of an image component, such as an edge, and a distribution over non-zero pixel intensities. With the nonlinearity (the nonlinear max combination rule), the idea is to target occlusions; dictionary elements correspond to image components that can occlude each other. There are major consequences of the model assumptions made by both (non)linear approaches, thus the main goal of this paper is to isolate and highlight differences between them. Parameter optimization is analytically and computationally intractable in our model, thus as a main contribution we design an exact Gibbs sampler for efficient inference which we can apply to higher dimensional data using latent variable preselection. Results on natural and artificial occlusion-rich data with controlled forms of sparse structure show that our model can extract a sparse set of edge-like components that closely match the generating process, which we refer to as interpretable components. Furthermore, the sparseness of the solution closely follows the ground-truth number of components/edges in the images. The linear model did not learn such edge-like components with any level of sparsity. This suggests that our model can adaptively well-approximate and characterize the meaningful generation process.
DOI: 10.1162/neco.1995.7.3.565
发表时间: 1995-05-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
DAYAN, P;ZEMEL, RS
通讯作者: ZEMEL, RS
DOI: 10.1016/j.neucom.2012.02.055
发表时间: 2014-04-23
期刊: NEUROCOMPUTING
影响因子: 6
作者:
Frolov, Alexander A.;Husek, Dusan;Polyakov, Pavel Y.
通讯作者: Polyakov, Pavel Y.
DOI: 10.1371/journal.pcbi.1003062
发表时间: 2013
影响因子: 4.3
作者:
Bornschein J;Henniges M;Lücke J
通讯作者: Lücke J
DOI: 10.1113/jphysiol.1959.sp006308
发表时间: 1959-01-01
影响因子: 5.5
作者:
HUBEL, DH;WIESEL, TN
通讯作者: WIESEL, TN
小鼠视觉皮层中高度选择性的接受场。
DOI: 10.1523/jneurosci.0623-08.2008
发表时间: 2008-07-23
期刊: The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子: --
作者:
Niell CM;Stryker MP
通讯作者: Stryker MP