How does the brain combine generative models and direct discriminative computations in high-level vision?

How does the brain combine generative models and direct discriminative computations in high-level vision?
复制标题

大脑如何将高级视觉中的生成模型和直接判别计算结合起来?

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
W. Shim
W. Shim
中科院分区:
--
文献类型:
--
作者:
Qing Yu;W. Shim

文献摘要

参考文献

被引文献

相似文献

图1:两种截然不同的视觉。分别植根于数学主义和理性主义概念的不同传统,将视觉感知的计算设想为一个基本上自下而上的前馈过程,提取行为相关的信息,或者作为一个推理过程,根据生成模型询问感官证据,该模型捕获有关世界中产生感官数据的过程的先验知识。当把这两种观点视为二分法时,几个理论上独立的维度(双箭头)被合并了。这个GAC的目标是解开维度,澄清它们之间的关系,了解偏好的算法如何依赖于视觉任务,并开发实验,帮助我们了解灵长类动物的大脑如何结合这两个概念的元素。我们的问题是灵长类动物的大脑如何在高级视觉中结合生成模型和直接判别计算。这两种方法都旨在从视觉数据x推断行为相关的潜在变量y。在概率设置中,后验p(y)的推断|x)被称为判别推理。这两种方法的区别在于如何实现歧视性推理。在生成方法中,采用潜变量和视觉输入的联合分布p(y,x)的模型。这个模型捕捉了世界上产生感官数据的过程的信息。然后,近似推理算法用于通过估计p(y)来推断给定图像的潜在项上的后验|x)= p(y,x)/p(x)。在直接判别方法中,从感觉数据到潜项p(y)上的后验数据的直接映射|x)是在不使用显式生成模型的情况下学习的。生成方法使世界结构的无监督学习成为可能,并承诺更好地推广到新的情况(统计效率)。直接判别式计算保证了更快的推断(计算效率),对于来自训练中经历的分布的新样本是准确的。在实践中,全后验的推断可能不现实,并且视觉系统在某些情况下可能满足于点估计。
Figure 1: Two contrasting visions of vision. Separate traditions rooted in an empiricist and a rationalist conception, respectively, have envisioned the computations underlying visual perception as either a largely bottom-up, feedforward process of extracting behaviorally relevant information or as an inference process that interrogates the sensory evidence in light of a generative model that captures prior knowledge about the processes in the world that give rise to the sensory data. Several theoretically independent dimensions (double arrows) are conflated when considering the two perspectives as a dichotomy. The goal of this GAC is to disentangle the dimensions and clarify how they relate to each other, to understand how the favoured algorithm depends on the visual task, and to develop experiments that will help us understand how the primate brain combines elements of both conceptions. Our question is how the primate brain combines generative models and direct discriminative computations in high-level vision. Both approaches aim at inferring behaviorally relevant latent variables y from visual data x. In a probabilistic setting, the inference of the posterior p(y|x) is known as discriminative inference. The two approaches differ in how discriminative inference is implemented. In the generative approach, a model of the joint distribution p(y,x) of the latent variables and the visual input is employed. This model captures information about the processes in the world that give rise to the sensory data. Approximate inference algorithms are then used to infer the posterior over the latents given an image by estimating p(y|x) = p(y,x)/p(x). In the direct discriminative approach, a direct mapping from the sensory data to the posterior over the latents p(y|x) is learned without the use of an explicit generative model. The generative approach enables unsupervised learning of the structure of the world and promises better generalization to novel situations (statistical efficiency). Direct discriminative computations promise faster inferences (computational efficiency) that are accurate for new samples from the distribution experienced in training. In practice, inference of the full posterior may not be realistic and the visual system may settle for point estimates in certain cases.
DOI: 10.1364/josaa.20.001434
发表时间: 2003-07-01
影响因子: 1.9
作者:
Lee, TS;Mumford, D
通讯作者: Mumford, D