How does the brain combine generative models and direct discriminative computations in high-level vision?
How does the brain combine generative models and direct discriminative computations in high-level vision?
复制标题
大脑如何将高级视觉中的生成模型和直接判别计算结合起来?
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
W. Shim
中科院分区:
文献类型:
--
作者:
Qing Yu;W. Shim
Figure 1: Two contrasting visions of vision. Separate traditions rooted in an empiricist and a rationalist conception, respectively, have envisioned the computations underlying visual perception as either a largely bottom-up, feedforward process of extracting behaviorally relevant information or as an inference process that interrogates the sensory evidence in light of a generative model that captures prior knowledge about the processes in the world that give rise to the sensory data. Several theoretically independent dimensions (double arrows) are conflated when considering the two perspectives as a dichotomy. The goal of this GAC is to disentangle the dimensions and clarify how they relate to each other, to understand how the favoured algorithm depends on the visual task, and to develop experiments that will help us understand how the primate brain combines elements of both conceptions. Our question is how the primate brain combines generative models and direct discriminative computations in high-level vision. Both approaches aim at inferring behaviorally relevant latent variables y from visual data x. In a probabilistic setting, the inference of the posterior p(y|x) is known as discriminative inference. The two approaches differ in how discriminative inference is implemented. In the generative approach, a model of the joint distribution p(y,x) of the latent variables and the visual input is employed. This model captures information about the processes in the world that give rise to the sensory data. Approximate inference algorithms are then used to infer the posterior over the latents given an image by estimating p(y|x) = p(y,x)/p(x). In the direct discriminative approach, a direct mapping from the sensory data to the posterior over the latents p(y|x) is learned without the use of an explicit generative model. The generative approach enables unsupervised learning of the structure of the world and promises better generalization to novel situations (statistical efficiency). Direct discriminative computations promise faster inferences (computational efficiency) that are accurate for new samples from the distribution experienced in training. In practice, inference of the full posterior may not be realistic and the visual system may settle for point estimates in certain cases.
DOI:
10.1364/josaa.20.001434
发表时间:
2003-07-01
影响因子:
1.9
作者:
Lee, TS;Mumford, D
通讯作者:
Mumford, D