Efficient and robust analysis-by-synthesis in vision : A computational framework , behavioral tests , and modeling neuronal representations

Efficient and robust analysis-by-synthesis in vision : A computational framework , behavioral tests , and modeling neuronal representations
复制标题

DOI:
--
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
Ilker Yildirim
Ilker Yildirim
中科院分区:
其他
文献类型:
--
作者:
Ilker Yildirim

文献摘要

被引文献

相似文献

即使在高度变化的视点和照明条件下,看一眼物体通常也足以识别它并恢复其形状和外观的精细细节。愿景如何如此丰富,同时又强大且快速?视觉的综合分析方法描述了我们感知的丰富性,但它通常被认为太脆弱,无法稳健地应用于真实场景,而且解释大脑中的感知也太慢。在这里,我们本着亥姆霍兹机的精神提出了一种综合分析的版本(Dayan、Hinton、Neal 和 Zemel,1995),通过将基于真实 3D 计算机图形引擎的生成模型与基于深度卷积网络的识别模型相结合(通过短暂运行 MCMC 推理进行微调),可以高效、稳健地实现该版本。我们在人脸识别领域测试了这种方法,并表明它满足了几个具有挑战性的需求:它可以从单个视图重建新面孔的近似形状和纹理,达到人类无法区分的水平;它定量地解释了人类在“硬”识别任务中的行为,这些任务挫败了传统的机器系统;它定性地匹配面部选择性大脑区域网络中的神经反应。与其他模型的比较可以让我们了解我们的模型是否成功。
A glance at an object is often sufficient to recognize it and recover fine details of its shape and appearance, even under highly variable viewpoint and lighting conditions. How can vision be so rich, but at the same time robust and fast? The analysis-by-synthesis approach to vision offers an account of the richness of our percepts, but it is typically considered too fragile to apply robustly to real scenes, and too slow to explain perception in the brain. Here we propose a version of analysisby-synthesis in the spirit of the Helmholtz machine (Dayan, Hinton, Neal, & Zemel, 1995) that can be implemented efficiently and robustly, by combining a generative model based on a realistic 3D computer graphics engine with a recognition model based on a deep convolutional network fine-tuned by brief runs of MCMC inference. We test this approach in the domain of face recognition and show that it meets several challenging desiderata: it can reconstruct the approximate shape and texture of a novel face from a single view, at a level indistinguishable to humans; it accounts quantitatively for human behavior in “hard” recognition tasks that foil conventional machine systems; and it qualitatively matches neural responses in a network of face-selective brain areas. Comparison to other models provides insights to the success of our model.