Efficient and robust analysis-by-synthesis in vision : A computational framework , behavioral tests , and modeling neuronal representations
Efficient and robust analysis-by-synthesis in vision : A computational framework , behavioral tests , and modeling neuronal representations
复制标题
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Ilker Yildirim
中科院分区:
文献类型:
--
作者:
Ilker Yildirim
A glance at an object is often sufficient to recognize it and recover fine details of its shape and appearance, even under highly variable viewpoint and lighting conditions. How can vision be so rich, but at the same time robust and fast? The analysis-by-synthesis approach to vision offers an account of the richness of our percepts, but it is typically considered too fragile to apply robustly to real scenes, and too slow to explain perception in the brain. Here we propose a version of analysisby-synthesis in the spirit of the Helmholtz machine (Dayan, Hinton, Neal, & Zemel, 1995) that can be implemented efficiently and robustly, by combining a generative model based on a realistic 3D computer graphics engine with a recognition model based on a deep convolutional network fine-tuned by brief runs of MCMC inference. We test this approach in the domain of face recognition and show that it meets several challenging desiderata: it can reconstruct the approximate shape and texture of a novel face from a single view, at a level indistinguishable to humans; it accounts quantitatively for human behavior in “hard” recognition tasks that foil conventional machine systems; and it qualitatively matches neural responses in a network of face-selective brain areas. Comparison to other models provides insights to the success of our model.