Unsupervised learning reveals interpretable latent representations for translucency perception.

Unsupervised learning reveals interpretable latent representations for translucency perception.
复制标题

无监督学习揭示了半透明感知的可解释的潜在表征。

DOI:
10.1371/journal.pcbi.1010878
复制
发表时间:
2023-02
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

人类不断地评估材料的外观,例如在不滑动的情况下踏上冰冷的道路。系统地发现与自然图像的材料推理有关的视觉特征。颜色以特定于比例的方式在模型的潜在空间中出现,我们可以将特定层的潜在代码嵌入到图像中,我们可以在早期的尺度上散发出对物体的范围潜在空间有选择地编码对这些层的半透明特征和操纵相干性地改变了透明度的外观,而无需改变对象的形状或身体颜色,我们发现潜在空间的中层层可以成功预测人类的半透明等级,这表明半透明的印象可以在中间的空间尺度上确定。人类的透明度感知在一起,我们的发现表明,学习自然图像的规模特定统计结构对于人类有效地表示跨环境的物质特性至关重要。 透明度是一种基本的视觉现象,支持我们与环境的相互作用。通过未标记的数据来综合清晰度的外观,并通过操纵潜在的尺度来捕获图像的统计结构,从而捕获图像的统计结构。比例尺可以预测人类的感知,从维度降低方法的基于像素的嵌入(例如,T-SNE)与我们的结果无关。想象。
Humans constantly assess the appearance of materials to plan actions, such as stepping on icy roads without slipping. Visual inference of materials is important but challenging because a given material can appear dramatically different in various scenes. This problem especially stands out for translucent materials, whose appearance strongly depends on lighting, geometry, and viewpoint. Despite this, humans can still distinguish between different materials, and it remains unsolved how to systematically discover visual features pertinent to material inference from natural images. Here, we develop an unsupervised style-based image generation model to identify perceptually relevant dimensions for translucent material appearances from photographs. We find our model, with its layer-wise latent representation, can synthesize images of diverse and realistic materials. Importantly, without supervision, human-understandable scene attributes, including the object’s shape, material, and body color, spontaneously emerge in the model’s layer-wise latent space in a scale-specific manner. By embedding an image into the learned latent space, we can manipulate specific layers’ latent code to modify the appearance of the object in the image. Specifically, we find that manipulation on the early-layers (coarse spatial scale) transforms the object’s shape, while manipulation on the later-layers (fine spatial scale) modifies its body color. The middle-layers of the latent space selectively encode translucency features and manipulation of such layers coherently modifies the translucency appearance, without changing the object’s shape or body color. Moreover, we find the middle-layers of the latent space can successfully predict human translucency ratings, suggesting that translucent impressions are established in mid-to-low spatial scale features. This layer-wise latent representation allows us to systematically discover perceptually relevant image features for human translucency perception. Together, our findings reveal that learning the scale-specific statistical structure of natural images might be crucial for humans to efficiently represent material properties across contexts. Translucency is an essential visual phenomenon, facilitating our interactions with the environment. Perception of translucent materials (i.e., materials that transmit light) is challenging to study due to the high perceptual variability of their appearance across different scenes. We present the first image-computable model that can predict human translucency judgments based on unsupervised learning from natural photographs of translucent objects. We train a deep image generation network to synthesize realistic translucent appearances from unlabeled data and learn a layer-wise latent representation that captures the statistical structure of images at multiple spatial scales. By manipulating specific layers of latent representation, we can independently modify certain visual attributes of the generated object, such as its shape, material, and color, without affecting the others. Particularly, we find the middle-layers of the latent space, which represent mid-to-low spatial scale features, can predict human perception. In contrast, the pixel-based embeddings from dimensionality reduction methods (e.g., t-SNE) do not correlate with perception. Our results suggest that scale-specific representation of visual information might be crucial for humans to perceive materials. We provide a systematic framework to discover perceptually relevant image features from natural stimuli for perceptual inference tasks and therefore valuable for understanding both human and computer vision.
DOI: 10.1016/j.isci.2022.103970
发表时间: 2022-03-18
期刊: iScience
影响因子: 5.8
作者:
Cheeseman JR;Fleming RW;Schmidt F
通讯作者: Schmidt F
DOI: 10.1016/j.cub.2011.10.036
发表时间: 2011-12-06
期刊: CURRENT BIOLOGY
影响因子: 9.2
作者:
Doerschner, Katja;Fleming, Roland W.;Kersten, Daniel
通讯作者: Kersten, Daniel
DOI: 10.1016/j.visres.2013.11.004
发表时间: 2014-01-01
期刊: VISION RESEARCH
影响因子: 1.8
作者:
Fleming, Roland W.
通讯作者: Fleming, Roland W.
DOI: 10.1167/17.3.17
发表时间: 2017-03-01
期刊: JOURNAL OF VISION
影响因子: 1.8
作者:
Chowdhury, Nahian S.;Marlow, Phillip J.;Kim, Juno
通讯作者: Kim, Juno
DOI: 10.1167/19.5.18
发表时间: 2019-05-01
期刊: JOURNAL OF VISION
影响因子: 1.8
作者:
Bi, Wenyan;Jin, Peiran;Xiao, Bei
通讯作者: Xiao, Bei