Capsule networks as recurrent models of grouping and segmentation

Capsule networks as recurrent models of grouping and segmentation
复制标题

DOI:
10.1371/journal.pcbi.1008017
复制
发表时间:
2020-07-01
影响因子:
4.3
通讯作者:
Herzog, Michael H.
Herzog, Michael H.
中科院分区:
生物学2区
文献类型:
--
作者:
Doerig, Adrien;Schmittwilken, Lynn;Herzog, Michael H.

文献摘要

被引文献

相似文献

经典地,视觉处理被描述为局部前馈计算的级联。前馈卷积神经网络(ffcnn)已经证明了这种模型的强大。然而,使用视觉拥挤作为一个控制良好的挑战,我们之前表明,没有经典的视觉模型,包括ffcnn,可以解释人类的整体形状处理。在这里,我们展示了胶囊神经网络(CapsNets),将ffcnn与循环分组和分割相结合,解决了这一挑战。我们还表明,ffcnn和标准循环cnn没有,这表明capnet的分组和分割能力是至关重要的。此外,我们提供了心理物理证据,证明分组和分割在人类中是反复实现的,并表明capnet很好地再现了这些结果。我们讨论了为什么递归似乎需要有效地实现分组和分割。总之,我们提供了相互加强的心理物理和计算证据,表明循环分组和分割过程对于理解视觉系统和创建利用全局形状计算的更好模型至关重要。
Classically, visual processing is described as a cascade of local feedforward computations. Feedforward Convolutional Neural Networks (ffCNNs) have shown how powerful such models can be. However, using visual crowding as a well-controlled challenge, we previously showed that no classic model of vision, including ffCNNs, can explain human global shape processing. Here, we show that Capsule Neural Networks (CapsNets), combining ffCNNs with recurrent grouping and segmentation, solve this challenge. We also show that ffCNNs and standard recurrent CNNs do not, suggesting that the grouping and segmentation capabilities of CapsNets are crucial. Furthermore, we provide psychophysical evidence that grouping and segmentation are implemented recurrently in humans, and show that CapsNets reproduce these results well. We discuss why recurrence seems needed to implement grouping and segmentation efficiently. Together, we provide mutually reinforcing psychophysical and computational evidence that a recurrent grouping and segmentation process is essential to understand the visual system and create better models that harness global shape computations.