Spatiotemporal information during unsupervised learning enhances viewpoint invariant object recognition.

Spatiotemporal information during unsupervised learning enhances viewpoint invariant object recognition.
复制标题

DOI:
10.1167/15.6.7
复制
发表时间:
2015-05
期刊:
影响因子:
1.8
通讯作者:
Moqian Tian;K. Grill-Spector
Moqian Tian;K. Grill-Spector
中科院分区:
医学4区
文献类型:
--
作者:
Moqian Tian;K. Grill-Spector

文献摘要

相似文献

识别对象是困难的,因为它既需要不同对象的链接视图,也需要区分具有相似外观的对象。有趣的是,人们可以通过自然的观看统计数据,在没有监督的情况下,学习在不同的视图中识别对象,没有反馈。然而,关于在无监督学习期间使用哪些信息来链接对象视图,存在着激烈的争论。具体地说,研究人员争论,在无监督学习期间,对象视图之间的时间接近、运动或时空连续性是否有益。在这里,我们解开了这些因素中的每一个在新的三维(3-D)对象的无监督学习中的作用。我们发现,在对横跨180°视角空间的24个物体进行无监督训练后,参与者在识别旋转中的3-D物体的能力方面显示出显著的提高。令人惊讶的是,与时间接近的训练相比,具有时空连续性或运动信息的无监督学习没有优势。然而,我们发现,当参与者接受横跨同一视点空间的三分之一的视点的训练时,通过时空连续性的无监督学习比通过时间邻近学习在新视点上的识别性能要好得多。这些结果表明,虽然仅仅通过观察呈现在时间邻近的对象的多个视图来获得视点不变的识别是可能的,但时空信息通过产生比通过时间关联学习具有更广泛的视点调节的表示来提高性能。我们的发现对物体识别理论和从例子中学习的计算算法的开发具有重要意义。
Recognizing objects is difficult because it requires both linking views of an object that can be different and distinguishing objects with similar appearance. Interestingly, people can learn to recognize objects across views in an unsupervised way, without feedback, just from the natural viewing statistics. However, there is intense debate regarding what information during unsupervised learning is used to link among object views. Specifically, researchers argue whether temporal proximity, motion, or spatiotemporal continuity among object views during unsupervised learning is beneficial. Here, we untangled the role of each of these factors in unsupervised learning of novel three-dimensional (3-D) objects. We found that after unsupervised training with 24 object views spanning a 180° view space, participants showed significant improvement in their ability to recognize 3-D objects across rotation. Surprisingly, there was no advantage to unsupervised learning with spatiotemporal continuity or motion information than training with temporal proximity. However, we discovered that when participants were trained with just a third of the views spanning the same view space, unsupervised learning via spatiotemporal continuity yielded significantly better recognition performance on novel views than learning via temporal proximity. These results suggest that while it is possible to obtain view-invariant recognition just from observing many views of an object presented in temporal proximity, spatiotemporal information enhances performance by producing representations with broader view tuning than learning via temporal association. Our findings have important implications for theories of object recognition and for the development of computational algorithms that learn from examples.