Finding Fallen Objects Via Asynchronous Audio-Visual Integration

Finding Fallen Objects Via Asynchronous Audio-Visual Integration
复制标题

DOI:
10.1109/cvpr52688.2022.01027
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Chuang Gan;Yi Gu;Siyuan Zhou;Jeremy Schwartz;S. Alter;James Traer;Dan Gutfreund;J. Tenenbaum;Josh H. McDermott;A. Torralba
Chuang Gan;Yi Gu;Siyuan Zhou;Jeremy Schwartz;S. Alter;James Traer;Dan Gutfreund;J. Tenenbaum;Josh H. McDermott;A. Torralba
中科院分区:
其他
文献类型:
--
作者:
Chuang Gan;Yi Gu;Siyuan Zhou;Jeremy Schwartz;S. Alter;James Traer;Dan Gutfreund;J. Tenenbaum;Josh H. McDermott;A. Torralba

文献摘要

相似文献

一个物体的外观和声音提供了其物理属性的补充反映。在许多设置线索从视觉和听觉异步到达,但必须集成,当我们听到一个对象掉在地板上,然后必须找到it. In本文中,我们介绍了一个设置中,研究多模态对象定位在3D虚拟环境。一个物体落在房间的某个地方。一个配备了摄像头和麦克风的嵌入式机器人代理必须通过将音频和视觉信号与基础物理知识相结合来确定什么物体被丢弃-以及在哪里。为了研究这个问题,我们已经生成了一个大规模的数据集-Fallen Objects数据集-它包括64个房间中30个物理对象类别的8000个实例。该数据集使用ThreeDWorld平台,可以模拟基于物理的撞击声和真实感环境中对象之间复杂的物理交互。作为应对这一挑战的第一步,我们基于模仿学习、强化学习和模块化规划开发了一套具体的代理基线,并对这一新任务的挑战进行了深入分析。此数据集可公开获取11项目页面:http://fallen-object.csail.mit.edu。
The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in which to study multi-modal object localization in 3D virtual environments. An object is dropped somewhere in a room. An embodied robot agent, equipped with camera and microphone, must determine what object has been dropped - and where - by combining audio and visual signals with knowledge of the underlying physics. To study this problem, we have generated a large-scale dataset - the Fallen Objects dataset - that includes 8000 instances of 30 physical object categories in 64 rooms. The dataset uses the ThreeDWorld Platform that can simulate physics-based impact sounds and complex physical interactions between objects in a photorealistic setting. As a first step toward addressing this challenge, we develop a set of embodied agent baselines, based on imitation learning, reinforcement learning, and modular planning, and perform an in-depth analysis of the challenge of this new task. This dataset is publicly available11Project page: http://fallen-object.csail.mit.edu.