Understanding 3D Object Interaction from a Single Image

Understanding 3D Object Interaction from a Single Image
复制标题

DOI:
10.1109/iccv51070.2023.01988
复制
发表时间:
2023-05
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Shengyi Qian;D. Fouhey
Shengyi Qian;D. Fouhey
中科院分区:
其他
文献类型:
--
作者:
Shengyi Qian;D. Fouhey

文献摘要

相似文献

人类可以很容易地将单个图像理解为描绘了允许交互的多个潜在对象。我们使用这种技能来规划我们与世界的互动,并在不参与互动的情况下加速理解新对象。在本文中,我们希望赋予机器类似的能力,使智能代理可以更好地探索3D场景或操纵对象。我们的方法是一个基于transformer的模型,预测物体的3D位置,物理属性和启示。为了支持这个模型,我们收集了一个包含互联网视频、以自我为中心的视频和室内图像的数据集来训练和验证我们的方法。我们的模型在我们的数据上产生了强大的性能,并且很好地推广到机器人数据。
Humans can easily understand a single image as depicting multiple potential objects permitting interaction. We use this skill to plan our interactions with the world and accelerate understanding new objects without engaging in interaction. In this paper, we would like to endow machines with the similar ability, so that intelligent agents can better explore the 3D scene or manipulate objects. Our approach is a transformer-based model that predicts the 3D location, physical properties and affordance of objects. To power this model, we collect a dataset with Internet videos, egocentric videos and indoor images to train and validate our approach. Our model yields strong performance on our data, and generalizes well to robotics data.