Where2Act: From Pixels to Actions for Articulated 3D Objects

Where2Act: From Pixels to Actions for Articulated 3D Objects
复制标题

DOI:
10.1109/iccv48922.2021.00674
复制
发表时间:
2021-01
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Kaichun Mo;L. Guibas;Mustafa Mukadam;A. Gupta;Shubham Tulsiani
Kaichun Mo;L. Guibas;Mustafa Mukadam;A. Gupta;Shubham Tulsiani
中科院分区:
其他
文献类型:
--
作者:
Kaichun Mo;L. Guibas;Mustafa Mukadam;A. Gupta;Shubham Tulsiani

文献摘要

被引文献

相似文献

视觉感知的基本目标之一是允许代理与他们的环境进行有意义的交互。在本文中,我们朝着这一长期目标迈出了一步--我们提取与基本动作相关的高度本地化的可操作信息,例如对具有可移动部件的铰接式对象进行推送或拉动。例如,给定一个抽屉,我们的网络预测对拉手施加拉力可以打开抽屉。我们提出、讨论和评估了给出图像和深度数据的新型网络体系结构,预测在每个像素上可能发生的动作集,以及关节部分在力的作用下可能移动的区域。我们提出了一个从交互中学习的框架,该框架具有在线数据采样策略,允许我们在模拟中训练网络(SAPIEN)并跨类别泛化。查看网站上的代码和数据发布。
One of the fundamental goals of visual perception is to allow agents to meaningfully interact with their environment. In this paper, we take a step towards that long-term goal – we extract highly localized actionable information related to elementary actions such as pushing or pulling for articulated objects with movable parts. For example, given a drawer, our network predicts that applying a pulling force on the handle opens the drawer. We propose, discuss, and evaluate novel network architectures that given image and depth data, predict the set of actions possible at each pixel, and the regions over articulated parts that are likely to move under the force. We propose a learning-from-interaction framework with an online data sampling strategy that allows us to train the network in simulation (SAPIEN) and generalizes across categories. Check the website for code and data release.