Egocentric Prediction of Action Target in 3D

Egocentric Prediction of Action Target in 3D
复制标题

DOI:
10.1109/cvpr52688.2022.02033
复制
发表时间:
2022-03
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yiming Li;Ziang Cao;Andrew Liang;Benjamin Liang;Luoyao Chen;Hang Zhao;Chen Feng
Yiming Li;Ziang Cao;Andrew Liang;Benjamin Liang;Luoyao Chen;Hang Zhao;Chen Feng
中科院分区:
其他
文献类型:
--
作者:
Yiming Li;Ziang Cao;Andrew Liang;Benjamin Liang;Luoyao Chen;Hang Zhao;Chen Feng

文献摘要

被引文献

相似文献

我们有兴趣尽早预测的目标位置,一个人的对象操作动作在3D工作空间从自我中心的视觉。它在人机协作等领域很重要,但尚未得到视觉和学习社区的足够关注。为了刺激对这个具有挑战性的以自我为中心的视觉任务的更多研究,我们提出了一个包含超过100万帧RGB-D和IMU流的大型多模态数据集,并基于我们的高质量2D和3D标签提供评估指标。同时,我们使用递归神经网络设计基线方法,并进行各种消融研究,以验证其有效性。我们的研究结果表明,这个新任务值得机器人、视觉和学习社区的研究人员进一步研究。
We are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received enough attention from vision and learning communities. To stimulate more research on this challenging egocentric vision task, we propose a large multimodality dataset of more than 1 million frames of RGB-D and IMU streams, and provide evaluation metrics based on our high-quality 2D and 3D labels from semi-automatic annotation. Meanwhile, we design baseline methods using recurrent neural networks and conduct various ablation studies to validate their effectiveness. Our results demonstrate that this new task is worthy of further study by researchers in robotics, vision, and learning communities.