Seeing the Unseen: Predicting the First-Person Camera Wearer’s Location and Pose in Third-Person Scenes

Seeing the Unseen: Predicting the First-Person Camera Wearer’s Location and Pose in Third-Person Scenes
复制标题

DOI:
10.1109/iccvw54120.2021.00384
复制
发表时间:
2021-10
期刊:
2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
影响因子:
--
通讯作者:
Yangming Wen;Krishna Kumar Singh;Markham H. Anderson;Wei-Pang Jan;Yong Jae Lee
Yangming Wen;Krishna Kumar Singh;Markham H. Anderson;Wei-Pang Jan;Yong Jae Lee
中科院分区:
其他
文献类型:
--
作者:
Yangming Wen;Krishna Kumar Singh;Markham H. Anderson;Wei-Pang Jan;Yong Jae Lee

文献摘要

相似文献

我们的目标是根据相机佩戴者的第一人称可穿戴相机捕捉到的内容来预测相机佩戴者在他/她的环境中的位置和姿势。为此,我们首先收集了一个新的数据集,在这个数据集中,相机佩戴者使用时间同步的第一人称和静止的第三人称相机在不同的场景中执行各种活动(例如,打开冰箱、阅读书籍)。然后,我们提出了一种新的深度网络架构,该架构以第一人称视频帧和空的第三人称场景图像(没有摄像头佩戴者)作为输入来预测摄像头佩戴者的位置和姿势。我们探索并比较了我们的方法与几个直观的基线,并在这个新颖的,具有挑战性的问题上展示了初步的有希望的结果。
Our goal is to predict the camera wearer’s location and pose in his/her environment based on what’s captured by the camera wearer’s first-person wearable camera. Toward this goal, we first collect a new dataset in which the camera wearer performs various activities (e.g., opening a fridge, reading a book) in different scenes with time-synchronized first-person and stationary third-person cameras. We then propose a novel deep network architecture, which takes as input the first-person video frames and empty third-person scene image (without the camera wearer) to predict the location and pose of the camera wearer. We explore and compare our approach with several intuitive baselines and show initial promising results on this novel, challenging problem.