MOVES: Manipulated Objects in Video Enable Segmentation

MOVES: Manipulated Objects in Video Enable Segmentation
复制标题

DOI:
10.1109/cvpr52729.2023.00613
复制
发表时间:
2023-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Richard E. L. Higgins;D. Fouhey
Richard E. L. Higgins;D. Fouhey
中科院分区:
其他
文献类型:
--
作者:
Richard E. L. Higgins;D. Fouhey

文献摘要

相似文献

我们的方法使用视频中的操作来学习理解手持物体和手与物体的接触。我们训练一个系统,获取单个RGB图像并产生像素嵌入,该像素嵌入可用于回答分组问题(这两个像素是否组合在一起)以及手部关联问题(这只手是否握住该像素)。我们观察真实视频数据中的人,而不是煞费苦心地注释分割掩码。我们发现,将极线几何与现代光流配对可以产生简单而有效的分组伪标号。给出人的分类,我们可以进一步将像素与手联系起来,以理解接触。我们的系统在手持目标任务和手持目标任务上取得了具有竞争力的结果。
Our method uses manipulation in video to learn to understand held-objects and hand-object contact. We train a system that takes a single RGB image and produces a pixel-embedding that can be used to answer grouping questions (do these two pixels go together) as well as hand-association questions (is this hand holding that pixel). Rather than painstakingly annotate segmentation masks, we observe people in realistic video data. We show that pairing epipolar geometry with modern optical flow produces simple and effective pseudo-labels for grouping. Given people segmentations, we can further associate pixels with hands to understand contact. Our system achieves competitive results on hand and hand-held object tasks.