Detecting Hands and Recognizing Physical Contact in the Wild

Detecting Hands and Recognizing Physical Contact in the Wild
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Supreeth Narasimhaswamy;Trung Nguyen;Minh Hoai
Supreeth Narasimhaswamy;Trung Nguyen;Minh Hoai
中科院分区:
其他
文献类型:
--
作者:
Supreeth Narasimhaswamy;Trung Nguyen;Minh Hoai

文献摘要

被引文献

相似文献

研究了无约束条件下手部检测和物理接触状态识别的新问题。考虑到需要在手的局部外观之外进行推理,这是一项具有挑战性的推理任务。缺乏指示手接触到对象的哪些对象或部分的训练注释进一步增加了任务的复杂性。为了解决这一问题,我们提出了一种基于MASK-RCNN的卷积网络,该网络可以共同学习定位手并预测手的物理接触。该网络使用来自另一个对象检测器的输出来获得场景中存在的对象的位置。它使用这些输出和手的位置,通过两种注意力机制来识别手的接触状态。第一种注意机制基于手和区域的亲和力,将手和物体包围在一起,并将该区域的特征密集地集中到手区域。第二注意模块自适应地从这一可信的接触区域中选择显著特征。为了开发和评估我们的方法的性能,我们引入了一个名为ContactHands的大规模数据集,其中包含用手位置和接触状态标注的不受约束的图像。所提出的网络,包括注意模块的参数,是端到端可训练的。与建立在Vanilla MASK-RCNN体系结构上并用于识别手部接触状态的基线网络相比,该网络获得了大约7%的相对改进。
We investigate a new problem of detecting hands and recognizing their physical contact state in unconstrained conditions. This is a challenging inference task given the need to reason beyond the local appearance of hands. The lack of training annotations indicating which object or parts of an object the hand is in contact with further complicates the task. We propose a novel convolutional network based on Mask-RCNN that can jointly learn to localize hands and predict their physical contact to address this problem. The network uses outputs from another object detector to obtain locations of objects present in the scene. It uses these outputs and hand locations to recognize the hand's contact state using two attention mechanisms. The first attention mechanism is based on the hand and a region's affinity, enclosing the hand and the object, and densely pools features from this region to the hand region. The second attention module adaptively selects salient features from this plausible region of contact. To develop and evaluate our method's performance, we introduce a large-scale dataset called ContactHands, containing unconstrained images annotated with hand locations and contact states. The proposed network, including the parameters of attention modules, is end-to-end trainable. This network achieves approximately 7\% relative improvement over a baseline network that was built on the vanilla Mask-RCNN architecture and trained for recognizing hand contact states.