Background Mixup Data Augmentation for Hand and Object-in-Contact Detection

Background Mixup Data Augmentation for Hand and Object-in-Contact Detection
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Koya Tango;Takehiko Ohkawa;Ryosuke Furuta;Yoichi Sato
Koya Tango;Takehiko Ohkawa;Ryosuke Furuta;Yoichi Sato
中科院分区:
其他
文献类型:
--
作者:
Koya Tango;Takehiko Ohkawa;Ryosuke Furuta;Yoichi Sato

文献摘要

相似文献

在每个视频帧中检测手部和接触物体的位置(手-物体检测)对于从视频中理解人类活动至关重要。为了训练目标检测器,一种被称为Mixup的方法已经被经验证明对于数据增强是有效的,该方法覆盖两个训练图像以减轻数据偏差。然而,在手部对象检测中,混合两个手部操作图像会产生意外的偏差,例如,手和对象在特定区域的集中降低了手部对象检测器识别对象边界的能力。我们提出了一种称为背景混合的数据增强方法,该方法利用了数据混合正则化,同时减少了手目标检测中的意外影响。我们不是将出现手和接触物体的两幅图像混合在一起,而是将目标训练图像与从外部图像源提取的无手和接触物体的背景图像混合,并使用混合图像训练检测器。实验表明,该方法在监督和半监督学习环境下均能有效地减少误报,提高手部目标检测的性能。
Detecting the positions of human hands and objects-in-contact (hand-object detection) in each video frame is vital for understanding human activities from videos. For training an object detector, a method called Mixup, which overlays two training images to mitigate data bias, has been empirically shown to be effective for data augmentation. However, in hand-object detection, mixing two hand-manipulation images produces unintended biases, e.g., the concentration of hands and objects in a specific region degrades the ability of the hand-object detector to identify object boundaries. We propose a data-augmentation method called Background Mixup that leverages data-mixing regularization while reducing the unintended effects in hand-object detection. Instead of mixing two images where a hand and an object in contact appear, we mix a target training image with background images without hands and objects-in-contact extracted from external image sources, and use the mixed images for training the detector. Our experiments demonstrated that the proposed method can effectively reduce false positives and improve the performance of hand-object detection in both supervised and semi-supervised learning settings.