Domain Adaptive Hand Keypoint and Pixel Localization in the Wild

Domain Adaptive Hand Keypoint and Pixel Localization in the Wild
复制标题

DOI:
10.48550/arxiv.2203.08344
复制
发表时间:
2022-03
期刊:
--
影响因子:
--
通讯作者:
Takehiko Ohkawa;Yu-Jhe Li;Qichen Fu;Rosuke Furuta;Kris Kitani;Yoichi Sato
Takehiko Ohkawa;Yu-Jhe Li;Qichen Fu;Rosuke Furuta;Kris Kitani;Yoichi Sato
中科院分区:
其他
文献类型:
--
作者:
Takehiko Ohkawa;Yu-Jhe Li;Qichen Fu;Rosuke Furuta;Kris Kitani;Yoichi Sato

文献摘要

相似文献

我们的目标是在新的成像条件下(例如,户外)当我们仅具有在非常不同的条件下拍摄的标记图像时(例如,室内)。在真实的世界中,重要的是针对两个任务训练的模型在各种成像条件下工作。然而,现有的标记手数据集所涵盖的变化是有限的。因此,有必要将在标记图像(源)上训练的模型适应于具有不可见成像条件的未标记图像(目标)。虽然自训练域自适应方法(即,以自我监督的方式从未标记的目标图像学习),但是当对目标图像的预测有噪声时,它们的训练可能会降低性能。为了避免这种情况,在自我训练期间为噪声预测分配低重要性(置信度)权重至关重要。在本文中,我们建议利用两个预测的分歧来估计这两个任务的目标图像的置信度。这些预测来自两个独立的网络,它们的分歧有助于识别噪声预测。为了将我们提出的置信度估计集成到自我训练中,我们提出了一个教师-学生框架,其中两个网络(教师)为网络(学生)提供监督以进行自我训练,并且教师通过知识蒸馏从学生那里学习。我们的实验表明,它的优越性超过国家的最先进的方法,在适应设置不同的照明,抓物体,背景和相机的观点。与最新的对抗性适应方法相比,我们的方法将HO 3D的多任务得分提高了4%。我们还验证了我们的方法在Ego 4D,自我中心的视频与户外成像条件的快速变化。
We aim to improve the performance of regressing hand keypoints and segmenting pixel-level hand masks under new imaging conditions (e.g., outdoors) when we only have labeled images taken under very different conditions (e.g., indoors). In the real world, it is important that the model trained for both tasks works under various imaging conditions. However, their variation covered by existing labeled hand datasets is limited. Thus, it is necessary to adapt the model trained on the labeled images (source) to unlabeled images (target) with unseen imaging conditions. While self-training domain adaptation methods (i.e., learning from the unlabeled target images in a self-supervised manner) have been developed for both tasks, their training may degrade performance when the predictions on the target images are noisy. To avoid this, it is crucial to assign a low importance (confidence) weight to the noisy predictions during self-training. In this paper, we propose to utilize the divergence of two predictions to estimate the confidence of the target image for both tasks. These predictions are given from two separate networks, and their divergence helps identify the noisy predictions. To integrate our proposed confidence estimation into self-training, we propose a teacher-student framework where the two networks (teachers) provide supervision to a network (student) for self-training, and the teachers are learned from the student by knowledge distillation. Our experiments show its superiority over state-of-the-art methods in adaptation settings with different lighting, grasping objects, backgrounds, and camera viewpoints. Our method improves by 4% the multi-task score on HO3D compared to the latest adversarial adaptation method. We also validate our method on Ego4D, egocentric videos with rapid changes in imaging conditions outdoors.