Learning Video-Independent Eye Contact Segmentation from?In-the-Wild Videos

Learning Video-Independent Eye Contact Segmentation from?In-the-Wild Videos
复制标题

从野外视频中学习与视频无关的眼神接触分割

DOI:
10.1007/978-3-031-26316-3_4
复制
发表时间:
2023
期刊:
Lecture Notes in Computer Science (ACCV2022)
影响因子:
--
通讯作者:
Sugano Yusuke
Sugano Yusuke
中科院分区:
--
文献类型:
--
作者:
Wu Tianyi;Sugano Yusuke

文献摘要

相似文献

人类的目光接触是一种非语言交流形式,对社会行为有很大的影响。由于目光接触目标的位置和大小在不同的视频中有所不同,因此学习通用的视频独立的目光接触检测器仍然是一项具有挑战性的任务。在这项工作中,我们解决的任务,单向眼神接触检测视频在野外。我们的目标是建立一个统一的模型,可以识别当一个人正在看他的凝视目标在任意输入视频。考虑到这需要时间序列的相对眼球运动信息,我们建议制定一个时间分割的任务。由于标记的训练数据的稀缺性,我们进一步提出了一种凝视目标发现方法来为未标记的视频生成伪标签,这使得我们能够使用野外视频以无监督的方式训练通用的目光接触分割模型。为了评估我们提出的方法,我们手动注释了一个由52个人类对话视频组成的测试数据集。实验结果表明,我们的目光接触分割模型优于以前的视频相关的目光接触检测器,可以达到71.88%的帧的准确率在我们的注释测试集。我们的代码和评估数据集可以在https://github上找到。com/ut-vision/Video-Independent-ECS。
Human eye contact is a form of non-verbal communication and can have a great influence on social behavior. Since the location and size of the eye contact targets vary across different videos, learning a generic video-independent eye contact detector is still a challenging task. In this work, we address the task of one-way eye contact detection for videos in the wild. Our goal is to build a unified model that can identify when a person is looking at his gaze targets in an arbitrary input video. Considering that this requires time-series relative eye movement information, we propose to formulate the task as a temporal segmentation. Due to the scarcity of labeled training data, we further propose a gaze target discovery method to generate pseudo-labels for unlabeled videos, which allows us to train a generic eye contact segmentation model in an unsupervised way using in-the-wild videos. To evaluate our proposed approach, we manually annotated a test dataset consisting of 52 videos of human conversations. Experimental results show that our eye contact segmentation model outperforms the previous video-dependent eye contact detector and can achieve 71.88% framewise accuracy on our annotated test set. Our code and evaluation dataset are available at https://github. com/ut-vision/Video-Independent-ECS.