Self-Supervised Learning of Face Representations for Video Face Clustering

Self-Supervised Learning of Face Representations for Video Face Clustering
复制标题

DOI:
10.1109/fg.2019.8756609
复制
发表时间:
2019-03
期刊:
2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019)
影响因子:
--
通讯作者:
Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen
Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen
中科院分区:
其他
文献类型:
--
作者:
Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen

文献摘要

被引文献

相似文献

分析电视剧和电影背后的故事通常需要了解角色是谁,他们在做什么。随着深脸模型的改进,这似乎是一个解决了的问题。然而,随着人脸检测器的改进,需要重新考虑聚类/识别,以解决面部外观日益多样化的问题。本文提出了一种基于非监督方法的视频人脸聚类方法。我们的重点是从使用深度预训练的人脸网络获得的表示中提取基本信息,身份。我们提出了一种自监督暹罗网络,它可以在不需要基于视频/跟踪的监督的情况下进行训练,因此也可以应用于图像收集。我们在三个视频人脸聚类数据集上对我们提出的方法进行了评估。实验表明,我们的方法在所有数据集上的性能都优于目前最先进的方法。视频人脸聚类缺乏一个通用的基准,因为目前的工作往往是使用不同的度量和/或不同的人脸轨迹集进行评估。数据集和代码可在https://github.com/vivoutlaw/SSIAM.上找到
Analyzing the story behind TV series and movies often requires understanding who the characters are and what they are doing. With improving deep face models, this may seem like a solved problem. However, as face detectors get better, clustering/identification needs to be revisited to address increasing diversity in facial appearance. In this paper, we address video face clustering using unsupervised methods. Our emphasis is on distilling the essential information, identity, from the representations obtained using deep pre-trained face networks. We propose a self-supervised Siamese network that can be trained without the need for video/track based supervision, and thus can also be applied to image collections. We evaluate our proposed method on three video face clustering datasets. The experiments show that our methods outperform current state-of-the-art methods on all datasets. Video face clustering is lacking a common benchmark as current works are often evaluated with different metrics and/or different sets of face tracks. The datasets and code are available at https://github.com/vivoutlaw/SSIAM.