Video Face Clustering With Self-Supervised Representation Learning

Video Face Clustering With Self-Supervised Representation Learning
复制标题

DOI:
10.1109/tbiom.2019.2947264
复制
发表时间:
2020-04
期刊:
IEEE Transactions on Biometrics, Behavior, and Identity Science
影响因子:
--
通讯作者:
Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen
Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen
中科院分区:
其他
文献类型:
--
作者:
Vivek Sharma;Makarand Tapaswi;M. Sarfraz;R. Stiefelhagen

文献摘要

被引文献

相似文献

角色是理解电视剧和电影中传达的故事的关键组成部分。随着高级深度人脸模型的兴起,识别人脸图像似乎是一个已解决的问题。然而,随着人脸检测器变得越来越好,聚类和识别需要重新审视,以解决面部外观日益增加的多样性。在本文中,我们提出了无监督的方法,用于视频人脸聚类的特征细化。我们的重点是从使用深度预训练的人脸网络获得的表示中提取基本信息,即身份。我们提出了一个自我监督的暹罗网络,可以训练,而不需要基于视频/跟踪的监督,也可以应用于图像采集。我们在三个视频人脸聚类数据集上评估了我们的方法。包括泛化研究在内的全面实验表明,我们的方法在所有数据集上都优于当前最先进的方法。数据集和代码可在https://github.com/vivoutlaw/SSIAM上获得。
Characters are a key component of understanding the story conveyed in TV series and movies. With the rise of advanced deep face models, identifying face images may seem like a solved problem. However, as face detectors get better, clustering and identification need to be revisited to address increasing diversity in facial appearance. In this paper, we propose unsupervised methods for feature refinement with application to video face clustering. Our emphasis is on distilling the essential information, identity, from the representations obtained using deep pre-trained face networks. We propose a self-supervised Siamese network that can be trained without the need for video/track based supervision, that can also be applied to image collections. We evaluate our methods on three video face clustering datasets. Thorough experiments including generalization studies show that our methods outperform current state-of-the-art methods on all datasets. The datasets and code are available at https://github.com/vivoutlaw/SSIAM.