Names and faces in the news

Names and faces in the news
复制标题

DOI:
10.1109/cvpr.2004.175
复制
发表时间:
2004-06
期刊:
Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004.
影响因子:
--
通讯作者:
Tamara L. Berg;A. Berg;Jaety Edwards;M. Maire;Ryan White;Y. Teh;E. Learned-Miller;D. Forsyth
Tamara L. Berg;A. Berg;Jaety Edwards;M. Maire;Ryan White;Y. Teh;E. Learned-Miller;D. Forsyth
中科院分区:
其他
文献类型:
--
作者:
Tamara L. Berg;A. Berg;Jaety Edwards;M. Maire;Ryan White;Y. Teh;E. Learned-Miller;D. Forsyth

文献摘要

被引文献

相似文献

我们表明,对于不准确和模糊标记的人脸图像的数据集,可以进行相当好的人脸聚类。我们的数据集是44,773张人脸图像,通过对大约50万张标题新闻图像应用人脸查找器获得。这个数据集比通常的人脸识别数据集更逼真,因为它包含了在野外拍摄的相对于相机的各种配置的人脸,采取了各种表情,并在颜色差异很大的照明下。每个人脸图像与一组名字相关联,这些名字是从相关联的字幕中自动提取的。许多但不是所有这样的集合都包含正确的名称。我们在适当的判别坐标下对人脸图像进行聚类。我们使用聚类过程来消除标记中的歧义,并识别错误标记的人脸。然后,合并过程识别指同一个人的名字的变体。所得到的表示可用于标记新闻图像中的人脸或由在场的个人来组织新闻图片。我们的程序的另一种观点是,我们的程序是一个清理噪声监督数据的过程。我们演示了如何使用熵度量来评估这类过程。
We show quite good face clustering is possible for a dataset of inaccurately and ambiguously labelled face images. Our dataset is 44,773 face images, obtained by applying a face finder to approximately half a million captioned news images. This dataset is more realistic than usual face recognition datasets, because it contains faces captured "in the wild" in a variety of configurations with respect to the camera, taking a variety of expressions, and under illumination of widely varying color. Each face image is associated with a set of names, automatically extracted from the associated caption. Many, but not all such sets contain the correct name. We cluster face images in appropriate discriminant coordinates. We use a clustering procedure to break ambiguities in labelling and identify incorrectly labelled faces. A merging procedure then identifies variants of names that refer to the same individual. The resulting representation can be used to label faces in news images or to organize news pictures by individuals present. An alternative view of our procedure is as a process that cleans up noisy supervised data. We demonstrate how to use entropy measures to evaluate such procedures.