Holistic Processing Develops Because it is Good

Holistic Processing Develops Because it is Good
复制标题

整体处理因其良好而得以发展

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Lingyun Zhang
Lingyun Zhang
中科院分区:
--
文献类型:
--
作者:
G. Cottrell;Lingyun Zhang

文献摘要

被引文献

相似文献

整体处理的发展是因为它是好的张凌云和Garrison W. Cottrell {Lingyun,gary}@cs.ucsd.edu UCSD计算机科学与工程9500 Gilman博士,La Jolla, CA 92093-0114 USA在目标和非目标(高误报)和大的,特定的特征不太可能被一般化到比它们所发现的图像更多的地方(高误报)。中等复杂度在误报和误报之间提供了一个很好的平衡。他们认为,这些特征是在V1编码简单特征之后,而在IT前皮层编码复杂物体视图之前表征的,中间复杂性特征是被选择用于视觉分类的自然结果。在本文中,我们研究了“人脸识别的最佳特征是什么?”Ullman等人使用互信息来衡量一个特征对一个类有多好[Ullman等人,2002]。他们的实验表明,中等复杂性的特征在面孔与非面孔、汽车与非汽车的任务中表现最好。我们对人脸识别和表情分类的任务感兴趣。我们将Ullman的带类别标签的高互信息特征查找技术应用到任务中。我们发现,大尺寸的特征传达了最多的关于面部身份的信息。局部特征,如眼睛和嘴巴,在大的面部区域的背景下,为身份提供了信息。然而,它们本身并不是很有用,特别是对于具有高度面部表情可变性的图像集。另一方面,眼睛和嘴巴周围的小尺寸特征包含相对较高的表情分类信息。这表明适当的特征大小取决于任务。我们认为,人脸的整体处理已经发展起来,因为这些特征是人脸识别的最佳选择。图1:通过最大化所传递信息的引入量提取的片段集。(a)人脸的特征。(b)训练集中的图像示例。(c)汽车的特征。(经作者许可,改编自[Ullman et al., 2002])。在本文中,我们研究了人脸识别的最佳特征。人脸识别是一项从属层次的任务[Diamond和Carey, 1986],被认为是整体的或配置的。整体加工通常被认为意味着整张脸的背景对部分的加工有重要贡献,并表明受试者在加工面部时使用某种全脸表征:他们难以单独识别面部的部分,并且在判断另一部分时难以忽略面部的部分[Carey和Diamond, 1977, Carey和Diamond, 1994, Farah等人,1995,Tanaka和Farah, 1993]。构形处理意味着受试者对部分之间的关系很敏感,例如,眼睛之间的间距。Ullman等人提出使用特征和类别之间的互信息度量来找到提供与分类问题相关的最多信息的特征[Ullman等人,2002]。他们的实验表明,在尺寸和分辨率上具有中等复杂性的特征最适合分类(人脸与非人脸,汽车与非汽车)。图1显示了他们发现的人脸和汽车的特征组合。直觉告诉我们,小而简单的特征很可能用于区分人脸和物体的特征是否也能用于识别?我们预计人脸识别将需要更具体的特征。结果表明,大尺寸特征对人脸身份分类效果最好。局部特征不像全局特征那样具有信息性,因为局部特征(如眼睛和嘴)在同一个人的不同图像上的差异(由于表情和其他面部因素)与个体之间的差异是可比较的。我们还表明,最适合表情分类的特征是小尺寸的。研究结果表明,人脸整体处理的发展仅仅是因为它对准确识别是有益的,甚至是必要的。方法使用FERET数据库中6个人的36张正面图像(每人6张)[Phillips et, 1998]。通过旋转、缩放和裁剪对图像进行对齐[Zhang和Cottrell, 2004]。图2显示了未化的人脸图像,其中每一行都是一个人。
Holistic Processing Develops Because it is Good Lingyun Zhang and Garrison W. Cottrell {lingyun,gary}@cs.ucsd.edu UCSD Computer Science and Engineering 9500 Gilman Dr., La Jolla, CA 92093-0114 USA to be found in both targets and non-targets (high false alarms) and large, specific features are unlikely to gen- eralize to more than the image they are found in (high misses). Intermediate complexity gives a good balance between the trade-offs of misses and false alarms. They suggested that these features are represented after the encoding of simple features in V1 but before the encod- ing of complex object views in anterior IT cortex, and that the features of intermediate complexity are the nat- ural result of being selected for visual classification. Abstract In this paper, we investigate the question, “what are the best features for face identification?” Ullman et al. used mutual information as measurement of how good a feature is for a class [Ullman et al., 2002]. Their ex- periments suggested that features of intermediate com- plexity are best in tasks of face vs. non-face and cars vs. non-cars. We are interested in the tasks of face identification and expression classification. We applied Ullman’s technique of finding features with high mu- tual information with category labels to the tasks. We found that features of large sizes convey the most in- formation about face identity. Local features such as eyes and mouth are informative for identity in the con- text of large face areas. Yet they are not very infor- mative by themselves, especially for an image set with high variability of facial expressions. On the other hand, small sized features around the eyes and mouth contain relatively high information for expression classification. This suggests that the appropriate feature sizes are task dependent. We suggest that holistic processing of faces has developed because these features are optimal for face identification. Figure 1: The set of fragments extracted by maximizing the Introduction amount of information delivered. (a) The features found for faces. (b) Examples of images in the training set. (c) The features found for cars. (adapted from [Ullman et al., 2002] with author’s permission). In this paper, we investigated what the best features are for face identification. Face identification is a sub- ordinate level task [Diamond and Carey, 1986] and is known to be holistic or configural. Holistic process- ing is typically taken to mean that the context of the whole face has an important contribution to process- ing the parts, and suggests that subjects use some kind of whole-face representation when processing faces: they have difficulty recognizing parts of the face in isolation, and they have difficulty ignoring parts of the face when making judgments about another part [Carey and Diamond, 1977, Carey and Diamond, 1994, Farah et al., 1995, Tanaka and Farah, 1993]. Configural processing means that subjects are sensitive to the re- lationships between the parts, e.g., the spacing between the eyes. Ullman et al. proposed using a measure of the mu- tual information between features and categories to find the features that provide the most information relevant to classification problems [Ullman et al., 2002]. Their experiments showed that features of intermediate com- plexity in size and resolution were best for classification (faces vs. non-faces, cars vs. non-cars). Figure 1 shows the combination of features they found for faces and cars. The intuition is that small, simple features are likely Will features that are good for telling faces from ob- jects be good for identification? We expect more specific features would be needed for face identification. Our re- sults show that features of large size are best for face identity classification. Local features are not as infor- mative as global ones because the variances of local fea- tures such as eyes and mouth across different images of the same person due to expressions and other fac- tors are comparable to those across individuals. We also show that the features optimal for expression classifica- tion are of small sizes. The result suggests that holistic processing for faces has been developed simply because it is good or even necessary for accurate identification. Methods Data Set 36 frontal images of 6 individuals (6 images each) from the FERET database were used [Phillips et al., 1998]. The images were aligned by rotating, scaling and crop- ping [Zhang and Cottrell, 2004]. Figure 2 shows the nor- malized face images, where each row is an individual.