Holistic Processing Develops Because it is Good
Holistic Processing Develops Because it is Good
复制标题
整体处理因其良好而得以发展
DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Lingyun Zhang
中科院分区:
文献类型:
--
作者:
G. Cottrell;Lingyun Zhang
Holistic Processing Develops Because it is Good Lingyun Zhang and Garrison W. Cottrell {lingyun,gary}@cs.ucsd.edu UCSD Computer Science and Engineering 9500 Gilman Dr., La Jolla, CA 92093-0114 USA to be found in both targets and non-targets (high false alarms) and large, specific features are unlikely to gen- eralize to more than the image they are found in (high misses). Intermediate complexity gives a good balance between the trade-offs of misses and false alarms. They suggested that these features are represented after the encoding of simple features in V1 but before the encod- ing of complex object views in anterior IT cortex, and that the features of intermediate complexity are the nat- ural result of being selected for visual classification. Abstract In this paper, we investigate the question, “what are the best features for face identification?” Ullman et al. used mutual information as measurement of how good a feature is for a class [Ullman et al., 2002]. Their ex- periments suggested that features of intermediate com- plexity are best in tasks of face vs. non-face and cars vs. non-cars. We are interested in the tasks of face identification and expression classification. We applied Ullman’s technique of finding features with high mu- tual information with category labels to the tasks. We found that features of large sizes convey the most in- formation about face identity. Local features such as eyes and mouth are informative for identity in the con- text of large face areas. Yet they are not very infor- mative by themselves, especially for an image set with high variability of facial expressions. On the other hand, small sized features around the eyes and mouth contain relatively high information for expression classification. This suggests that the appropriate feature sizes are task dependent. We suggest that holistic processing of faces has developed because these features are optimal for face identification. Figure 1: The set of fragments extracted by maximizing the Introduction amount of information delivered. (a) The features found for faces. (b) Examples of images in the training set. (c) The features found for cars. (adapted from [Ullman et al., 2002] with author’s permission). In this paper, we investigated what the best features are for face identification. Face identification is a sub- ordinate level task [Diamond and Carey, 1986] and is known to be holistic or configural. Holistic process- ing is typically taken to mean that the context of the whole face has an important contribution to process- ing the parts, and suggests that subjects use some kind of whole-face representation when processing faces: they have difficulty recognizing parts of the face in isolation, and they have difficulty ignoring parts of the face when making judgments about another part [Carey and Diamond, 1977, Carey and Diamond, 1994, Farah et al., 1995, Tanaka and Farah, 1993]. Configural processing means that subjects are sensitive to the re- lationships between the parts, e.g., the spacing between the eyes. Ullman et al. proposed using a measure of the mu- tual information between features and categories to find the features that provide the most information relevant to classification problems [Ullman et al., 2002]. Their experiments showed that features of intermediate com- plexity in size and resolution were best for classification (faces vs. non-faces, cars vs. non-cars). Figure 1 shows the combination of features they found for faces and cars. The intuition is that small, simple features are likely Will features that are good for telling faces from ob- jects be good for identification? We expect more specific features would be needed for face identification. Our re- sults show that features of large size are best for face identity classification. Local features are not as infor- mative as global ones because the variances of local fea- tures such as eyes and mouth across different images of the same person due to expressions and other fac- tors are comparable to those across individuals. We also show that the features optimal for expression classifica- tion are of small sizes. The result suggests that holistic processing for faces has been developed simply because it is good or even necessary for accurate identification. Methods Data Set 36 frontal images of 6 individuals (6 images each) from the FERET database were used [Phillips et al., 1998]. The images were aligned by rotating, scaling and crop- ping [Zhang and Cottrell, 2004]. Figure 2 shows the nor- malized face images, where each row is an individual.