Face recognition: A convolutional neural-network approach

Face recognition: A convolutional neural-network approach
复制标题

DOI:
10.1109/72.554195
复制
发表时间:
1997-01-01
影响因子:
--
通讯作者:
Back, AD
Back, AD
中科院分区:
其他
文献类型:
--
作者:
Lawrence, S;Giles, CL;Back, AD

文献摘要

被引文献

相似文献

面部代表复杂的多维有意义的视觉刺激,开发面部识别的计算模型很困难。我们提出了一种混合神经网络解决方案,与其他方法相比具有优势。该系统结合了局部图像采样、自组织映射(SOM)神经网络和卷积神经网络。 SOM 将图像样本量化到拓扑空间中,其中原始空间中附近的输入也在输出空间中附近,从而为图像样本中的微小变化提供降维和不变性,而卷积神经网络为平移、旋转、缩放和变形提供部分不变性。卷积网络在分层的层集中连续提取更大的特征。我们使用 Karhunen-Loeve (KL) 变换代替 SOM,并使用多层感知器 (MLP) 代替卷积网络来呈现结果。 KL 变换的性能几乎一样(错误率分别为 5.3% 和 3.8%)。 MLP 的表现非常差(错误率为 40%,错误率为 3.8%)。该方法能够快速分类,只需要快速近似归一化和预处理,并且当训练数据库中每人的图像数量从一张到五张不等时,在数据库上始终表现出比特征脸方法更好的分类性能。对于每人五幅图像,所提出的方法和特征脸分别产生 3.8% 和 10.5% 的误差。识别器提供了对其输出的置信度测量,并且当拒绝少至 10% 的示例时,分类误差接近于零,该识别器使用包含 40 个个体的 400 张图像的数据库,其中包含相当高的表情、姿势和面部细节的可变性。我们分析计算复杂性并讨论如何将新类添加到经过训练的识别器中。
Faces represent complex multidimensional meaningful visual stimuli and developing a computational model for face recognition is difficult. We present a hybrid neural-network solution which compares favorably with other methods. The system combines local image sampling, a self-organizing map (SOM) neural network, and a convolutional neural network. The SOM provides a quantization of the image samples into a topological space where inputs that are nearby in the original space are also nearby in the output space, thereby providing dimensionality reduction and invariance to minor changes in the image sample, and the convolutional neural network provides for partial invariance to translation, rotation, scale, and deformation. The convolutional network extracts successively larger features in a hierarchical set of layers. We present results using the Karhunen-Loeve (KL) transform in place of the SOM, and a multilayer perceptron (MLP) in place of the convolutional network. The KL transform performs almost as well (5.3% error versus 3.8%). The MLP performs very poorly (40% error versus 3.8%). The method is capable of rapid classification, requires only fast approximate normalization and preprocessing, and consistently exhibits better classification performance than the eigenfaces approach on the database considered as the number of images per person in the training database is varied from one to five. With five images per person the proposed method and eigenfaces result in 3.8% and 10.5% error, respectively. The recognizer provides a measure of confidence in its output and classification error approaches zero when rejecting as few as 10% of the examples, the use a database of 400 images of 40 individuals which contains quite a high degree of variability in expression, pose, and facial details. We analyze computational complexity and discuss how new classes could be added to the trained recognizer.