Robust Automatic Human Identification Using Face, Mouth, and Acoustic Information

Robust Automatic Human Identification Using Face, Mouth, and Acoustic Information
复制标题

使用面部、嘴巴和声音信息进行稳健的自动人体识别

DOI:
--
复制
发表时间:
2005
期刊:
Analysis and Modeling of Faces and Gestures
影响因子:
--
通讯作者:
R. Reilly
R. Reilly
中科院分区:
--
文献类型:
--
作者:
N. Fox;R. Gross;J. Cohn;R. Reilly

文献摘要

被引文献

相似文献

关于个人身份的歧视性信息是多模态的。然而,大多数人识别系统是单峰的,例如使用面部外观。为了利用不同模式的信息和增加模式识别的鲁棒性测试信号退化的互补性,我们开发了一个多专家生物特征识别系统,结合了三个专家的信息:面部,视觉语音和音频。该系统采用多模态融合在一个自动的无监督的方式,适应本地的性能和输出的可靠性,每个专家。自动选择专家权重,使得组合分数的可靠性度量最大化。为了测试系统对训练/测试失配的鲁棒性,我们使用了广泛的高斯噪声和JPEG压缩来分别降低音频和视频信号。实验在XM 2 VTS数据库上进行。在所有比较中,多模式专家系统的表现都优于单个专家。在测试的严重音频和视觉不匹配水平下,音频、嘴部、面部和三专家融合准确率分别为37.1%、48%、75%和92.7%,比表现最好的专家相对提高了23.6%。
Discriminatory information about person identity is multimodal. Yet, most person recognition systems are unimodal, e.g. the use of facial appearance. With a view to exploiting the complementary nature of different modes of information and increasing pattern recognition robustness to test signal degradation, we developed a multiple expert biometric person identification system that combines information from three experts: face, visual speech, and audio. The system uses multimodal fusion in an automatic unsupervised manner, adapting to the local performance and output reliability of each of the experts. The expert weightings are chosen automatically such that the reliability measure of the combined scores is maximized. To test system robustness to train/test mismatch, we used a broad range of Gaussian noise and JPEG compression to degrade the audio and visual signals, respectively. Experiments were carried out on the XM2VTS database. The multimodal expert system out performed each of the single experts in all comparisons. At severe audio and visual mismatch levels tested, the audio, mouth, face, and tri-expert fusion accuracies were 37.1%, 48%, 75%, and 92.7% respectively, representing a relative improvement of 23.6% over the best performing expert.