Audio-visual person recognition: an evaluation of data fusion strategies

Audio-visual person recognition: an evaluation of data fusion strategies
复制标题

视听人物识别:数据融合策略的评估

DOI:
10.1049/cp:19970414
复制
发表时间:
1997
期刊:
影响因子:
--
通讯作者:
J. Mason
J. Mason
中科院分区:
--
文献类型:
--
作者:
C. Chibelushi;F. Deravi;J. Mason

文献摘要

被引文献

相似文献

视听人识别承诺更高的识别精度比在任何一个领域的孤立识别。为了达到这一目标,应特别注意结合听觉和视觉感官形式的策略。本文提出了一个比较评估的三个决策层数据融合技术的个人身份识别。在不匹配的训练和测试噪声条件下,贝叶斯推理和Dempster-Shafer理论优于可能性理论。对于这些不匹配的噪声条件,所有三种技术都会导致集成度受损。在匹配的训练和测试噪声条件下,这三种技术产生类似的错误率,接近更准确的两种感觉方式,并显示出在低噪声水平下增强整合的迹象。本文还表明,自动识别同卵双胞胎是可能的,唇缘传达了高层次的说话人身份信息。
Audio-visual person recognition promises higher recognition accuracy than recognition in either domain in isolation. To reach this goal, special attention should be given to the strategies for combining the acoustic and visual sensory modalities. The paper presents a comparative assessment of three decision level data fusion techniques for person identification. Under mismatched training and test noise conditions, Bayesian inference and Dempster-Shafer theory are shown to outperform possibility theory. For these mismatched noise conditions, all three techniques result in compromising integration. Under matched training and test noise conditions, the three techniques yield similar error rates approaching the more accurate of the two sensory modalities, and show signs of leading to enhancing integration at low acoustic noise levels. The paper also shows that automatic identification of identical twins is possible, and that lip margins convey a high level of speaker identity information.