An evaluation of sound source identification with RWCP sound scene database in real acoustic environments

An evaluation of sound source identification with RWCP sound scene database in real acoustic environments
复制标题

真实声环境下RWCP声音场景数据库声源识别评估

DOI:
10.1109/icme.2002.1035570
复制
发表时间:
2002
期刊:
Proceedings. IEEE International Conference on Multimedia and Expo
影响因子:
--
通讯作者:
Satoshi Nakamura
Satoshi Nakamura
中科院分区:
--
文献类型:
--
作者:
T. Nishiura;Satoshi Nakamura

文献摘要

被引文献

相似文献

对于免提语音接口来说,高质量地捕捉远处的语音非常重要。麦克风阵列是实现此目的的理想选择。然而,这种方法需要定位目标说话者。传统的多声源环境下的说话人定位方法不仅难以准确定位多个声源,而且难以在已知的多个声源位置中定位目标说话人。为了解决这些问题,我们提出了一种由两种算法组成的新的说话者定位方法。一种算法是基于 CSP(互功率谱相位)分析的多声源定位。另一种算法用于在定位的多个声源中进行声源识别,以实现讲话者定位。我们特别关注后者的统计声源识别,其中具有基于 GMM(高斯混合模型)的统计语音和环境声音模型以及用于讲话者定位的麦克风阵列。我们特别使用真实声学环境中的 RWCP 声音场景数据库(RWCP-DB)评估了所提出算法的性能。
It is very important for a hands-free speech interface to capture distant speech with high quality. A microphone array is an ideal candidate for this purpose. However, this approach requires localizing the target talker. Conventional talker localization methods in multiple sound source environments not only have difficulty localizing the multiple sound sources accurately, but also have difficulty localizing the target talker among known multiple sound source positions. To cope with these problems, we propose a new talker localization method consisting of two algorithms. One algorithm is for multiple sound source localization based on CSP (cross-power spectrum phase) analysis. The other algorithm is for sound source identification among localized multiple sound sources towards talker localization. We particularly focus on the latter statistical sound source identification among localized multiple sound sources with statistical speech and environmental sound models based on GMMs (Gaussian mixture models) and a microphone array towards talker localization. We especially evaluate the performance of the proposed algorithms with the RWCP sound scene database in real acoustic environments (RWCP-DB).