An evaluation of sound source identification with RWCP sound scene database in real acoustic environments
An evaluation of sound source identification with RWCP sound scene database in real acoustic environments
复制标题
真实声环境下RWCP声音场景数据库声源识别评估
DOI:
10.1109/icme.2002.1035570
复制
发表时间:
2002
期刊:
影响因子:
--
通讯作者:
Satoshi Nakamura
中科院分区:
文献类型:
--
作者:
T. Nishiura;Satoshi Nakamura
It is very important for a hands-free speech interface to capture distant speech with high quality. A microphone array is an ideal candidate for this purpose. However, this approach requires localizing the target talker. Conventional talker localization methods in multiple sound source environments not only have difficulty localizing the multiple sound sources accurately, but also have difficulty localizing the target talker among known multiple sound source positions. To cope with these problems, we propose a new talker localization method consisting of two algorithms. One algorithm is for multiple sound source localization based on CSP (cross-power spectrum phase) analysis. The other algorithm is for sound source identification among localized multiple sound sources towards talker localization. We particularly focus on the latter statistical sound source identification among localized multiple sound sources with statistical speech and environmental sound models based on GMMs (Gaussian mixture models) and a microphone array towards talker localization. We especially evaluate the performance of the proposed algorithms with the RWCP sound scene database in real acoustic environments (RWCP-DB).