Multiple speaker tracking using a microphone array by combining auditory processing and a gaussian mixture cardinalized probability hypothesis density filter

Multiple speaker tracking using a microphone array by combining auditory processing and a gaussian mixture cardinalized probability hypothesis density filter
复制标题

通过结合听觉处理和高斯混合基数概率假设密度滤波器,使用麦克风阵列进行多说话人跟踪

DOI:
--
复制
发表时间:
2011
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
G. Fink
G. Fink
中科院分区:
--
文献类型:
--
作者:
A. Plinge;D. Hauschildt;Marius H. Hennecke;G. Fink

文献摘要

被引文献

相似文献

跟踪说话人是智能环境中的一个重要应用。由于两个主要原因,使用麦克风阵列的声学跟踪是一项具有挑战性的任务:一方面,多个人可能同时发言,因此说话者的数量随着时间的推移而变化;另一方面,由于混响语音的性质,所提供的位置假设包含许多间隙和混乱。在该方法中,“瞥见模型”是通过神经生物学中稳健但稀疏的位置假设的螺旋计算结合高斯混合基数概率假设密度过滤器来实现的。通过对来自多个频段的位置假设迭代应用该滤波器,获得了良好的结果。使用从录音中得出的统计语音模型,使用能够实时执行的实现来跟踪具有显著混响的会议室中的多个发言者。
Tracking speakers is an important application in smart environments. Acoustic tracking using microphone arrays is a challenging task due to two major reasons: On the one hand, multiple persons may speak simultaneously and thus the number of speakers varies over time; on the other hand, due to the nature of reverberated speech, the provided position hypotheses contain many gaps and clutter. In the proposed approach, the "glimpsing model" is realized by neurobiologically in spired calculation of robust but sparse position hypotheses in combination with a Gaussian mixture cardinalized probability hypothesis density filter. By iteratively applying the filter to the position hypotheses from multiple frequency bands, good results are achieved. Using a statistical speech model derived from recordings, a real-time capable implementation is used to track multiple speakers in a conference room with significant reverberation.