Speaker indexing and speech enhancement in real meetings / conversations

Speaker indexing and speech enhancement in real meetings / conversations
复制标题

真实会议/对话中的发言者索引和语音增强

DOI:
10.1109/icassp.2008.4517554
复制
发表时间:
2008
期刊:
2008 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
S. Makino
S. Makino
中科院分区:
--
文献类型:
--
作者:
S. Araki;Masakiyo Fujimoto;K. Ishizuka;H. Sawada;S. Makino

文献摘要

被引文献

相似文献

本文提出了一种说话人索引方法,使用少量的麦克风来估计谁说话时。我们提出的说话人索引是通过使用噪声鲁棒的语音活动检测器(VAD),QCC-PHAT的到达方向(DOA)估计器,和DOA分类器实现的。使用估计的说话人索引信息,我们还可以用最大信噪比(MaxSNR)波束形成器增强每个说话人的话语。本文将我们的系统应用于真实的记录会议/谈话记录在一个房间中的混响时间为350毫秒,并评估性能的标准措施:日志错误率(DER)。即使对于真实的会话,其中有许多发言人轮流和重叠,说话人错误的时间是非常小的与我们提出的系统。我们计划在ICASSP 2008上展示一个实时的说话人索引系统。
This paper presents a speaker indexing method that uses a small number of microphones to estimate who spoke when. Our proposed speaker indexing is realized by using a noise robust voice activity detector (VAD), a QCC-PHAT based direction of arrival (DOA) estimator, and a DOA classifier. Using the estimated speaker indexing information, we can also enhance the utterances of each speaker with a maximum signal-to-noise-ratio (MaxSNR) beamformer. This paper applies our system to real recorded meetings / conversations recorded in a room with a reverberation time of 350 ms, and evaluates the performance by a standard measure: the diarization error rate (DER). Even for the real conversations, which have many speaker turn-takings and overlaps, the speaker error time was very small with our proposed system. We are planning to demonstrate a real-time speaker indexing system at ICASSP2008.