Multi-Channel VAD for Transcription of Group Discussion

Multi-Channel VAD for Transcription of Group Discussion
复制标题

DOI:
10.21437/interspeech.2021-200
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Osamu Ichikawa;Kaito Nakano;T. Nakayama;H. Shirouzu
Osamu Ichikawa;Kaito Nakano;T. Nakayama;H. Shirouzu
中科院分区:
其他
文献类型:
--
作者:
Osamu Ichikawa;Kaito Nakano;T. Nakayama;H. Shirouzu

文献摘要

相似文献

人们正在尝试通过给参加课堂小组作业的学生佩戴麦克风来可视化学习过程,随后使用自动语音识别(ASR)系统来可视化他们的讲话。然而,即使使用具有噪声鲁棒性的近距离麦克风,附近学生的声音也经常与输出语音数据混合。为了解决这一挑战,在本文中,我们建议使用多通道语音活动检测(VAD)来确定目标说话者的语音片段,同时还参考连接到组中其他说话者的麦克风的输出语音。利用中学生在小组作业课上的实际语音进行的评估实验表明,与传统技术单通道VAD(49.5%)相比,我们提出的方法显着提高了误帧率(38.7%)。我们认为,分布式麦克风阵列和深度学习等传统方法在某种程度上依赖于说话者位置的时间平稳性。然而,所提出的方法本质上是一个 VAD 过程,因此工作稳健。它是真实课堂环境中实用且经过验证的解决方案。
Attempts are being made to visualize the learning process by attaching microphones to students participating in group works conducted in classrooms, and subsequently, their speech using an automatic speech recognition (ASR) system. However, the voices of nearby students frequently become mixed with the output speech data, even when using close-talk microphones with noise robustness. To resolve this challenge, in this paper, we propose using multi-channel voice activity detection (VAD) to determine the speech segments of a target speaker while also referencing the output speech from the microphones attached to the other speakers in the group. The conducted evaluation experiments using the actual speech of middle school students during group work lessons showed that our proposed method significantly improves the frame error rate (38.7%) compared to that of the conventional technology, single-channel VAD (49.5%). In our view, conventional approaches, such as distributed microphone arrays and deep learning, are somewhat dependent on the temporal stationarity of the speakers’ positions. However, the proposed method is essentially a VAD process and thus works robustly. It is the practical and proven solution in a real classroom environment.