An analysis of environment, microphone and data simulation mismatches in robust speech recognition

An analysis of environment, microphone and data simulation mismatches in robust speech recognition
复制标题

DOI:
10.1016/j.csl.2016.11.005
复制
发表时间:
2017-11-01
影响因子:
4.3
通讯作者:
Marxer, Ricard
Marxer, Ricard
中科院分区:
计算机科学3区
文献类型:
--
作者:
Vincent, Emmanuel;Watanabe, Shinji;Marxer, Ricard

文献摘要

被引文献

相似文献

语音增强和自动语音识别 (ASR) 最常在匹配(或多条件)设置中进行评估,其中训练数据的声学条件与测试数据的声学条件匹配(或覆盖)。很少有研究系统地评估训练数据和测试数据之间声学不匹配的影响,特别是关于最近的语音增强和最先进的 ASR 技术。在本文中,我们在 CHiME-3 数据集的背景下研究这个问题,该数据集包含使用基于 6 通道平板电脑的麦克风阵列记录的处于挑战性嘈杂环境中的说话者所说的句子。我们对该数据集上发布的各种信号增强、特征提取和 ASR 后端技术的结果进行了批判性分析,并进行了许多新实验,以便分别评估不同噪声环境、不同数量和位置的麦克风或模拟数据与真实数据对语音增强和 ASR 性能的影响。我们表明,除了最小方差无失真响应(MVDR)波束形成之外,大多数算法在真实数据和模拟数据上的表现一致,并且可以从模拟数据的训练中受益。我们还发现,在不同噪声环境和不同麦克风上进行训练几乎不会影响 ASR 性能,特别是当训练数据中存在多种环境时:只有麦克风的数量有显着影响。基于这些结果,我们引入了 CHiME-4 语音分离和识别挑战赛,它重新审视了 CHiME-3 数据集,并通过减少可用于测试的麦克风数量使其更具挑战性。 (C) 2016 Elsevier Ltd. 保留所有权利。
Speech enhancement and automatic speech recognition (ASR) are most often evaluated in matched (or multi-condition) settings where the acoustic conditions of the training data match (or cover) those of the test data. Few studies have systematically assessed the impact of acoustic mismatches between training and test data, especially concerning recent speech enhancement and state-of-the-art ASR techniques. In this article, we study this issue in the context of the CHiME-3 dataset, which consists of sentences spoken by talkers situated in challenging noisy environments recorded using a 6-channel tablet based microphone array. We provide a critical analysis of the results published on this dataset for various signal enhancement, feature extraction, and ASR backend techniques and perform a number of new experiments in order to separately assess the impact of different noise environments, different numbers and positions of microphones, or simulated vs. real data on speech enhancement and ASR performance. We show that, with the exception of minimum variance distortionless response (MVDR) beamforming, most algorithms perform consistently on real and simulated data and can benefit from training on simulated data. We also find that training on different noise environments and different microphones barely affects the ASR performance, especially when several environments are present in the training data: only the number of microphones has a significant impact. Based on these results, we introduce the CHiME-4 Speech Separation and Recognition Challenge, which revisits the CHiME-3 dataset and makes it more challenging by reducing the number of microphones available for testing. (C) 2016 Elsevier Ltd. All rights reserved.