The third 'CHIME' speech separation and recognition challenge: Analysis and outcomes

The third 'CHIME' speech separation and recognition challenge: Analysis and outcomes
复制标题

DOI:
10.1016/j.csl.2016.10.005
复制
发表时间:
2017-11-01
影响因子:
4.3
通讯作者:
Watanabe, Shinji
Watanabe, Shinji
中科院分区:
计算机科学3区
文献类型:
--
作者:
Barker, Jon;Marxer, Ricard;Watanabe, Shinji

文献摘要

被引文献

相似文献

本文介绍了CHIME-3挑战的设计和结果,这是第一个开放的语音识别评估,旨在针对日益相关的多通道、移动设备语音识别场景。这篇论文有两个目的。首先,它为这项挑战提供了明确的参考,包括对任务设计、数据采集和基线系统的全面说明,以及对提交的26个系统的说明和评价。最好的系统对基准的每个阶段进行了重新设计,导致字错误率从33.4%降至5.8%。通过跨系统进行比较,确定了实现强大性能所必需的技术。其次,本文考虑了从评估中得出结论的问题,这些评估使用的是在噪声环境中直接记录的语音。由此产生的材料带来的挑战程度很难控制,也很难完全描述。我们试图通过在每个会话和每个话语的基础上将各种估计的信号属性与典型的系统性能相关联来剖析各种“困难轴”。我们发现了依赖于信噪比和信道质量的强有力的证据。系统对扬声器运动程度的变化不那么敏感。本文最后讨论了CHME-3的结果与未来移动语音识别评估的设计相关的问题。(C)2016爱思唯尔有限公司。保留所有权利。
This paper presents the design and outcomes of the CHiME-3 challenge, the first open speech recognition evaluation designed to target the increasingly relevant multichannel, mobile-device speech recognition scenario. The paper serves two purposes. First, it provides a definitive reference for the challenge, including full descriptions of the task design, data capture and baseline systems along with a description and evaluation of the 26 systems that were submitted. The best systems re-engineered every stage of the baseline resulting in reductions in word error rate from 33.4% to as low as 5.8%. By comparing across systems, techniques that are essential for strong performance are identified. Second, the paper considers the problem of drawing conclusions from evaluations that use speech directly recorded in noisy environments. The degree of challenge presented by the resulting material is hard to control and hard to fully characterise. We attempt to dissect the various 'axes of difficulty' by correlating various estimated signal properties with typical system performance on a per session and per utterance basis. We find strong evidence of a dependence on signal-to-noise ratio and channel quality. Systems are less sensitive to variations in the degree of speaker motion. The paper concludes by discussing the outcomes of CHiME-3 in relation to the design of future mobile speech recognition evaluations. (C) 2016 Elsevier Ltd. All rights reserved.