Directly Comparing the Listening Strategies of Humans and Machines

Directly Comparing the Listening Strategies of Humans and Machines
复制标题

DOI:
10.1109/taslp.2020.3040545
复制
发表时间:
2016-09
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Michael I. Mandel
Michael I. Mandel
中科院分区:
其他
文献类型:
--
作者:
Michael I. Mandel

文献摘要

被引文献

相似文献

自动语音识别(ASR)在许多干净的语料库上已经达到了人类的水平,但在嘈杂的环境中仍然不如人类听众。本文研究了这种表现差异是否可能是由于每个听众在做出决定时所利用的时频区域的差异,以及使用不同声学模型(AMs)和语言模型(lm)的asr如何改变这些“重要”区域。我们将重要区域定义为频谱图中的时间-频率点,当听者正确识别噪声中的话语时,这些点往往是可听的。本研究的证据表明,与传统的高斯混合模型(GMM) AM相比,神经网络AM关注的区域更类似于人类的区域(捕获某些高能量区域)。我们的分析还表明,神经网络AM还没有捕捉到人类听众使用的所有线索,例如沉默和高言语能量之间的某些转换。我们还发现,重要时频区域的差异往往会追踪测试句子中特定单词的准确性差异,这表明存在联系。由于这种联系,调整ASR以关注人类使用的相同区域可能会提高其对噪声的泛化。
Automatic speech recognition (ASR) has reached human performance on many clean speech corpora, but it remains worse than human listeners in noisy environments. This paper investigates whether this difference in performance might be due to a difference in the time-frequency regions that each listener utilizes in making their decisions and how these “important” regions change for ASRs using different acoustic models (AMs) and language models (LMs). We define important regions as time-frequency points in a spectrogram that tend to be audible when the listener correctly recognizes that utterance in noise. The evidence from this study indicates that a neural network AM attends to regions that are more similar to those of humans (capturing certain high-energy regions) than those of a traditional Gaussian mixture model (GMM) AM. Our analysis also shows that the neural network AM has not yet captured all the cues that human listeners utilize, such as certain transitions between silence and high speech energy. We also find that differences in important time-frequency regions tend to track differences in accuracy on specific words in a test sentence, suggesting a connection. Because of this connection, adapting an ASR to attend to the same regions humans use might improve its generalization in noise.