Predicting error rates for unknown data in automatic speech recognition

Predicting error rates for unknown data in automatic speech recognition
复制标题

DOI:
10.1109/icassp.2017.7953174
复制
发表时间:
2017-03
期刊:
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
B. Meyer;Sri Harish Reddy Mallidi;H. Kayser;H. Hermansky
B. Meyer;Sri Harish Reddy Mallidi;H. Kayser;H. Hermansky
中科院分区:
其他
文献类型:
--
作者:
B. Meyer;Sri Harish Reddy Mallidi;H. Kayser;H. Hermansky

文献摘要

被引文献

相似文献

在本文中,我们研究了在存在未知噪声类型(训练过程中未曾见过)的情况下预测自动语音识别中的单词错误率的方法。性能测量对从神经网络获得的音素后验图进行操作。我们将平均帧熵作为基线测量与平均时间距离(M-Measure)和语音事件的数量进行比较。后者是通过从干净的训练数据中学习典型的音素激活来获得的,随后将其用作后验图(MaP)的特定于音素的匹配过滤器。当过滤后超过阈值时,我们将其注册为语音事件。对于使用 10 种未知噪声类型和各种信噪比的测试集,我们发现 M-Measure 和 MaP 产生的预测准确度是基线测量的两倍。当排除包含语音片段的噪声类型时,预测误差为 3.1%,而基线测量的预测误差为 15.0%。
In this paper we investigate methods to predict word error rates in automatic speech recognition in the presence of unknown noise types, which have not been seen during training. The performance measures operate on phoneme posteriorgrams that are obtained from neural nets. We compare average frame-wise entropy as a baseline measure to the mean temporal distance (M-Measure) and to the number of phonetic events. The latter is obtained by learning typical phoneme activations from clean training data, which are later applied as phoneme-specific matched filters to posteriorgrams (MaP). When exceeding a threshold after filtering, we register this as phonetic event. For test sets using 10 unknown noise types and a wide range of signal-to-noise ratios, we find M-Measure and MaP to produce predictions twice as accurate as the baseline measure. When excluding noise types that contain speech segments, a prediction error of 3.1% is achieved, compared to 15.0% for the baseline measure.