Predicting error rates for unknown data in automatic speech recognition
Predicting error rates for unknown data in automatic speech recognition
复制标题
DOI:
10.1109/icassp.2017.7953174
复制
发表时间:
2017-03
期刊:
影响因子:
--
通讯作者:
B. Meyer;Sri Harish Reddy Mallidi;H. Kayser;H. Hermansky
中科院分区:
文献类型:
--
作者:
B. Meyer;Sri Harish Reddy Mallidi;H. Kayser;H. Hermansky
In this paper we investigate methods to predict word error rates in automatic speech recognition in the presence of unknown noise types, which have not been seen during training. The performance measures operate on phoneme posteriorgrams that are obtained from neural nets. We compare average frame-wise entropy as a baseline measure to the mean temporal distance (M-Measure) and to the number of phonetic events. The latter is obtained by learning typical phoneme activations from clean training data, which are later applied as phoneme-specific matched filters to posteriorgrams (MaP). When exceeding a threshold after filtering, we register this as phonetic event. For test sets using 10 unknown noise types and a wide range of signal-to-noise ratios, we find M-Measure and MaP to produce predictions twice as accurate as the baseline measure. When excluding noise types that contain speech segments, a prediction error of 3.1% is achieved, compared to 15.0% for the baseline measure.