Analysis on individual differences in automatic transcription of spontaneous presentations

Analysis on individual differences in automatic transcription of spontaneous presentations
复制标题

自发陈述自动转录的个体差异分析

DOI:
10.1109/icassp.2002.5743821
复制
发表时间:
2002
期刊:
2002 IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
S. Furui
S. Furui
中科院分区:
--
文献类型:
--
作者:
T. Shinozaki;S. Furui

文献摘要

被引文献

相似文献

本文报告了自发呈现语音识别性能的个体差异的分析。50名男性发言者每次发言的10分钟,总共500分钟,已被自动识别用于分析。相关和回归分析被施加到单词识别的准确性和各种扬声器属性。一组有限的扬声器属性,包括说话率,词汇率和修复率被认为是最显着的,以产生个别差异的单词准确性。无监督MLLR说话人自适应对提高词汇准确率效果很好,但没有改变个体差异的结构。大约一半的变异在单词准确性解释了回归模型使用有限的三个属性。
This paper reports an analysis of individual differences in spontaneous presentation speech recognition performances. Ten minutes from each presentation given by 50 male speakers, for a total of 500 minutes, has been automatically recognized for the analysis. Correlation and regression analyses were applied to the word recognition accuracy and various speaker attributes. A restricted set of the speaker attributes comprising the speaking rate, the out of vocabulary rate and the repair rate was found to be most significant to yield individual differences in the word accuracy. Unsupervised MLLR speaker adaptation worked well for improving the word accuracy but did not change the structure of the individual differences. Approximately half of the variance in the word accuracy was explained by a regression model using the limited set of three attributes.