Towards automatic transcription of spontaneous presentations

Towards automatic transcription of spontaneous presentations
复制标题

实现自发演示的自动转录

DOI:
10.21437/eurospeech.2001-129
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
S. Furui
S. Furui
中科院分区:
--
文献类型:
--
作者:
T. Shinozaki;Chiori Hori;S. Furui

文献摘要

被引文献

相似文献

本文报道了与 1999 年启动的“自发演讲”国家项目相关的识别自发演讲演讲的各种调查。由 10 名男性演讲者发表的、持续约 4.5 小时的演讲演讲已被识别。实验结果表明,基于实际自发语音语料库的声学和语言建模比基于朗读语音的传统建模更有效。根据语速、填充词数量、修复次数等,识别准确度在说话人与说话人之间存在很大差异。已证实声学模型的无监督说话人自适应可有效提高识别准确度。然而,自发语音的识别准确率仍然较低,仍然存在大量的研究问题。
This paper reports various investigations on recognizing spontaneous presentation speech in connection with the “Spontaneous Speech” national project started in 1999. Presentation speech uttered by 10 male speakers of approximately 4.5 hours duration has been recognized. Experimental results show that acoustic and language modeling based on an actual spontaneous speech corpus is far more effective than conventional modeling based on read speech. The recognition accuracy has a wide speaker-tospeaker variability according to the speaking rate, the number of fillers, the number of repairs, etc. It was confirmed that unsupervised speaker adaptation of acoustic models was effective to improve the recognition accuracy. The recognition accuracy for spontaneous speech is, however, still rather low, and there remains a large number of research issues.