The IBM 2006 speech transcription system for european parliamentary speeches

The IBM 2006 speech transcription system for european parliamentary speeches
复制标题

IBM 2006 年欧洲议会演讲语音转录系统

DOI:
10.21437/interspeech.2006-369
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
Alvaro Soneiro
Alvaro Soneiro
中科院分区:
--
文献类型:
--
作者:
B. Ramabhadran;Olivier Siohan;L. Mangu;Geoffrey Zweig;Martin Westphal;Henrik Schulz;Alvaro Soneiro

文献摘要

被引文献

相似文献

TC-STAR是欧盟资助的语音翻译项目,用于转录,翻译和合成欧洲议会全体会议演讲(EPPS)。本文介绍了IBM的英语和西班牙语语音识别系统提交的TC-STAR 2006评估。该提交的技术进步包括用于输入音频的自动分段和说话者聚类的两种不同算法;基于跨这两种分段方案的交叉适应的系统架构和通过使用随机决策树状态绑定生成系统集合的系统组合;语音识别输出的自动标点;以及结合额外的35小时的域内EPPS声学训练数据。这些进步使2006年英语开发测试集的TC-STAR 2005评估中表现最好的系统的错误率相对降低了30%,并成为2006年英语评估中表现最好的系统之一,单词错误率为8.3%。索引术语:语音识别,自动分割,交叉适应,随机决策树,TC-STAR。
TC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM’s English and Spanish speech recognition systems submitted to the TC-STAR 2006 Evaluation. The technical advances in this submission include two different algorithms for automatic segmentation and speaker clustering of the input audio; a system architecture that is based on cross-adaptation across these two segmentation schemes and system combination through generation of an ensemble of systems using randomized decision tree state-tying; automatic punctuation of the speech recognition output; and the incorporation of an additional 35 hours of in-domain EPPS acoustic training data. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2005 Evaluation on the 2006 English development test set, and produced one of the best performing systems on the 2006 evaluation in English with a word error rate of 8.3%. Index Terms: speech recognition, automatic segmentation, crossadaptation, randomized decision trees, TC-STAR.