The IBM 2006 speech transcription system for european parliamentary speeches
The IBM 2006 speech transcription system for european parliamentary speeches
复制标题
IBM 2006 年欧洲议会演讲语音转录系统
DOI:
10.21437/interspeech.2006-369
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
Alvaro Soneiro
中科院分区:
文献类型:
--
作者:
B. Ramabhadran;Olivier Siohan;L. Mangu;Geoffrey Zweig;Martin Westphal;Henrik Schulz;Alvaro Soneiro
TC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM’s English and Spanish speech recognition systems submitted to the TC-STAR 2006 Evaluation. The technical advances in this submission include two different algorithms for automatic segmentation and speaker clustering of the input audio; a system architecture that is based on cross-adaptation across these two segmentation schemes and system combination through generation of an ensemble of systems using randomized decision tree state-tying; automatic punctuation of the speech recognition output; and the incorporation of an additional 35 hours of in-domain EPPS acoustic training data. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2005 Evaluation on the 2006 English development test set, and produced one of the best performing systems on the 2006 evaluation in English with a word error rate of 8.3%. Index Terms: speech recognition, automatic segmentation, crossadaptation, randomized decision trees, TC-STAR.