Web-based language modelling for automatic lecture transcription
Web-based language modelling for automatic lecture transcription
复制标题
基于网络的自动讲座转录语言模型
DOI:
10.21437/interspeech.2007-266
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
R. Baecker
中科院分区:
文献类型:
--
作者:
Cosmin Munteanu;Gerald Penn;R. Baecker
Universities have long relied on written text to share knowledge. As more lectures are made available on-line, these must be accompanied by textual transcripts in order to provide the same access to information as textbooks. While Automatic Speech Recognition (ASR) is a cost-effective method to deliver transcriptions, its accuracy for lectures is not yet satisfactory. One approach for improving lecture ASR is to build smaller, topic-dependent Language Models (LMs) and combine them (through LM interpolation or hypothesis space combination) with general-purpose, large-vocabulary LMs. In this paper, we propose a simple solution for lecture ASR with similar or better Word Error Rate reductions (as well as topic-specific keyword identification accuracies) than combination-based approaches. Our method eliminates the need for two types of LMs by exploiting the lecture slides to collect a web corpus appropriate for modelling both the conversational and the topic-specific styles of lectures.