LANGUAGE MODEL ADAPTATION FOR BROADCAST NEWS TRANSCRIPTION

LANGUAGE MODEL ADAPTATION FOR BROADCAST NEWS TRANSCRIPTION
复制标题

广播新闻转录的语言模型自适应

DOI:
--
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
M. Adda
M. Adda
中科院分区:
--
文献类型:
--
作者:
Langzhou Chen;J. Gauvain;L. Lamel;G. Adda;M. Adda

文献摘要

被引文献

相似文献

本文报告了广播新闻转录任务的语言模型自适应。针对这一任务的语言模型适应具有挑战性,因为任何特定节目或其部分的主题事先都是未知的,并且通常与多个主题相关。语言模型自适应的问题之一是从音频信号中提取可靠的主题信息,特别是在存在识别错误​​的情况下。在这项工作中,我们利用信息检索中使用的技术从单词识别器中提取主题信息,然后使用这些信息从大型通用文本语料库中自动选择适应数据。已经使用适应数据研究了两种自适应语言模型,即基于混合的模型和基于 MAP 的模型。使用 LIMSI 普通话广播新闻转录系统进行的实验表明,通过结合两种适应方法,相对字符错误率降低了 4.3%。
This paper reports on language model adaptation for the broa dcast news transcription task. Language model adaptation fo r this task is challenging in that the subject of any particular sho w or portion thereof is unknown in advance and is often related to more than one topic. One of the problems in language model adaptation is the extraction of reliable topic information from the audio signal, particularly in the presence of recognition e rrors. In this work, we draw upon techniques used in information retri eval to extract topic information from the word recognizerhypot heses, which are then used to automatically select adaptation data from a large general text corpus. Two adaptive language models, a mixture-based model and a MAP-based model, have been investigated using the adaptation data. Experiments carried out with the LIMSI Mandarin broadcast news transcription system giv es a relative character error rate reduction of 4.3% by combinin g both adaptation methods.