A novel method of language modeling for automatic captioning in TC video teleconferencing.

A novel method of language modeling for automatic captioning in TC video teleconferencing.
复制标题

TC 视频电话会议中自动字幕的语言建模新方法。

DOI:
10.1109/titb.2006.885549
复制
发表时间:
2007
期刊:
IEEE transactions on information technology in biomedicine : a publication of the IEEE Engineering in Medicine and Biology Society
影响因子:
--
通讯作者:
Schopp,Laura
Schopp,Laura
中科院分区:
--
文献类型:
--
作者:
Zhang,Xiaojia;Zhao,Yunxin;Schopp,Laura

文献摘要

相似文献

我们正在开发一种基于大词汇会话语音识别的远程医疗远程会诊视频会议(TC-VTC)自动字幕系统。在TC-VTC中,医生的言语中包含了大量不常用的医学术语,且风格自然。由于数据不足,我们采用混合语言建模,模型由多个医疗和非医疗领域的数据集训练而成。本文针对混合语言模型(LM)提出了一种新的建模和估计方法。组件LM从单个数据集训练,类n-gram LM从域内数据集训练,词n-gram LM从域外数据集训练,它们被插值到混合LM中。对于类lm,语义类别用于医学术语、名称和数字的类定义。采用前向权值调整贪心算法估计混合LM的插值权值。本文提出的混合域内类lm和域外词lm、类的语义定义以及FWA的权重估计算法在TC-VTC任务上是有效的。与使用传统期望最大化算法估计权值的词LMs混合方法相比,所提出的方法使五位医生的测试集的困惑度降低了21%,这转化为字幕准确性的提高
We are developing an automatic captioning system for teleconsultation video teleconferencing (TC-VTC) in telemedicine, based on large vocabulary conversational speech recognition. In TC-VTC, doctors' speech contains a large number of infrequently used medical terms in spontaneous styles. Due to insufficiency of data, we adopted mixture language modeling, with models trained from several datasets of medical and nonmedical domains. This paper proposes novel modeling and estimation methods for the mixture language model (LM). Component LMs are trained from individual datasets, with class n-gram LMs trained from in-domain datasets and word n-gram LMs trained from out-of-domain datasets, and they are interpolated into a mixture LM. For class LMs, semantic categories are used for class definition on medical terms, names, and digits. The interpolation weights of a mixture LM are estimated by a greedy algorithm of forward weight adjustment (FWA). The proposed mixing of in-domain class LMs and out-of-domain word LMs, the semantic definitions of word classes, as well as the weight-estimation algorithm of FWA are effective on the TC-VTC task. As compared with using mixtures of word LMs with weights estimated by the conventional expectation-maximization algorithm, the proposed methods led to a 21% reduction of perplexity on test sets of five doctors, which translated into improvements of captioning accuracy