Progress in the CU-HTK broadcast news transcription system

Progress in the CU-HTK broadcast news transcription system
复制标题

CU-HTK广播新闻转录系统研究进展

DOI:
10.1109/tasl.2006.878264
复制
发表时间:
2006
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
S. Tranter
S. Tranter
中科院分区:
--
文献类型:
--
作者:
M. Gales;Do Yeong Kim;P. Woodland;R. Chan;D. Mrva;R. Sinha;S. Tranter

文献摘要

被引文献

相似文献

多年来,广播新闻转录一直是一个具有挑战性的研究领域。在过去的几年中,大量粗略转录的声学训练数据和先进的模型训练技术的可用为极大地降低这项任务的错误率提供了机会。本文描述了利用这些发展的BN转录系统的设计和性能。首先,讨论了使用低监督训练数据和先进的声学建模技术的效果。然后利用这些新模型详细介绍了一个实时广播新闻识别系统的设计。由于已经发现系统组合在性能上产生了很大的收益,接下来将描述允许组合多个识别输出的一系列框架。这包括使用多种类型的声学模型和多个分段。作为对比,还描述了由多个站点开发的允许跨站点组合的系统--“SuperEARS”系统。使用几个最新的BN开发和评估测试集来评估各种模型和识别配置。这些新的BN转录系统与CU-HTK 2003 BN系统相比,可以获得超过25%的收益
Broadcast news (BN) transcription has been a challenging research area for many years. In the last couple of years, the availability of large amounts of roughly transcribed acoustic training data and advanced model training techniques has offered the opportunity to greatly reduce the error rate on this task. This paper describes the design and performance of BN transcription systems which make use of these developments. First, the effects of using lightly supervised training data and advanced acoustic modeling techniques are discussed. The design of a real-time broadcast news recognition system is then detailed using these new models. As system combination has been found to yield large gains in performance, a range of frameworks that allow multiple recognition outputs to be combined are next described. These include the use of multiple types of acoustic models and multiple segmentations. As a contrast a system developed by multiple sites allowing cross-site combination, the "SuperEARS" system, is also described. The various models and recognition configurations are evaluated using several recent BN development and evaluation test sets. These new BN transcription systems can give gains of over 25% relative to the CU-HTK 2003 BN system