A system for automatic alignment of broadcast media captions using weighted finite-state transducers

A system for automatic alignment of broadcast media captions using weighted finite-state transducers
复制标题

DOI:
10.1109/asru.2015.7404861
复制
发表时间:
2015-12
期刊:
2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)
影响因子:
--
通讯作者:
P. Bell;S. Renals
P. Bell;S. Renals
中科院分区:
其他
文献类型:
--
作者:
P. Bell;S. Renals

文献摘要

相似文献

我们在2015年MGB挑战赛中描述了我们的广播媒体字幕对齐系统。先前生成的字幕与媒体数据的精确时间对准在广播公司的字幕生成过程中是重要的。然而,由于音频的高度多样化、经常有噪声的内容,并且因为字幕经常不是所讲的实际单词的逐字表示,因此该任务具有挑战性。我们的系统采用了一个两遍的方法与适当的约束加权有限状态传感器(WFST),使良好的对齐,即使当音频质量将是传统的ASR的挑战。该系统在MGB Challenge开发集上获得了0.8965的f分数。
We describe our system for alignment of broadcast media captions in the 2015 MGB Challenge. A precise time alignment of previously-generated subtitles to media data is important in the process of caption generation by broadcasters. However, this task is challenging due to the highly diverse, often noisy content of the audio, and because the subtitles are frequently not a verbatim representation of the actual words spoken. Our system employs a two-pass approach with appropriately constrained weighted finite state transducers (WFSTs) to enable good alignment even when the audio quality would be challenging for conventional ASR. The system achieves an f-score of 0.8965 on the MGB Challenge development set.