Impact of audio segmentation and segment clustering on automated transcription accuracy of large spoken archives

Impact of audio segmentation and segment clustering on automated transcription accuracy of large spoken archives
复制标题

音频分割和片段聚类对大型语音档案自动转录准确性的影响

DOI:
10.21437/eurospeech.2003-714
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
H. Nock
H. Nock
中科院分区:
--
文献类型:
--
作者:
B. Ramabhadran;Jing Huang;U. Chaudhari;G. Iyengar;H. Nock

文献摘要

被引文献

相似文献

本文讨论了音频分割和段聚类对大型口语档案自动转录精度的影响。这项工作是正在进行的MALACH项目的一部分,该项目正在开发先进的技术,以支持访问世界上最大的视频口述历史数字档案,这些视频口述历史以多种语言收集自52000多名大屠杀幸存者和证人。我们提出了几种纯音频和视听分割方案,其中包括两个新的方案:第一个是迭代和纯音频,第二个使用视听同步。与大多数以前的工作,我们评估这些计划的识别精度的影响。英语访谈的结果表明,自动分割方案的性能相比(令人发指的昂贵和不切实际的冗长)手动分割时,使用一个单一的通过解码策略的基础上说话人独立的模型。然而,当使用具有自适应的多通道解码策略时,结果对初始音频分段和用于在自适应之前对分段进行聚类的方案都敏感:由于“说话者不纯”分段的发生,我们的最佳自动分段和聚类方案的组合具有比手动音频分段和聚类差8%(相对于)的错误率。
This paper addresses the influence of audio segmentation and segment clustering on automatic transcription accuracy for large spoken archives. The work formspart of the ongoing MALACH project, which is developing advanced techniques for supporting access to the world’s largest digital archive of video oral histories collected in many languages from over 52000 survivors and witnesses of the Holocaust. We present several audio-only and audio-visual segmentation schemes, including two novel schemes: the first is iterative and audio-only, the second uses audio-visual synchrony. Unlike most previous work, we evaluate these schemes in terms of their impact upon recognition accuracy. Results on English interviews show the automatic segmentation schemes give performance comparable to (exhorbitantly expensive and impractically lengthy) manual segmentation when using a single pass decoding strategy based on speaker-independent models. However, when using a multiple pass decoding strategy with adaptation, results are sensitive to both initial audio segmentation and the scheme for clustering segments prior to adaptation: the combination of our best automatic segmentation and clustering scheme has an error rate 8% worse (relative) to manual audio segmentation and clustering due to the occurrence of “speaker-impure” segments.