Impact of audio segmentation and segment clustering on automated transcription accuracy of large spoken archives
Impact of audio segmentation and segment clustering on automated transcription accuracy of large spoken archives
复制标题
音频分割和片段聚类对大型语音档案自动转录准确性的影响
DOI:
10.21437/eurospeech.2003-714
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
H. Nock
中科院分区:
文献类型:
--
作者:
B. Ramabhadran;Jing Huang;U. Chaudhari;G. Iyengar;H. Nock
This paper addresses the influence of audio segmentation and segment clustering on automatic transcription accuracy for large spoken archives. The work formspart of the ongoing MALACH project, which is developing advanced techniques for supporting access to the world’s largest digital archive of video oral histories collected in many languages from over 52000 survivors and witnesses of the Holocaust. We present several audio-only and audio-visual segmentation schemes, including two novel schemes: the first is iterative and audio-only, the second uses audio-visual synchrony. Unlike most previous work, we evaluate these schemes in terms of their impact upon recognition accuracy. Results on English interviews show the automatic segmentation schemes give performance comparable to (exhorbitantly expensive and impractically lengthy) manual segmentation when using a single pass decoding strategy based on speaker-independent models. However, when using a multiple pass decoding strategy with adaptation, results are sensitive to both initial audio segmentation and the scheme for clustering segments prior to adaptation: the combination of our best automatic segmentation and clustering scheme has an error rate 8% worse (relative) to manual audio segmentation and clustering due to the occurrence of “speaker-impure” segments.