Podcastle: collaborative training of acoustic models on the basis of wisdom of crowds for podcast transcription

Podcastle: collaborative training of acoustic models on the basis of wisdom of crowds for podcast transcription
复制标题

Podcastle:基于群体智慧的声学模型协作训练,用于播客转录

DOI:
--
复制
发表时间:
2009
期刊:
Interspeech
影响因子:
--
通讯作者:
Masataka Goto
Masataka Goto
中科院分区:
--
文献类型:
--
作者:
J. Ogata;Masataka Goto

文献摘要

被引文献

相似文献

本文提出了提高播客自动转录的声学模型训练技术。声学建模的一个典型方法是创建一个特定任务的语料库,其中包括数百(甚至数千)小时的语音数据及其准确的转录。然而,这种方法在播客转录任务中是不切实际的,因为手动生成涵盖所有不同类型播客内容的大量语音转录将过于昂贵和耗时。为了解决这个问题,我们引入了基于群体智慧的声学模型的协同训练,即播客语音数据的转录是由我们的网络服务PodCastle上的匿名用户生成的。然后,我们利用RSS元数据描述了一个播客相关的声学建模系统,以处理播客语音数据中声学条件的差异。通过对实际播客语音数据的实验结果,验证了所提声学模型训练的有效性。
This paper presents acoustic-model-training techniques for improving automatic transcription of podcasts. A typical approach for acoustic modeling is to create a task-specific corpus including hundreds (or even thousands) of hours of speech data and their accurate transcriptions. This approach, however, is impractical in podcast-transcription task because manual generation of the transcriptions of the large amounts of speech covering all the various types of podcast contents will be too costly and time consuming. To solve this problem, we introduce collaborative training of acoustic models on the basis of wisdom of crowds, i.e., the transcriptions of podcast-speech data are generated by anonymous users on our web service PodCastle. We then describe a podcast-dependent acoustic modeling system by using RSS metadata to deal with the differences of acoustic conditions in podcast speech data. From our experimental results on actual podcast speech data, the effectiveness of the proposed acoustic model training was confirmed.