Unsupervised submodular subset selection for speech data

Unsupervised submodular subset selection for speech data
复制标题

语音数据的无监督子模子集选择

DOI:
--
复制
发表时间:
2014
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
J. Bilmes
J. Bilmes
中科院分区:
--
文献类型:
--
作者:
K. Wei;Yuzong Liu;K. Kirchhoff;J. Bilmes

文献摘要

被引文献

相似文献

我们对选择培训电话识别器的声学数据子集进行了比较研究。将数据选择问题作为一个约束的次体优化问题解决。这种方法的先前应用需要以监督方式训练的转录或声学模型。在本文中,我们开发并评估了一种新颖且完全无监督的方法,并将其应用于Timit数据。结果表明,我们的方法始终优于许多基线方法,同时计算非常有效,并且不需要标记。
We conduct a comparative study on selecting subsets of acoustic data for training phone recognizers. The data selection problem is approached as a constrained submodular optimization problem. Previous applications of this approach required transcriptions or acoustic models trained in a supervised way. In this paper we develop and evaluate a novel and entirely unsupervised approach, and apply it to TIMIT data. Results show that our method consistently outperforms a number of baseline methods while being computationally very efficient and requiring no labeling.