Towards Spontaneous Speech Translation

Towards Spontaneous Speech Translation
复制标题

走向自发语音翻译

DOI:
10.1016/j.procs.2016.04.032
复制
发表时间:
1994
影响因子:
4.7
通讯作者:
A. Waibel
A. Waibel
中科院分区:
医学3区
文献类型:
--
作者:
M. Woszczyna;N. Aoki;Finn Dag Buø;N. Coccaro;Keiko Horiguchi;T. Kemp;A. Lavie;A. McNair;T. Polzin;I. Rogina;C. Rosé;T. Schultz;B. Suhm;M. Tomita;A. Waibel

文献摘要

被引文献

相似文献

在这项工作中,我们使用无监督线性判别分析(LDA)来支持零资源场景下的声学单元发现。其思想是自动找到特征向量到子空间的映射,该映射更适合于基于DPGMM(DPGMM)的聚类,而不需要监督。有监督声学建模通常利用诸如LDA之类的特征变换来最小化类内可区分性、最大化类间可区分性并从跨越较大上下文的高维特征中提取相关信息。由于需要类标签,因此很难在类甚至数量未知的零资源环境中使用该技术。为了解决这个问题,我们在标准特征上使用DPGMM聚类的第一次迭代来生成数据的标签,作为学习适当变换的基础。第二聚类对变换后的特征进行操作。在给定无监督数据的情况下,无监督LDA的应用明显地导致了更好的聚类结果。我们表明,改进的输入特征始终优于我们的基准输入特征。
In this work we make use of unsupervised linear discriminant analysis (LDA) to support acoustic unit discovery in a zero resource scenario. The idea is to automatically find a mapping of feature vectors into a subspace that is more suitable for Dirichlet process Gaussian mixture model (DPGMM) based clustering, without the need of supervision. Supervised acoustic modeling typically makes use of feature transformations such as LDA to minimize intra-class discriminability, to maximize inter-class discriminability and to extract relevant informations from high-dimensional features spanning larger contexts. The need of class labels makes it difficult to use this technique in a zero resource setting where the classes and even their amount are unknown. To overcome this issue we use a first iteration of DPGMM clustering on standard features to generate labels for the data, that serve as basis for learning a proper transformation. A second clustering operates on the transformed features. The application of unsupervised LDA demonstrably leads to better clustering results given the unsupervised data. We show that the improved input features consistently outperform our baseline input features.