Unsupervised optimal phoneme segmentation: Objectives, algorithm and comparisons

Unsupervised optimal phoneme segmentation: Objectives, algorithm and comparisons
复制标题

DOI:
10.1109/icassp.2008.4518528
复制
发表时间:
2008-05
期刊:
2008 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Y. Qiao;Naoya Shimomura;N. Minematsu
Y. Qiao;Naoya Shimomura;N. Minematsu
中科院分区:
其他
文献类型:
--
作者:
Y. Qiao;Naoya Shimomura;N. Minematsu

文献摘要

被引文献

相似文献

音素分割是许多语音识别和合成研究中的一个基本问题。无监督音素分割假设不了解语言内容和声学模型,因此提出了一个具有挑战性的问题。这里的基本问题是什么是最佳分割。本文将最优分割问题表述为概率框架。利用统计学和信息论分析,我们开发了三种不同的目标函数,即误差平方和(SSE)、对数行列式(LD)和率失真(RD)。特别地,RD函数源自信息率失真理论,可以与人类信号感知机制相关。我们引入了时间约束的凝聚聚类算法来找到最佳分割。我们还提出了一种使用积分函数来实现该算法的有效方法。我们在TIMIT数据库上进行实验来比较上述三个目标函数。结果表明,率失真实现了最佳性能,并表明我们的方法优于最近发布的无监督分割方法。
Phoneme segmentation is a fundamental problem in many speech recognition and synthesis studies. Unsupervised phoneme segmentation assumes no knowledge on linguistic contents and acoustic models, and thus poses a challenging problem. The essential question here is what is the optimal segmentation. This paper formulates the optimal segmentation problem into a probabilistic framework. Using statistics and information theory analysis, we develop three different objective functions, namely, summation of square error (SSE), log determinant (LD) and rate distortion (RD). Specially, RD function is derived from information rate distortion theory and can be related to human signal perception mechanism. We introduce a time-constrained agglomerative clustering algorithm to find the optimal segmentations. We also propose an efficient method to implement the algorithm by using integration functions. We carry out experiments on TIMIT database to compare the above three objective functions. The results show that rate distortion achieves the best performance and indicate that our method outperforms the recently published unsupervised segmentation methods.