Convex weighting criteria for speaking rate estimation.

Convex weighting criteria for speaking rate estimation.
复制标题

DOI:
10.1109/taslp.2015.2434213
复制
发表时间:
2015-09
期刊:
IEEE/ACM transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Liss J
Liss J
中科院分区:
其他
文献类型:
--
作者:
Jiao Y;Berisha V;Tu M;Liss J

文献摘要

被引文献

相似文献

直接从语音波形中估计语速是语音信号处理中的一个长期问题。在本文中,我们提出的说话率估计问题,估计时间密度函数的积分在一个给定的时间间隔内的说话率。与许多现有的方法相比,我们避免了更困难的任务,检测语音信号中的单个音素,我们避免了诸如阈值化时间包络来估计元音的数量的算法。相反,所提出的方法的目的是学习一个最佳的加权函数,可以直接应用于语音信号中的时间-频率特征,以产生时间密度函数。我们提出了两个凸成本函数的学习加权函数和自适应策略,以定制的方法,以一个特定的扬声器使用最少的训练。TIMIT语料库,构音障碍的语音语料库,和ICSI开关板自发语音语料库上的算法进行评估。结果表明,所提出的方法优于三个竞争的方法对健康和构音障碍的语音。此外,对于自发语音速率估计,结果显示估计的说话速率和地面真值之间的高度相关性。
Speaking rate estimation directly from the speech waveform is a long-standing problem in speech signal processing. In this paper, we pose the speaking rate estimation problem as that of estimating a temporal density function whose integral over a given interval yields the speaking rate within that interval. In contrast to many existing methods, we avoid the more difficult task of detecting individual phonemes within the speech signal and we avoid heuristics such as thresholding the temporal envelope to estimate the number of vowels. Rather, the proposed method aims to learn an optimal weighting function that can be directly applied to time-frequency features in a speech signal to yield a temporal density function. We propose two convex cost functions for learning the weighting functions and an adaptation strategy to customize the approach to a particular speaker using minimal training. The algorithms are evaluated on the TIMIT corpus, on a dysarthric speech corpus, and on the ICSI Switchboard spontaneous speech corpus. Results show that the proposed methods outperform three competing methods on both healthy and dysarthric speech. In addition, for spontaneous speech rate estimation, the result show a high correlation between the estimated speaking rate and ground truth values.