Sinusoidal modeling for nonstationary voiced speech based on a local vector transform

Sinusoidal modeling for nonstationary voiced speech based on a local vector transform
复制标题

DOI:
10.1121/1.2431581
复制
发表时间:
2007-03-01
影响因子:
2.4
通讯作者:
Yano, Masafumi
Yano, Masafumi
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Ito, Masashi;Yano, Masafumi

文献摘要

被引文献

相似文献

浊音语音信号可以表示为其瞬时频率和幅值随时间连续变化的正弦分量之和。从输入中确定这些参数,时变特性是算法的关键误差源,算法在局部分析段内假设它们是平稳的。为了克服这一问题,提出了一种新的方法——局部向量变换(LVT),该方法可以确定非平稳正弦波的瞬时频率和幅值。该方法不假设局部平稳性。在合成语音信号和自然语音信号的参数确定中检验了LVT的有效性。确定一阶谐波分量的瞬时频率的精度几乎等于时间校正瞬时频率法的精度,并且高于频谱拾峰、自相关和倒谱法的精度。LVT也能准确地确定瞬时幅值,而其他算法存在较大误差。由确定的参数通过LVT重建的信号与浊音的相应分量吻合较好。结果表明,该方法对时变语音信号的分析是有效的。(c) 2007年美国声学学会。
A voiced speech signal can be expressed as a sum of sinusoidal components of which instantaneous frequency and amplitude continuously vary with time. Determining these parameters from the input, the time-varying characteristics are crucial error sources for the algorithms, which assume their stationarity within a local analysis segment. To overcome this problem, a new method is proposed, local vector transform (LVT), which can determine instantaneous frequency and amplitude for nonstationary sinusoids. The method does not assume the local stationarity. The effectiveness of LVT Was examined in parameter determination for synthesized and naturally uttered speech signals. The instantaneous frequency for the first harmonic component was determined with an accuracy almost equal to that of the time-corrected instantaneous frequency method and higher accuracy than that of spectral peak-picking, autocorrelation, and cepstrum. The instantaneous amplitude was also determined accurately by LVT while considerable errors were left in the other algorithms. The signal reconstructed from, the determined parameters by LVT agreed well with the corresponding component of voiced speech. These results suggest that the method is effective for analyzing time-varying voiced speech signals. (c) 2007 Acoustical Society of America.