A full-band adaptive harmonic representation of speech

A full-band adaptive harmonic representation of speech
复制标题

语音的全频带自适应谐波表示

DOI:
10.21437/interspeech.2012-138
复制
发表时间:
2012
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Y. Stylianou
Y. Stylianou
中科院分区:
--
文献类型:
--
作者:
G. Degottex;Y. Stylianou

文献摘要

被引文献

相似文献

本文提出了一种全频带自适应谐波模型(aHM),该模型能够准确地重建语音的平稳和非平稳部分。该模型不需要任何浊音/浊音的决定,也不需要对音高轮廓的准确估计。其鲁棒性基于先前提出的自适应准谐波模型(aQHM),该模型提供了一种频率校正机制和基函数对输入信号特性的自适应。该方法克服了基于aQHM的初始方法在检测随时间变化的频率轨迹方面的局限性,特别是在中高频,通过采用带限迭代过程来重新估计基频。听力测试表明,aHM重建的语音与原始信号基本没有区别,优于标准正弦模型(SM)和基于aqhm的方法,而且重建所需的参数比SM少。
In this paper we present a full-band Adaptive Harmonic Model (aHM) that is able to accurately reconstruct stationary and non stationary parts of speech. The model does not require any voiced/unvoiced decision, neither an accurate estimation of the pitch contour. Its robustness is based on the previously suggested adaptive Quasi-Harmonic model (aQHM), which provides a mechanism for frequency correction and adaptivity of its basis functions to the characteristics of the input signal. The suggested method overcomes limitations of the initial method based on aQHM in detecting frequency tracks over time, especially at mid and high frequencies, by employing a bandlimited iterative procedure for the re-estimation of the fundamental frequency. Listening tests show that reconstructed speech using aHM is mainly indistinguishable from the original signal, outperforming standard sinusoidal models (SM) and the aQHMbased method, while it uses less parameters for the reconstruction than SM.