Single and Multiple F0 Contour Estimation Through Parametric Spectrogram Modeling of Speech in Noisy Environments

Single and Multiple F0 Contour Estimation Through Parametric Spectrogram Modeling of Speech in Noisy Environments
复制标题

DOI:
10.1109/tasl.2007.894510
复制
发表时间:
2007
期刊:
IEEE Trans. Speech Audio Process.
影响因子:
--
通讯作者:
Jonathan Le Roux;H. Kameoka;Nobutaka Ono;A. Cheveigné;S. Sagayama
Jonathan Le Roux;H. Kameoka;Nobutaka Ono;A. Cheveigné;S. Sagayama
中科院分区:
其他
文献类型:
--
作者:
Jonathan Le Roux;H. Kameoka;Nobutaka Ono;A. Cheveigné;S. Sagayama

文献摘要

被引文献

相似文献

本文提出了一种新颖的 0 轮廓估计算法,该算法基于从功率谱导出的语音部分的精确参数描述。该算法能够在各种噪声环境中执行,并能够估计同信道并发语音的 0。语音频谱被建模为由表示为样条曲线的公共 0 轮廓控制的频谱簇序列。这些簇是使用新的 EM 算法公式通过功率密度的无监督二维时频聚类获得的,同时估计它们的公共 0 轮廓。为整个话语提取平滑的 0 轮廓,将其浊音部分连接在一起。噪声模型用于处理非谐波背景噪声,否则会干扰语音谐波部分的聚类。我们在多个任务上与现有方法进行比较,评估我们的算法,结果表明:1)它在干净的单说话人语音上具有竞争力,2)它在存在噪声的情况下优于现有方法,3)它在估计同道并发语音的多个 0 轮廓方面优于现有方法。
This paper proposes a novel 0 contour estimation algorithm based on a precise parametric description of the voiced parts of speech derived from the power spectrum. The algorithm is able to perform in a wide variety of noisy environments as well as to estimate the 0s of cochannel concurrent speech. The speech spectrum is modeled as a sequence of spectral clusters governed by a common 0 contour expressed as a spline curve. These clusters are obtained by an unsupervised 2-D time-frequency clustering of the power density using a new formulation of the EM algorithm, and their common 0 contour is estimated at the same time. A smooth 0 contour is extracted for the whole utterance, linking together its voiced parts. A noise model is used to cope with nonharmonic background noise, which would otherwise interfere with the clustering of the harmonic portions of speech. We evaluate our algorithm in comparison with existing methods on several tasks, and show 1) that it is competitive on clean single-speaker speech, 2) that it outperforms existing methods in the presence of noise, and 3) that it outperforms existing methods for the estimation of multiple 0 contours of cochannel concurrent speech.