Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant tracking

Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant tracking
复制标题

使用时变加权线性预测对语音信号进行准闭相位分析,以实现准确的共振峰跟踪

DOI:
--
复制
发表时间:
2016
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
P. Alku
P. Alku
中科院分区:
--
文献类型:
--
作者:
Dhananjaya N. Gowda;Manu Airaksinen;P. Alku

文献摘要

被引文献

相似文献

时间加权线性预测的最新研究表明,语音信号的准闭相(QCP)分析提供了更好的声道和声门源的建模。准闭相分析给出了声门循环的闭相上的更多权重,同时不强调通常预测不佳的显著激励的瞬间周围的区域。然而,包括QCP分析在内的所有传统分析技术都是在短时间间隔内进行的。它们不对声道系统或声门源施加任何连续性约束。这种约束通常在稍后阶段施加,以随时间平滑或跟踪估计的特征。时变线性预测(TVLP)提供了一个框架,用于对声道形状施加长期连续性约束的语音进行建模。在本文中,我们提出了一种新的方法,准确的建模和跟踪的声道共振相结合的优势,QCP分析与TVLP。Forcenter跟踪实验表明,在包括不同语音类型和广泛的基频范围在内的各种条件下,与传统的LP或TVLP方法相比,性能得到了一致的改善。
Recent research on temporally weighted linear prediction shows that quasi closed phase (QCP) analysis of speech signals provides better modeling of the vocal tract and the glottal source. Quasi closed phase analysis gives more weightage on the closed phase of the glottal cycle, at the same time deemphasizing the region around the instant of significant excitation which is often poorly predicted. However, all the traditional analysis techniques including the QCP analysis is performed over short intervals of time. They do not impose any continuity constraints either on the vocal tract system or the glottal source. Such constraints are often imposed at a later stage to either smooth or track the estimated features over time. Time varying linear prediction (TVLP) provides a framework for modeling speech with a long-term continuity constraint imposed on the vocal tract shape. In this paper, we propose a new method for accurate modeling and tracking of the vocal tract resonances by integrating the advantages of a QCP analysis with that of TVLP. Formant tracking experiments show consistent improvement in performance over traditional LP or TVLP methods under a variety of conditions including different voice types and over a wide range of fundamental frequency.