Tone feature extraction through parametric modeling and analysis-by-synthesis-based pattern matching

Tone feature extraction through parametric modeling and analysis-by-synthesis-based pattern matching
复制标题

通过参数化建模和基于综合分析的模式匹配提取音调特征

DOI:
10.1109/icassp.2003.1198719
复制
发表时间:
2003
期刊:
2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03).
影响因子:
--
通讯作者:
H. Kawai
H. Kawai
中科院分区:
--
文献类型:
--
作者:
Jinfu Ni;H. Kawai

文献摘要

被引文献

相似文献

针对大规模语音语料库的自动韵律标注问题,采用函数基频(F/sub 0/)模型从汉语F/sub 0/轮廓中提取声调的波峰和滑音特征。基于F/sub 0/模型对四个词汇声调进行建模并以参数形式表示,首先使用LBG(Linde-Buzo-Gray)算法对基线声调模式进行聚类,然后进行基于合成分析的模式匹配,从观察到的F/sub 0/轮廓和语音标签中估计潜在的声调峰值和声调模式类型。在确定音调峰值之后,重新估计音调滑动特征。在对来自8名母语者的968个话语的开放测试中,94%的自动估计的标签与手动标签一致。实验结果表明,该方法适用于F/sub 0/轮廓平滑和色调验证。
A functional fundamental frequency (F/sub 0/) model is applied to extract tone peak and gliding features from Mandarin F/sub 0/ contours aiming at automatic prosodic labeling of a large scale speech corpus. Modeling four lexical tones and representing them in a parametric form based on the F/sub 0/ model, we first cluster baseline tone patterns using the LBG (Linde-Buzo-Gray) algorithm, then perform analysis-by-synthesis-based pattern matching to estimate underlying tone peaks and tone pattern types from observed F/sub 0/ contours and phonetic labels with lexical tones. Tone gliding features are re-estimated after the determination of tone peaks. 94% of the automatically estimated labels were consistent with the manual labels in an open test of 968 utterances from eight native speakers. Also, experimental results indicate that the proposed method is applicable for F/sub 0/ contour smoothing and tone verification.