The voice source in speech production: data, analysis and models

The voice source in speech production: data, analysis and models
复制标题

语音生成中的语音源:数据、分析和模型

DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Yen
Yen
中科院分区:
--
文献类型:
--
作者:
A. Alwan;Yen

文献摘要

参考文献

被引文献

相似文献

关于语音质量的语音源分析对于理解人类语音产生系统是必不可少的,这可以导致更好的语音建模以改进广泛的应用。然而,由于声带的位置,分析声源往往受到缺乏直接观测以校准算法的阻碍。 本文研究了两种语音源和音质分析方法。在第一种方法中,通过分析声带高速成像的声门区波形来提取声源波形。这些直接观测导致了一种新的源模型的开发,与现有模型相比,该模型更准确。然后提出了一种码本搜索技术来从声学数据中估计源信号。一些模型参数的结果是令人满意的,例如开商和打开速度。然而,误差分析表明,该算法需要合理的共振峰频率约束,在某些情况下可能很难自动获得。 在第二种方法中,语音源相关测量被用于三种语音质量应用:语音源分析、自动性别分类和韵律分析。在声源分析中,在声源模型参数的背景下检查声学测量,所述声源模型参数是通过模型拟合声门Arca波形获得的。结果表明,模型参数与相关的声学指标,如不对称系数和谐噪比指标之间具有相关性。研究还表明,模型参数和相关的声学测量受语音质量类型(压音、正常和喘息)的影响。在性别分类中,与声源相关的测量在较年轻(10-14岁)的说话者中被发现更有帮助,而传统的音高和共振峰频率特征不那么有用。对韵律的分析表明,除其他因素外,与音高重音相关的特征不一定集中在目标音节上,而取决于其他韵律事件的位置。
Analysis of the voice source with respect to voice quality is essential to the understanding of the human speech production system, which can lead to better speech modeling for improving a vast range of applications. However, due to the position of the vocal folds, analyzing the source is often hampered by the lack of direct observations with which to calibrate algorithms. In this dissertation, two approaches to voice source and voice quality analysis were pursued. In the first approach, the source waveform was extracted by analyzing the glottal area waveforms from high-speed imaging of the vocal folds. These direct observations led to the development of a new source model, which is more accurate compared to existing models. A codebook search technique was then proposed to estimate the source signal from the acoustic data. Results were promising for a number of model parameters such as the open quotient and speed of opening. However, error analysis showed that the algorithm required reasonable formant-frequency constraints which may be difficult to obtain automatically in some cases. In the second approach, voice source related measures were used in three voice quality applications: voice source analysis, automatic gender classification and prosody analysis. In voice source analysis, acoustic measures were examined in the context of the voice source model parameters obtained from model-fitting the glottal arca waveforms. Results showed that correlations could be made between model parameters and the related acoustic measures, such as the asymmetry coefficient and harmonic-to-noise ratio measures. It was also shown that the model parameters and related acoustic measures were affected by the type of voice quality (pressed, normal and breathy). In gender classification, voice source related measures were found to be more helpful in younger (10-14 year old) speakers, where traditional pitch and formant frequency features were less useful. Analysis of prosody showed that, amongst other things, features correlated to pitch accents were not necessarily centered at the target syllable, and depended on the position of other prosodic events.
DOI: 10.1121/1.401663
发表时间: 1991-10
期刊: The Journal of the Acoustical Society of America
影响因子: --
作者:
Ke Wu;D. Childers
通讯作者: Ke Wu;D. Childers
DOI: 10.1044/jshr.3704.769
发表时间: 1994-08-01
期刊: JOURNAL OF SPEECH AND HEARING RESEARCH
影响因子: --
作者:
HILLENBRAND, J;CLEVELAND, RA;ERICKSON, RL
通讯作者: ERICKSON, RL