GENDER RECOGNITION FROM SPEECH .2. FINE ANALYSIS

GENDER RECOGNITION FROM SPEECH .2. FINE ANALYSIS
复制标题

DOI:
10.1121/1.401664
复制
发表时间:
1991-10-01
影响因子:
2.4
通讯作者:
WU, K
WU, K
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
CHILDERS, DG;WU, K

文献摘要

被引文献

相似文献

本研究的目的是调查数字语音处理和模式识别技术在语音中自动识别性别的潜在有效性。第一部分粗略分析[K.Wu和D.G.Childers,J.Acoust.SoC。上午好。90,1828-1840(1991)]考察了各种特征向量和距离度量,以确定它们是否适合从元音、清音和浊音中识别说话人的性别。一种基于从元音提取的特征向量的识别方案使用52个说话人(27个男性和25个女性)的数据库实现了对说话人性别的100%正确识别。本文对元音的特征进行了详细的分析,包括共振峰的频率、带宽和幅度,以及说话人发声的基频。精细分析使用了音调同步闭相分析技术。通过采用可变遗忘因子的闭相加权递推最小二乘法提取频率、带宽和幅度等共振峰的详细特征。电声门图信号被用来定位语音信号的闭相部分。采用双因素方差分析(ANOVA)检验性别特征之间的差异。采用模式识别的方法评价了元音分组特征的相对重要性。研究得到了许多有趣的结果,包括第二共振峰频率比基本频率对性别的识别略好,正确识别率分别为98.1%和96.2%。统计检验表明,女性说话者的频谱比男性说话者的频谱有更大的斜率(或倾斜度)。结果表明,在基频和声道共振特征中嵌入了冗余的性别信息。观察到女性声音的特征向量比男性声音的特征向量具有更高的组内变异。本研究中的数据也被用于复制Peterson和Barney的部分[J.Acoust。SoC。上午好。24,175-184(1952)]男性和女性说话者的元音研究。
The purpose of this research was to investigate the potential effectiveness of digital speech processing and pattern recognition techniques for the automatic recognition of gender from speech. In part I Coarse Analysis [K. Wu and D. G. Childers, J. Acoust. Soc. Am. 90, 1828-1840 (1991)] various feature vectors and distance measures were examined to determine their appropriateness for recognizing a speaker's gender from vowels, unvoiced fricatives, and voiced fricatives. One recognition scheme based on feature vectors extracted from vowels achieved 100% correct recognition of the speaker's gender using a database of 52 speakers (27 male and 25 female). In this paper a detailed, fine analysis of the characteristics of vowels is performed, including formant frequencies, bandwidths, and amplitudes, as well as speaker fundamental frequency of voicing. The fine analysis used a pitch synchronous closed-phase analysis technique. Detailed formant features, including frequencies, bandwidths, and amplitudes, were extracted by a closed-phase weighted recursive least-squares method that employed a variable forgetting factor, i.e., WRLS-VFT. The electroglottograph signal was used to locate the closed-phase portion of the speech signal. A two-way statistical analysis of variance (ANOVA) was performed to test the differences between gender features. The relative importance of grouped vowel features was evaluated by a pattern recognition approach. Numerous interesting results were obtained, including the fact that the second formant frequency was a slightly better recognizer of gender than fundamental frequency, giving 98.1% versus 96.2% correct recognition, respectively. The statistical tests indicated that the spectra for female speakers had a steeper slope (or tilt) than that for males. The results suggest that redundant gender information was imbedded in the fundamental frequency and vocal tract resonance characteristics. The feature vectors for female voices were observed to have higher within-group variations than those for male voices. The data in this study were also used to replicate portions of the Peterson and Barney [J. Acoust. Soc. Am. 24, 175-184 (1952)] study of vowels for male and female speakers.