A model of acoustic interspeaker variability based on the concept of formant-cavity affiliation.

A model of acoustic interspeaker variability based on the concept of formant-cavity affiliation.
复制标题

基于共振峰腔隶属关系概念的声学扬声器间变异性模型。

DOI:
10.1121/1.1631946
复制
发表时间:
2004
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
G. Bailly
G. Bailly
中科院分区:
--
文献类型:
--
作者:
L. Apostol;P. Perrier;G. Bailly

文献摘要

参考文献

被引文献

相似文献

提出了一种语音共振峰模式的说话人间变异性模型。据推测,这种变异性起源于扬声器之间存在的差异,在各自的长度的前,后声道腔。为了从声学语音信号的频谱描述中表征说话者之间的这些声道差异,根据共振峰-腔联系的概念,将每个共振峰解释为特定声道腔的共振。因此,其频率可以直接与相应的腔长相关,并且可以基于对应于相同谐振的共振峰的频率比提出从扬声器A到扬声器B的转换模型。为了最小化为每个说话者记录的声音的数量,以便执行该说话者变换,仅对三个极端的基本元音[i,a,u]精确地计算频率比,并且通过插值函数对剩余的元音进行近似。该方法是通过其能力,以转换(F1,F2)共振峰模式的8个口头元音发音的5名男性扬声器到(F1,F2)模式的发音模型所产生的声道相应的元音。由此产生的共振峰模式进行比较,在文献中发表的归一化技术提供的。所提出的方法被认为是有效的,但也观察和讨论了一些限制。这些限制可以与共振峰腔附属模型本身或与扬声器特定声道几何形状在横截面方向上的可能影响相关联,该模型可能没有考虑到这一点。
A method is proposed to model the interspeaker variability of formant patterns for oral vowels. It is assumed that this variability originates in the differences existing among speakers in the respective lengths of their front and back vocal-tract cavities. In order to characterize, from the spectral description of the acoustic speech signal, these vocal-tract differences between speakers, each formant is interpreted, according to the concept of formant-cavity affiliation, as a resonance of a specific vocal-tract cavity. Its frequency can thus be directly related to the corresponding cavity length, and a transformation model can be proposed from a speaker A to a speaker B on the basis of the frequency ratios of the formants corresponding to the same resonances. In order to minimize the number of sounds to be recorded for each speaker in order to carry out this speaker transformation, the frequency ratios are exactly computed only for the three extreme cardinal vowels [i, a, u] and they are approximated for the remaining vowels through an interpolation function. The method is evaluated through its capacity to transform the (F1,F2) formant patterns of eight oral vowels pronounced by five male speakers into the (F1,F2) patterns of the corresponding vowels generated by an articulatory model of the vocal tract. The resulting formant patterns are compared to those provided by normalization techniques published in the literature. The proposed method is found to be efficient, but a number of limitations are also observed and discussed. These limitations can be associated with the formant-cavity affiliation model itself or with a possible influence of speaker-specific vocal-tract geometry in the cross-sectional direction, which the model might not have taken into account.
DOI: 10.1121/1.2024119
发表时间: 1989
期刊: The Journal of the Acoustical Society of America
影响因子: --
作者:
James D. Miller
通讯作者: James D. Miller