MODELING THE PERCEPTION OF CONCURRENT VOWELS - VOWELS WITH THE SAME FUNDAMENTAL-FREQUENCY

MODELING THE PERCEPTION OF CONCURRENT VOWELS - VOWELS WITH THE SAME FUNDAMENTAL-FREQUENCY
复制标题

DOI:
10.1121/1.397684
复制
发表时间:
1989-01-01
影响因子:
2.4
通讯作者:
SUMMERFIELD, Q
SUMMERFIELD, Q
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
ASSMANN, PF;SUMMERFIELD, Q

文献摘要

被引文献

相似文献

在一系列关于从多接受者波形中提取语音信息的研究中,首次研究了听众识别同时合成元音对的能力。元音对的两个成员具有相同的起始和偏移时间和恒定的100赫兹基频。听众识别这两个元音的准确率明显高于偶然。(a)级联形成峰合成和(b)加性谐波合成产生的元音的正确响应和混淆模式相似,加性谐波合成用一个振幅相等的单对谐波取代了最低的三个共振峰中的每一个。为了选择一个合适的模型来描述听众的表现,我们评估了四种模式匹配过程。每个人都预测了(i)任何一个元音被选为两个回答之一的概率,以及(ii)任何一对元音被选的概率。这些可能性是通过测量双元音与单元音参考模式的听觉激发模式的相似性来估计的。单个响应中高达88%的方差和成对响应中高达67%的方差可以通过在激励模式中突出显示谱峰和谱肩的程序来解释。对激励模式的所有区域分配均匀权重的程序给出的预测结果较差。这些发现支持了一个假设,即听觉系统在识别元音时特别注意频谱峰值的频率,也可能是肩部的频率。这种策略的一个优点是,当频谱形状的其他方面被竞争声音掩盖时,频谱峰和肩部可以指示共振峰的频率。
The ability of listeners of identify pairs of simultaneous synthetic vowels has been investigated in the first of a series of studies on the extraction of phonetic information from multiple-taker waveforms. Both members of the vowel pair had the same onset and offset times and a constant fundamental frequency of 100 Hz. Listeners identified both vowels with an accuracy significantly greater than chance. The pattern of correct responses and confusions was similar for vowels generated by (a) cascade formant synthesis and (b) additive harmonic synthesis that replaced each of the lowest three formants with a single pair of harmonics of equal amplitude. In order to choose an appropriate model for describing listeners'' performance, four pattern-matching procedures were evaluated. Each predicted the probability that (i) any individual vowel would be selected as one of the two responses, and (ii) any pair of vowels would be selected. These probabilities were estimated from measures of the similarities of the auditory excitation patterns of the double vowels to those of single-vowel reference patterns. Up to 88% of the variance in individual responses and up to 67% of the variance in pairwise responses could be accounted for by procedures that highlighted spectral peaks and shoulders in the excitation pattern. Procedures that assigned uniform weight to all regions of the excitation pattern gave poorer predictions. These findings support the hypothesis that the auditory system pays particular attention to the frequencies of spectral peaks, and possibly also of shoulders, when identifying vowels. One virtue of this strategy is that the spectral peaks and shoulders can indicate the frequencies of formants when other aspects of spectral shape are obscured by competing sounds.