MODELING THE PERCEPTION OF CONCURRENT VOWELS - VOWELS WITH DIFFERENT FUNDAMENTAL FREQUENCIES

MODELING THE PERCEPTION OF CONCURRENT VOWELS - VOWELS WITH DIFFERENT FUNDAMENTAL FREQUENCIES
复制标题

DOI:
10.1121/1.399772
复制
发表时间:
1990-08-01
影响因子:
2.4
通讯作者:
SUMMERFIELD, Q
SUMMERFIELD, Q
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
ASSMANN, PF;SUMMERFIELD, Q

文献摘要

被引文献

相似文献

如果两个具有不同基频 (fo''s) 的元音同时单声道呈现,听众经常会听到两个讲话者以不同的音调发出不同的元音。本文描述了对听觉和感知过程的四种计算模型的评估,这可能是这种能力的基础。每个模型涉及四个阶段:(i)使用“听觉”滤波器组进行频率分析,(ii)确定刺激中存在的音调,(iii)通过对与每个音调相关的能量进行分组以创建两个派生频谱模式来分离竞争语音源,以及(iv)对派生频谱模式进行分类以预测听众元音识别响应的概率。 “位置”模型通过分析滤波器组通道上的均方根电平分布来执行音调确定和频谱分离操作。 “地点-时间”模型通过分析每个通道中波形的周期性来执行这些操作。在“线性”版本中,地点和地点时间模型直接对滤波器产生的波形进行操作。在它们的“非线性”版本中,类似的操作被应用于附加级的输出,该附加级对滤波后的波形应用了压缩非线性。与其他三个模型相比,非线性地点时间模型提供了对并发合成元音对的 fo 的最准确估计,并且最接近于预测听众对此类刺激的识别反应。尽管该模型有一些局限性,但结果与使用地点时间分析来隔离竞争声源的想法是一致的。
If two vowels with different fundamental frequencies (fo''s) are presented simultaneously and monaurally, listeners often hear two talkers producing different vowels on different pitches. This paper describes the evaluation of four computational models of the auditory and perceptual processes which may underlie this ability. Each model involves four stages: (i) frequency analysis using an "auditory" filter bank, (ii) determination of the pitches present in the stimulus, (iii) segregation of the competing speech sources by grouping energy associated with each pitch to create two derived spectral patterns, and (iv) classification of the derived spectral patterns to predict the probabilities of listeners'' vowel-identification responses. The "place" models carry out the operations of pitch determination and spectral segregation by analyzing the distribution of rms levels across the channels of the filter bank. The "place-time" models carry out these operations by analyzing the periodicities in the waveforms in each channel. In their "linear" versions, the place and place-time models operate directly on the waveforms emerging from the filters. In their "nonlinear" versions, analogous operations are applied to the output of an additional stage which applied a compressive nonlinearity to the filtered waveforms. Compared to the other three models, the nonlinear place-time model provides the most accurate estimates of the fo''s of pairs of concurrent synthetic vowels and comes closest to predicting the identification responses of listeners to such stimuli. Although the model has several limitations, the results are compatible with the idea that a place-time analysis is used to segregate competing sound sources.