Segregating information about the size and shape of the vocal tract using a time-domain auditory model: The stabilised wavelet-Mellin transform

Segregating information about the size and shape of the vocal tract using a time-domain auditory model: The stabilised wavelet-Mellin transform
复制标题

DOI:
10.1016/s0167-6393(00)00085-6
复制
发表时间:
2002-03-01
影响因子:
3.2
通讯作者:
Patterson, RD
Patterson, RD
中科院分区:
计算机科学3区
文献类型:
--
作者:
Irino, T;Patterson, RD

文献摘要

被引文献

相似文献

我们听到男性和女性的元音发音大致相同,尽管不同群体的声道长度有很大差异。同时,我们可以识别说话人群体。这表明,听觉系统可以提取和分离关于声道大小的信息和关于其形状的信息。随着声道长度的增加或减少,声道脉冲反应的持续时间会扩大或缩小。有一种变换,即梅林变换,它不受时间膨胀的影响;它将在时间尺度上不同的脉冲响应映射到单个分布上,并将大小信息单独编码为标量常量。本文研究了梅林变换在元音归一化中的应用。在听觉系统中,声音最初在耳蜗处接受一种形式的小波分析,然后在每个频率通道中,周期性声音产生的重复模式似乎通过一种形式的时间间隔计算而稳定下来。结果就像一个间隔直方图的二维阵列,它被称为听觉图像。本文证明了存在一种二维形式的Mellin变换,它可以将不同大小的元音的听觉图像转换为不变的Mellin图像(MI),从而便于提取和分离与给定元音类型相关的大小和形状信息。在信号处理方面。声音的MI是声音的稳定小波变换的Mellin变换。我们认为,MI提供了一个很好的元音听觉正常化模型,这为从耳蜗皮层到皮层的听觉处理提供了一个很好的框架。(C)2002 Elsevier Science B.V.保留所有权利。
We hear vowels pronounced by men and women as approximately the same although the length of the vocal tract varies considerably from group to group. At the same time, we can identify the speaker group. This suggests that the auditory system can extract and separate information about the size of the vocal-tract from information about its shape. The duration of the impulse response of the vocal tract expands or contracts as the length of the vocal tract increases or decreases. There is a transform, the Mellin transform, that is immune to the effects of time dilation; it maps impulse responses that differ in temporal scale onto a single distribution and encodes the size information separately as a scalar constant. In this paper we investigate the use of the Mellin transform for vowel normalisation. In the auditory system, sounds are initially subjected to a form of wavelet analysis in the cochlea and then, in each frequency channel, the repeating patterns produced by periodic sounds appear to be stabilised by a form of time-interval calculation. The result is like a two-dimensional array of interval histograms and it is referred to as an auditory image. In this paper, we show that there is a two-dimensional form of the Mellin transform that can convert the auditory images of vowel sounds from vocal tracts with different sizes into an invariant Mellin image (MI) and, thereby, facilitate the extraction and separation of the size and shape information associated with a given vowel type. In signal processing terms. the MI of a sound is the Mellin transform of a stabilised wavelet transform of the sound. We suggest that the MI provides a good model of auditory vowel normalisation, and that this provides a good framework for auditory processing from cochlea to cortex. (C) 2002 Elsevier Science B.V. All rights reserved.