Book Review: Auditory Scene Analysis: The Perceptual Organization of Sound

Book Review: Auditory Scene Analysis: The Perceptual Organization of Sound
复制标题

DOI:
10.1080/14640749208401289
复制
发表时间:
1992-01
影响因子:
1.7
通讯作者:
Q. Summerfield
Q. Summerfield
中科院分区:
心理学4区
文献类型:
--
作者:
Q. Summerfield

文献摘要

被引文献

相似文献

这个世界充满了声音的来源。当我写这篇评论的时候,我能听到文字处理机的嗡嗡声,风中门的吱吱声,远处飞机的隆隆声,附近汽车的经过,小鸟的啁啾声,我的邻居在他家门口说话,他儿子的高保真音响的音乐,还有隔壁房间收音机里有人说话。虽然每一种来源都产生一种特定的气压变化模式,但当它们到达我的耳朵时,这些变化已经加在一起了,但我清楚地感知到每一种来源。听众使用什么样的感知分组和分离原则来划分这样的声音混合?哪些原则自动适用于所有的声音?哪些是专门为特定类别的声音,如语音?在音乐创作中,这些原则是如何被利用的?这些是这本冗长、学术性强、但可读性强的书的主要关注点。布雷格曼的方法是功能性的而不是生理性的,是经验性的而不是计算性的。他提供了一个全面的审查和解释知觉实验到1989年左右,所以他的书早于最近试图实施听觉分组原则的计算模型,并找到一个生理基板。一个重要的区别贯穿全书。Bregman认为听觉的分组和分离有两种原则:“基于图式的”和“原始的”。基于模式的原则特定于特定类型的源。它们是由听者学习的,它们的应用是在注意力控制下的。一个例子可能是使用的知识,一种乐器的音色,以遵循其部分在合奏。另一个示例可以是使用语音知识来将声学线索整合到语音感知中。相比之下,原始的分组原则是天生的,通过进化习得的。它们自动利用声音和声源的基本物理特性。举例来说:谐振器的尺寸通常变化缓慢;它们经常在很宽的频率范围内同时产生能量;当它们振动时,它们在离散的频率范围内产生能量。
The world is full of sources of sound. As I write this review, I can hear the humming of the word processor, the creaking of a door in the wind, the distant rumble of an aeroplane, the passage of a car close by, a bird twittering, my neighbour talking on his doorstep, music from his son’s hi-fi, and someone speaking on the radio in the next room. Although each source generates a particular pattern of changes in air-pressure, the changes have summed together by the time they reach my ears, yet I perceive each source distinctly. What principles of perceptual grouping and segregation do listeners use to partition such mixtures of sound? Which principles are applied automatically to all sounds? Which are specialized for particular classes of sound, such as speech? In what ways have the principles been exploited in musical composition? These are the major concerns of this lengthy, scholarly, but readable book. Bregman’s approach is functional not physiological, empirical not computational. He provides a comprehensive review and interpretation of perceptual experiments up to about 1989, so his book pre-dates recent attempts to implement auditory grouping principles in computational models and to find a physiological substrate for them. One important distinction is sustained throughout the book. Bregman argues that there are two kinds of principle for auditory grouping and segregation: “schema-based’’ and “primitive”. Schema-based principles are specific to particular types of source. They are learnt by listeners, and their application is under attentional control. One example may be the use of the knowledge of the timbre of an instrument to follow its part in an ensemble. Another example may be the use of phonetic knowledge to integrate acoustic cues in speech perception. Primitive grouping principles, in contrast, are innate, learnt through evolution. They automatically exploit fundamental physical properties of sounds and sound sources. For example: the sizes of resonators generally change slowly; they often generate energy simultaneously over a wide frequency range; when they vibrate, they create energy at the discrete