Understanding speech in the presence of other speech: Perceptual mechanisms for auditory scene analysis in human listeners
Understanding speech in the presence of other speech: Perceptual mechanisms for auditory scene analysis in human listeners
批准号:
ES/K004905/1
负责人:
Brian Roberts
金额:
$45.47万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --
中文摘要
我们理所当然地认为,在日常生活中,我们可以与其他人交谈,即使有明显的努力,也能被理解。然而,孤立地听到特定说话者的讲话是相当不寻常的;语音通常是在存在干扰声音的情况下听到的,例如其他说话者的声音。因此,负责我们听觉的人类听觉系统面临着识别到达我们耳朵的声音的哪些部分来自同一环境来源的挑战。这包括将那些来自一个来源(例如,一个说话者的声音)的声音元素与来自其他来源的声音元素分开,并以大脑中更高级别的过程(例如,涉及我们对语音的理解的那些)可以解释的方式对它们进行分组。如果没有“听觉场景分析”这个问题的解决方案,我们对语言(和其他声音)的感知就不会与产生它们的事件相对应。人类在进化过程中暴露在各种复杂的倾听环境中,因此我们通常非常成功地在其他说话者在场的情况下理解一个人的讲话。这与开发监听机器的尝试形成了鲜明对比,当面对复杂的监听环境时,例如开放式办公室或拥挤的派对,监听机器往往会灾难性地失败。有听力障碍的人类听众也发现这些环境非常困难,即使使用助听器或人工耳蜗技术的最新发展。到目前为止,大多数关于听觉场景分析的研究都集中在相对简单的声音上,并确定了一些感知分组和分离声音元素的一般原则。然而,至少根据目前的理解,这些原则似乎不足以完全解释言语的知觉分组。这是因为语音信号由不同且快速变化的声音流组成。我们母语的语言也是一种非常熟悉的刺激,所以到成年时,我们的听觉系统已经有很多年的时间来了解它的言语特有特性。这些特性还有助于成功地对语音进行感知分组。理解语音所需的大部分信息是由语音信号频谱中的几个宽峰的频率随时间的变化而携带的,称为共振峰。这个项目的目的是调查人类听者如何能够将适当的共振峰组合在一起,并拒绝其他共振峰,以便我们想要听的说话者的演讲能够被理解。我们将使用人类听者的感知实验来做到这一点,在这些实验中,我们测量目标语音的可理解性(例如,测量为正确报告的字数)在各种条件下如何变化。该项目将探索通用分组因素(即那些适用于各种声音的因素)和特定于语音的分组因素的作用,包括与语音清晰度(即我们说话时舌头、嘴唇和下巴移动的方式)和我们的语言规则相关的更高级别的限制。我们的方法是生成具有精确控制属性的人工语音类刺激,将目标语音与精心设计的提供可选分组可能性的“竞争对手”混合,并测量操纵这些竞争对手的属性如何影响人类听者识别混合中的目标语音的能力。该项目的结果不仅将提高我们对人类听者如何区分语音和干扰声音的理解,还将有助于完善计算机听力模型。这样的改进反过来将提供改善助听器和自动语音识别器等设备在嘈杂环境中运行时的性能的方法。
英文摘要
We take it for granted that we can converse with other people in daily life and be understood with little, if any, noticeable effort. However, it is fairly unusual to hear the speech of a particular talker in isolation; speech is typically heard in the presence of interfering sounds, such as the voices of other talkers. The human auditory system, which is responsible for our sense of hearing, therefore faces the challenge of identifying which parts of the sounds reaching our ears have originated from the same environmental source. This involves separating those sound elements coming from one source (e.g., the voice of one talker) from those arising from other sources, and grouping them in ways that can be interpreted by higher-level processes in the brain (such as those involved in our understanding of speech). Without a solution to this "auditory scene analysis" problem, our perceptions of speech (and other sounds) would not correspond to the events that produced them. Humans have been exposed to a variety of complex listening environments over the course of evolution, and so we are generally very successful at understanding the speech of one person in the presence of other talkers. This contrasts with attempts to develop listening machines, which often fail catastrophically when confronted with complex listening environments, such as an open-plan office or a crowded party. Human listeners with hearing impairment also find these environments very difficult, even when using the latest developments in hearing-aid or cochlear-implant technology.So far, most research on auditory scene analysis has focussed on relatively simple sounds and has identified a number of general principles for the perceptual grouping and separation of sound elements. However, at least as currently understood, these principles seem inadequate to explain fully the perceptual grouping of speech. This is because the speech signal consists of a diverse and rapidly changing stream of sounds. The speech of our native language is also a highly familiar stimulus, and so by adulthood our auditory system has had many years to learn about its speech-specific properties. These properties may also assist in the successful perceptual grouping of speech.Much of the information necessary to understand speech is carried by the changes in frequency over time of a few broad peaks in the frequency spectrum of the speech signal, known as formants. The aim of this project is to investigate how human listeners presented with speech sound mixtures are able to group together the appropriate formants, and to reject others, such that the speech of the talker we want to listen to can be understood. We will do so using perceptual experiments with human listeners, in which we measure how the intelligibility of target speech (measured, for example, as the number of words reported correctly) changes under a variety of conditions. The project will explore the roles of general-purpose grouping factors (i.e., those that apply to a wide variety of sounds) and of speech-specific grouping factors, including higher-level constraints associated with the articulation of speech (i.e., the way our tongue, lips, and jaw move when we speak) and with the rules of our language. Our approach is to generate artificial speech-like stimuli with precisely controlled properties, to mix target speech with carefully designed "competitors" that offer alternative grouping possibilities, and to measure how manipulating the properties of these competitors affects the ability of human listeners to recognise the target speech in the mixture. The results of this project will not only improve our understanding of how human listeners separate speech from interfering sounds, but will also help to refine computer models of listening. Such refinements will in turn provide ways of improving the performance of devices such as hearing aids and automatic speech recognisers when they operate in noisy environments.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Informational masking and the effects of differences in fundamental frequency and fundamental-frequency contour on phonetic integration in a formant ensemble
信息掩蔽以及基频和基频轮廓的差异对共振峰合奏中语音整合的影响
DOI:
10.1121/1.4949932
发表时间:
2016
期刊:
Journal of the Acoustical Society of America
影响因子:
2.4
作者:
[Summers R]
通讯作者:
Summers R
DOI:
10.1037/xhp0000038
发表时间:
2015-06
期刊:
JOURNAL OF EXPERIMENTAL PSYCHOLOGY-HUMAN PERCEPTION AND PERFORMANCE
影响因子:
2.1
作者:
[Roberts, Brian, Summers, Robert J., Bailey, Peter J.]
通讯作者:
Bailey, Peter J.
Informational masking and the effects of differences in fundamental frequency and fundamental-frequency contour on phonetic integration in a formant ensemble.
信息掩蔽以及基频和基频轮廓的差异对共振峰合奏中语音整合的影响。
DOI:
10.1016/j.heares.2016.10.026
发表时间:
2017
期刊:
Hearing research
影响因子:
2.8
作者:
[Summers RJ]
通讯作者:
Summers RJ
Informational masking of monaural target speech by a single contralateral formant.
单个对侧共振峰对单耳目标语音的信息掩蔽。
DOI:
10.1121/1.4919344
发表时间:
2015
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
[Roberts B]
通讯作者:
Roberts B
Across-formant integration and speech intelligibility: Effects of acoustic source properties in the presence and absence of a contralateral interferer.
跨共振峰积分和语音清晰度:存在和不存在对侧干扰源时声源特性的影响。
DOI:
10.1121/1.4960595
发表时间:
2016
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
[Summers RJ]
通讯作者:
Summers RJ
RAPID: Securing the LUMCON natural history collection, a vital Gulf Coast resource
-
批准号:2203268
-
项目类别:Standard Grant
-
资助金额:$19.91万
-
财政年份:2022
-
负责人:Brian Roberts
-
依托单位:
EAR-Climate: Collaborative Research: Methane Dynamics Across Microbe-to-Landscape Scales in Coastal Wetlands
-
批准号:2218581
-
项目类别:Continuing Grant
-
资助金额:$94.49万
-
财政年份:2022
-
负责人:Brian Roberts
-
依托单位:
REU Site: Interdisciplinary Research Experiences in Louisiana's Changing Coastal Environments
-
批准号:2150358
-
项目类别:Continuing Grant
-
资助金额:$52.92万
-
财政年份:2022
-
负责人:Brian Roberts
-
依托单位:
REU Site: Interdisciplinary Research Experiences in Changing Coastal Environments
-
批准号:1757887
-
项目类别:Continuing Grant
-
资助金额:$23.81万
-
财政年份:2018
-
负责人:Brian Roberts
-
依托单位:
Collaborative Research: A RAPID response to Hurricane Harvey's impacts on coastal carbon cycle, metabolic balance and ocean acidification
-
批准号:1760687
-
项目类别:Standard Grant
-
资助金额:$4.27万
-
财政年份:2017
-
负责人:Brian Roberts
-
依托单位:
Interference in spoken communication: Evaluating the corrupting and disrupting effects of other voices
-
批准号:ES/N014383/1
-
项目类别:Research Grant
-
资助金额:$42.04万
-
财政年份:2016
-
负责人:Brian Roberts
-
依托单位:
REU Site: Interdisciplinary Research Experiences in Changing Coastal Environments
-
批准号:1063036
-
项目类别:Continuing Grant
-
资助金额:$22.26万
-
财政年份:2011
-
负责人:Brian Roberts
-
依托单位:
Collaborative Research: RAPID: The 2011 Atchafalaya River Flood and a possible altered system state for the Atchafalaya River Delta Estuary
-
批准号:1141354
-
项目类别:Standard Grant
-
资助金额:$9.26万
-
财政年份:2011
-
负责人:Brian Roberts
-
依托单位:
RAPID: Effects of oiling and hydrologic remediation on baldcypress swamp elevation and ecosystem processes in the context of the BP Deepwater Horizon Oil Spill
-
批准号:1049838
-
项目类别:Standard Grant
-
资助金额:$16.34万
-
财政年份:2010
-
负责人:Brian Roberts
-
依托单位:
The perceptual organization of speech: Contributions of general and speech-specific factors
-
批准号:EP/F016484/1
-
项目类别:Research Grant
-
资助金额:$47.3万
-
财政年份:2008
-
负责人:Brian Roberts
-
依托单位:
国内基金
海外基金
儿童植入耳蜗后听觉行为与言语发展进程的关联性研究
-
批准号:81170916
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2011
-
负责人:刘莎
-
依托单位:
儿童植入人工耳蜗后开放式听觉言语发育特性研究
-
批准号:30872859
-
项目类别:面上项目
-
资助金额:30.0万元
-
批准年份:2008
-
负责人:刘莎
-
依托单位: