课题基金 / 基金详情

Probabilistic Auditory Scene Analysis

Probabilistic Auditory Scene Analysis
概率听觉场景分析
批准号:
EP/G050821/1
负责人:
Richard Turner
金额:
$29.57万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --

项目摘要

项目成果

Richard Turner的其他基金

相似基金

相关文献

中文摘要
翻译
听觉环境通常非常复杂。例如,鸡尾酒会包括许多来源;玻璃杯的丁当声;许多客人的闲聊声;背景音乐的声音。然而,我们的听觉系统可以理解这种上升;它可以计算出有多少声源,并确定每个声源对场景的贡献。值得注意的是,它可以用一个麦克风的信息来做到这一点。听觉神经科学的一个主要目标是了解听觉系统是如何实现这一壮举的。一般来说,听觉场景分析被认为有三个阶段。第一个阶段在生理学上是很好理解的,那就是将传入的声音转换成时频表示。这揭示了在特定时间的局部能量。在第二阶段,心理物理学证据表明,原始的分组原则是用来分组局部地区的频谱时间能量bridging从一个共同的来源。通过使用简单的刺激-如音调和噪音-一长串的原始分组原则已经阐明。例如,良好的连续性原则识别平滑变化的功能与一个单一的来源和突变作为一个签名的分离源。在听觉场景分析的最后阶段,称为基于模式的分组,更高层次的知识,如音乐或语音的结构,被用来绑定频谱-时间能量组成流,使得每个源有一个流。一个重要的问题是听觉皮层在听觉场景分析中的作用,因为它还没有很好地建立起来。另一个问题是所建立的基元分组规则列表的通用性和完整性。因为尽管这些原理成功地解释了对简单声音的感知,但对于自然声音的描述是否成功和相关还不清楚。本计画旨在透过模拟工作、心理物理学实验及神经记录实验来解决这些问题。新的观点是把原始的分组原则看作是听觉场景的潜变量模型中的推理。潜变量模型描述的是一个听觉场景,比如鸡尾酒会上遇到的场景,是如何由潜在的噪音组成的,比如叮当作响的眼镜和喋喋不休的客人。它还包括对这些来源的统计数据的描述,比如说,眼镜的叮当声往往是孤立的、高频的事件,而喋喋不休则更恒定、频率更低。这个想法是,大脑试图利用统计数据的先验知识来推断这些潜在的来源。新的概率推理工具可以使这些直觉具体化,这种新的视角称为概率场景分析,它有两个主要的优点:一个是实用性的,一个是理论性的。实际的优点是,声音的统计特性可以用来产生复杂的刺激,但控制结构,用于实验。理论上的好处是,连续分组规则的列表,以及它们权衡的方式,现在来自声音的统计数据;不再需要启发式实现。这使我们能够预测实验的结果。特别是,心理物理学实验旨在解决听觉分组如何在合成纹理(例如雨,风,水等)中运作。and whether是否this is consistent一致with the probabilistic概率account帐户.此外,神经记录实验将探讨听觉皮层在听觉场景分析中的作用,以及它代表声音的高水平统计量的假设,如缓慢变化的调制成分。
英文摘要
Auditory environments are typically very complicated. For example, thecocktail party comprises many sources; the chinking of glasses; thechattering of the many guests; the sound of backgroundmusic. Nevertheless, our auditory system can make sense of such ascene; it can work out how many acoustic sources there are anddetermine the individual contributions to the scene fromeach. Remarkably, it can do this using the information from a singlemicrophone. A major goal of auditory neuroscience is to understandhow the auditory system achieves this feat.Broadly speaking, it is thought that there are three stages toauditory scene analysis. The first stage is well understoodphysiologically and that is to convert the incoming sound into atime-frequency representation. This reveals the local energy in afrequency band at a particular time. In the second stage,psychophysical evidence suggests that primitive grouping principlesare used to group local regions of spectral-temporal energy arisingfrom a common source. By using simple stimuli - like tones and noise -a long list of primitive grouping principles have been elucidated. Forexample, the principle of good continuation identifies smoothlyvarying features with a single source and abrupt changes as asignature of separate sources. In the final stage of auditory sceneanalysis, called schema-based grouping, higher level knowledge, likethe structure of music or speech, is used to bind the groups ofspectral-temporal energy into streams so that there is one stream foreach source.There are many outstanding questions with this framework. Oneimportant open question is the role that auditory cortex plays inauditory scene analysis as it is not well established. Anotherconcerns the generality and completeness of the established list ofprimitive grouping rules. For although the principles successfullycharacterise perception of simple sounds it is unclear how successfuland relevant the description is for natural sounds. This project aims to resolve these questions though modelling work,psychophysics experiments and neural recording experiments. The newidea is to view the primitive grouping principles as arising frominference in a latent variable model of auditory scenes. A latentvariable model is a description of how an auditory scene, like thatencountered at a coctail party, is composed of latent auditorysources, like the chinking glasses and chattering guests. It alsoincludes a description of the statistics of these sources, like thefact that the chinking glasses tend to be isolated, high frequencyevents whist the chattering rather more constant and lower infrequency. The idea is that the brain is trying to infer these latentsources using prior knowledge of their statistics. New tools ofprobabilistic inference can make these intuitions concrete.This new perspective, called probabilistic scene analysis, has twomain advantages; one practical and one theoretical. The practicaladvantage is that a statistical characterisation of sounds can be usedto produce stimuli with complicated, but controlled structure, for usein experiments. The theoretical benefit is that the list of primitivegrouping rules, and the manner in which they trade off, are nowderived from the statistics of sounds; Heuristic implementation is nolonger required. This enables us to predict the results of theexperiments. In particular, the psychophysics experiments are aimedat resolving both how auditory grouping operates in synthetic auditorytextures (e.g. rain, wind, water etc.) and whether this is consistentwith the probabilistic account. Furthermore, the neural recordingexperiments will investigate the role of auditory cortex in auditoryscene analysis, and the hypothesis that it is representing high levelstatistics of sounds like slowly varying modulatory components.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.17863/cam.15597
发表时间: 2015-04
期刊:
影响因子: --
作者: [A. G. Matthews;J. Hensman;Richard E. Turner;Zoubin Ghahramani]
通讯作者: A. G. Matthews;J. Hensman;Richard E. Turner;Zoubin Ghahramani
DOI: 10.17863/cam.21348
发表时间: 2016-02
期刊:
影响因子: --
作者: [T. Bui;D. Hernández-Lobato;José Miguel Hernández-Lobato;Yingzhen Li;Richard E. Turner]
通讯作者: T. Bui;D. Hernández-Lobato;José Miguel Hernández-Lobato;Yingzhen Li;Richard E. Turner
DOI: 10.1080/14697688.2013.851402
发表时间: 2013-11
期刊: Quantitative Finance
影响因子: 1.3
作者: [H. Christensen;Richard E. Turner;Simon I. Hill;S. Godsill]
通讯作者: H. Christensen;Richard E. Turner;Simon I. Hill;S. Godsill
Neural Adaptive Sequential Monte Carlo
神经自适应序列蒙特卡罗
DOI: 10.48550/arxiv.1506.03338
发表时间: 2015
期刊:
影响因子: --
作者: [Gu S]
通讯作者: Gu S
Machine Learning for Tomorrow: Efficient, Flexible, Robust and Automated
  • 批准号:
    EP/T005637/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $208.89万
  • 财政年份:
    2020
  • 负责人:
    Richard Turner
  • 依托单位:
Nanoporous polymer particles and gels containing functionalized semi-rigid copolymer structures
Machine Learning for Hearing Aids: Intelligent Processing and Fitting
  • 批准号:
    EP/M026957/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $72.04万
  • 财政年份:
    2015
  • 负责人:
    Richard Turner
  • 依托单位:
Unifying audio signal processing and machine learning: a fundamental framework for machine hearing
  • 批准号:
    EP/L000776/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $12.37万
  • 财政年份:
    2013
  • 负责人:
    Richard Turner
  • 依托单位:
海外基金