课题基金 / 基金详情

Integrating sound and context recognition for acoustic scene analysis

Integrating sound and context recognition for acoustic scene analysis
集成声音和上下文识别以进行声学场景分析
批准号:
EP/R01891X/1
负责人:
Emmanouil Benetos
金额:
$12.47万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在过去十年中,生成的音频数据量急剧增加,从用户生成的内容,视听档案中的记录,到在城市,自然或家庭环境中捕获的传感器数据。需要检测和识别环境记录中的声音事件(例如敲门声,玻璃破碎)以及识别音频记录的上下文(例如火车站,会议),这导致了一个新的研究领域的出现:声学场景分析。声学场景分析的新兴应用包括智能家居和智能城市的声音识别技术开发、安全/监控、音频检索和存档、环境辅助生活以及自动生物多样性评估等,但目前的声音识别技术无法适应不同的环境或情况(例如,在办公室环境中的声音识别,假设特定的房间属性、工作时间、室外噪声和天气条件)。如果关于上下文的信息是可用的,则其通常由用于整个音频流的单个标签来表征,而不考虑复杂且不断变化的环境,例如当使用手持设备进行记录时,其中上下文可以由多个时变因素组成,并且可以由多个标签来表征。该项目将通过调查和开发上下文感知声音识别技术来解决上述缺点。我们假设音频流的上下文由几个随时间变化的因素组成,这些因素可以被视为不同环境和情况的组合;不断变化的上下文反过来通知系统要识别的声音的类型和属性。基于信号处理和机器学习理论,将研究和开发上下文和声音识别的方法。该项目的主要贡献将是一个算法框架,该框架联合识别基于音频的上下文和声音事件,适用于具有多个声源和时变环境的复杂音频流。将使用在城市和家庭环境中记录的复杂音频流以及使用模拟音频数据来评估拟议的软件框架,以便仔细控制上下文和声音属性,并获得准确注释的好处。为了进一步推动情境感知声音识别系统的研究,将结合声学场景和事件的检测和分类(DCASE)公众挑战赛组织公众评估任务。本项目开展的研究以商业和公共部门中声音和音频背景识别技术的广泛潜在受益者以及此类技术的用户和从业人员为目标。除了声学场景分析,我们相信这种新方法将推动更广泛的音频和声学领域,从而为相关领域创建上下文感知系统,包括音乐和语音技术以及助听器。
英文摘要
The amount of audio data being generated has dramatically increased over the past decade, spanning from user-generated content, recordings in audiovisual archives, to sensor data captured in urban, nature or domestic environments. The need to detect and identify sound events in environmental recordings (e.g. door knock, glass break) as well as to recognise the context of an audio recording (e.g. train station, meeting) has led to the emergence of a new field of research: acoustic scene analysis. Emerging applications of acoustic scene analysis include the development of sound recognition technologies for smart homes and smart cities, security/surveillance, audio retrieval and archiving, ambient assisted living, and automatic biodiversity assessment.However, current sound recognition technologies cannot adapt to different environments or situations (e.g. sound identification in an office environment, assuming specific room properties, working hours, outdoor noise and weather conditions). If information about context is available, it is typically characterised by a single label for an entire audio stream, not taking into account complex and ever-changing environments, for example when recording using hand-held devices, where context can consist of multiple time-varying factors and can be characterised by more than a single label. This project will address the aforementioned shortcomings by investigating and developing technologies for context-aware sound recognition. We assume that the context of an audio stream consists of several time-varying factors that can be viewed as a combination of different environments and situations; the ever-changing context in turn informs the types and properties of sounds to be recognised by the system. Methods for context and sound recognition will be investigated and developed, based on signal processing and machine learning theory. The main contribution of the project will be an algorithmic framework that jointly recognises audio-based context and sound events, applied to complex audio streams with several sound sources and time-varying environments. The proposed software framework will be evaluated using complex audio streams recorded in urban and domestic environments, as well as using simulated audio data in order to carefully control contextual and sound properties and have the benefit of accurate annotations. In order to further promote the study of context-aware sound recognition systems, a public evaluation task will be organised in conjunction with the public challenge on Detection and Classification of Acoustic Scenes and Events (DCASE). Research carried out in this project targets a wide range of potential beneficiaries in the commercial and public sector for sound and audio-based context recognition technologies, as well as users and practitioners of such technologies. Beyond acoustic scene analysis, we believe this new approach will advance the broader fields of audio and acoustics, leading to the creation of context-aware systems for related fields, including music and speech technology and hearing aids.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.33682/sm6r-8p49
发表时间: 2019-10
期刊:
影响因子: --
作者: [Arjun Pankajakshan;Helen L. Bear;Emmanouil Benetos;Events]
通讯作者: Arjun Pankajakshan;Helen L. Bear;Emmanouil Benetos;Events
Towards Joint Sound Scene and Polyphonic Sound Event Recognition
走向联合声音场景和和弦声音事件识别
DOI: 10.21437/interspeech.2019-2169
发表时间: 2019
期刊:
影响因子: --
作者: [Bear H]
通讯作者: Bear H
City Classification from Multiple Real-World Sound Scenes
根据多个真实世界声音场景进行城市分类
DOI: 10.1109/waspaa.2019.8937271
发表时间: 2019
期刊:
影响因子: --
作者: [Bear H]
通讯作者: Bear H
DOI: --
发表时间: 2018-09
期刊:
影响因子: --
作者: [Helen L. Bear;Emmanouil Benetos]
通讯作者: Helen L. Bear;Emmanouil Benetos
国内基金
海外基金
通用声场空间信息捡拾与重放方法的研究
  • 批准号:
    11174087
  • 项目类别:
    面上项目
  • 资助金额:
    70.0万元
  • 批准年份:
    2011
  • 负责人:
    谢菠荪
  • 依托单位: