III: Medium: Spatial Sound Scene Description
III: Medium: Spatial Sound Scene Description
批准号:
1955357
负责人:
Juan Bello
金额:
$99.99万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2024-06-30
中文摘要
声音包含了丰富的关于周围环境的信息。如果你闭着眼睛站在城市人行道上倾听,你会听到周围发生的事情的声音:鸟儿在叫,松鼠在跑,人们在说话,门在开,救护车在疾驰,卡车在空转。此外,您还可能能够感知到每个声源的位置、去向和移动速度。该项目将建立创新的技术,使计算机能够从声音中提取丰富的信息。通过识别存在哪些声源,并估计每个声源的空间位置和移动,声音传感技术将能够更好地描述具有麦克风功能的日常设备的环境,例如智能手机、耳机、智能扬声器、助听器、家庭摄像头和混合现实耳机。对于听力受损的人来说,开发的技术有可能提醒他们注意城市或家庭环境中的危险情况。对于城市机构来说,声学传感器将能够更准确地量化城市环境中的交通、建筑和其他活动。对于生态学家来说,这项技术可以帮助他们更准确地监测和研究野生动物。此外,这一信息补充了计算机视觉可以感知到的东西,因为声音可以包括关于不易看到的事件的信息,例如小来源(例如昆虫)、远处的来源(例如远处的手提钻)或简单地隐藏在另一个物体后面的来源(例如建筑物角落附近的救护车)。该项目还包括有100多名公立学校学生和教师参加的外联活动,以及对博士后、研究生和本科生的培训和指导。该项目将开发空间声音场景描述的计算模型:即根据生物和物体在真实环境中发出的声音来估计它们的类别、空间位置、方向和运动速度。研究人员的目标是,他们的解决方案在各种声音场景和传感条件下都是稳健的:噪音、稀疏、自然、城市、室内、室外、具有不同的源组成、具有未知源、具有移动的源、具有移动的传感器等。虽然当前的方法显示出希望,但它们在现实世界的条件下仍然远远不能稳健,因此无法支持上述任何场景。这些缺陷源于重要的数据问题,如缺乏空间注释的真实世界音频数据,以及过度依赖质量差、不切实际的合成数据;以及方法问题,如过度依赖监督学习和未能捕获解决方案空间的结构。该项目计划一种将创新的数据收集策略与尖端的机器学习解决方案相结合的方法。首先,提出了一种利用物理模型和生成模型对声景数据集进行概率合成的新框架。其目标是大幅增加强标签空间音频数据的数量、真实性和多样性。其次,它通过结合高质量的现场录音、众包、新颖的VR/AR多模式注释策略和公民科学家的大规模注释来收集和注释真实声音场景的新数据集。第三,提出了基于大量未标记音频数据训练的深度自监督表征学习策略。第四,这些表示模块与分层预测模型配对,其中分层的顶层/底层对应于场景描述的较粗/较细级别。最后,该项目包括与三个行业合作伙伴的合作,以探索建议的解决方案所支持的应用。该项目将产生新的方法和开放源码软件库,用于空间声音场景生成、注释、表示学习和声音事件检测/定位/跟踪;以及空间音频录音、空间声音场景注释、合成孤立声音和合成空间声音场景的新开放数据集。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Sound is rich with information about the surrounding environment. If you stand on a city sidewalk with your eyes closed and listen, you will hear the sounds of events happening around you: birds chirping, squirrels scurrying, people talking, doors opening, an ambulance speeding, a truck idling. In addition, you will also likely be able to perceive the location of each sound source, where it’s going, and how fast it’s moving. This project will build innovative technologies to allow computers to extract this rich information out of sound. By not only identifying which sound sources are present but also estimating the spatial location and movement of each sound source, sound sensing technology will be able to better describe our environments with microphone-enabled everyday devices, e.g. smartphones, headphones, smart speakers, hearing-aids, home camera, and mixed-reality headsets. For hearing impaired individuals, the developed technologies have the potential to alert them to dangerous situations in urban or domestic environments. For city agencies, acoustic sensors will be able to more accurately quantify traffic, construction, and other activities in urban environments. For ecologists, this technology can help them more accurately monitor and study wildlife. In addition, this information complements what computer vision can sense, as sound can include information about events that are not easily visible, such as sources that are small (e.g., insects), far away (e.g., a distant jackhammer), or simply hidden behind another object (e.g., an incoming ambulance around a building's corner). This project also includes outreach activities involving over 100 public school students and teachers, as well as the training and mentoring of postdoctoral, graduate and undergraduate students. This project will develop computational models for spatial sound scene description: that is, estimating the class, spatial location, direction and speed of movement of living beings and objects in real environments by the sounds they make. The investigators aim for their solutions to be robust across a wide range of sound scenes and sensing conditions: noisy, sparse, natural, urban, indoors, outdoors, with varying compositions of sources, with unknown sources, with moving sources, with moving sensors, etc. While current approaches show promise, they are still far from robust in real-world conditions and thus unable to support any of the above scenarios. These shortcomings stem from important data issues such as a lack of spatially annotated real-world audio data, and an over-reliance on poor quality, unrealistic synthesized data; as well as methodological issues such as excessive dependence on supervised learning and a failure to capture the structure of the solution space. This project plans an approach mixing innovative data collection strategies with cutting-edge machine learning solutions. First, it advances a novel framework for the probabilistic synthesis of soundscape datasets using physical and generative models. The goal is to substantially increase the amount, realism and diversity of strongly-labeled spatial audio data. Second, it collects and annotates new datasets of real sound scenes via a combination of high-quality field recordings, crowdsourcing, novel VR/AR multimodal annotation strategies and large-scale annotation by citizen scientists. Third, it puts forward novel deep self-supervised representation learning strategies trained on vast quantities of unlabeled audio data. Fourth, these representation modules are paired with hierarchical predictive models, where the top/bottom levels of the hierarchy correspond to coarser/finer levels of scene description. Finally, the project includes collaborations with three industrial partners to explore applications enabled by the proposed solutions. The project will result in novel methods and open source software libraries for spatial sound scene generation, annotation, representation learning, and sound event detection/localization/tracking; and new open datasets of spatial audio recordings, spatial sound scene annotations, synthesized isolated sounds, and synthesized spatial soundscapes.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/icassp43922.2022.9747669
发表时间:
2021-10
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Ho-Hsiang Wu;Prem Seetharaman;Kundan Kumar;J. Bello]
通讯作者:
Ho-Hsiang Wu;Prem Seetharaman;Kundan Kumar;J. Bello
Sound Event Detection in Urban Audio with Single and Multi-Rate Pcen
使用单速率和多速率 Pcen 进行城市音频中的声音事件检测
DOI:
10.1109/icassp39728.2021.9414697
发表时间:
2021
期刊:
2021
影响因子:
--
作者:
[Ick, Christopher, McFee, Brian]
通讯作者:
McFee, Brian
Few-Shot Musical Source Separation
少镜头音乐源分离
DOI:
10.1109/icassp43922.2022.9747536
发表时间:
2022
期刊:
Speech and Signal Processing (ICASSP
影响因子:
--
作者:
[Wang, Yu, Stoller, Daniel, Bittner, Rachel M., Pablo Bello, Juan]
通讯作者:
Pablo Bello, Juan
Micarraylib: Software for the Reproducible Aggregation, Standardization, and Signal Processing of Microphone Array Datasets. Detection and Classification of Acoustic Scenes and Events
Micarraylib:用于麦克风阵列数据集的可重复聚合、标准化和信号处理的软件。
DOI:
--
发表时间:
2021
期刊:
2021
影响因子:
--
作者:
[Roman, I. R., Bello, J.P.]
通讯作者:
Bello, J.P.
Analyzing the Effect of Equal-Angle Spatial Discretization on Sound Event Localization and Detection
分析等角空间离散对声音事件定位和检测的影响
DOI:
--
发表时间:
2022
期刊:
Proceedings of the 7th Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022
影响因子:
--
作者:
[Kushwaha, S.S.]
通讯作者:
Kushwaha, S.S.
共 9 条
PFI-TT: Acoustic Continuous Condition Monitoring of Manufacturing Machinery
-
批准号:1827523
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2018
-
负责人:Juan Bello
-
依托单位:
I-Corps: Embedded Machine Listening for Smart Acoustic Monitoring
-
批准号:1759592
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2017
-
负责人:Juan Bello
-
依托单位:
BIGDATA: Collaborative Research: IA: BirdVox: Automating Acoustic Monitoring of Migrating Bird Species
-
批准号:1633259
-
项目类别:Standard Grant
-
资助金额:$61.24万
-
财政年份:2016
-
负责人:Juan Bello
-
依托单位:
CPS: Frontier: SONYC: A Cyber-Physical System for Monitoring, Analysis and Mitigation of Urban Noise Pollution
-
批准号:1544753
-
项目类别:Continuing Grant
-
资助金额:$462.82万
-
财政年份:2016
-
负责人:Juan Bello
-
依托单位:
CAREER: Analyzing the Sequential Structure of Music Audio
-
批准号:0844654
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2009
-
负责人:Juan Bello
-
依托单位:
海外基金