Scaper: A library for soundscape synthesis and augmentation

Scaper: A library for soundscape synthesis and augmentation
复制标题

DOI:
10.1109/waspaa.2017.8170052
复制
发表时间:
2017-10
期刊:
2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
影响因子:
--
通讯作者:
J. Salamon;D. MacConnell;M. Cartwright;P. Li;J. Bello
J. Salamon;D. MacConnell;M. Cartwright;P. Li;J. Bello
中科院分区:
其他
文献类型:
--
作者:
J. Salamon;D. MacConnell;M. Cartwright;P. Li;J. Bello

文献摘要

被引文献

相似文献

环境记录中的声音事件检测(SED)是机器听力研究的一个关键课题,应用于智能城市的噪声监测、自动驾驶汽车、监控、生物声学监测以及大型多媒体集合的索引。为SED开发新的解决方案通常依赖于强标记音频记录的可用性,其中注释包括每个事件的开始,偏移和来源。手动生成这种精确的注释非常耗时,因此具有强标签的SED的现有数据集非常稀少且大小有限。为了解决这个问题,我们提出了Scaper,一个用于音景合成和增强的开源库。给定一组孤立的声音事件,Scaper充当高级音序器,可以从单个概率定义的“规范”生成多个音景。为了增加输出的可变性,Scaper支持对每个事件单独应用音频转换,例如音高转换和时间拉伸。为了说明该库的潜力,我们生成了一个包含10,000个音景的数据集,并使用它来比较两种最先进算法的性能,包括按音景特征进行的细分。我们还描述了如何使用Scaper来生成音频刺激的音频标签众包实验,并结束了讨论Scaper的局限性和潜在的应用。
Sound event detection (SED) in environmental recordings is a key topic of research in machine listening, with applications in noise monitoring for smart cities, self-driving cars, surveillance, bioa-coustic monitoring, and indexing of large multimedia collections. Developing new solutions for SED often relies on the availability of strongly labeled audio recordings, where the annotation includes the onset, offset and source of every event. Generating such precise annotations manually is very time consuming, and as a result existing datasets for SED with strong labels are scarce and limited in size. To address this issue, we present Scaper, an open-source library for soundscape synthesis and augmentation. Given a collection of iso-lated sound events, Scaper acts as a high-level sequencer that can generate multiple soundscapes from a single, probabilistically defined, “specification”. To increase the variability of the output, Scaper supports the application of audio transformations such as pitch shifting and time stretching individually to every event. To illustrate the potential of the library, we generate a dataset of 10,000 sound-scapes and use it to compare the performance of two state-of-the-art algorithms, including a breakdown by soundscape characteristics. We also describe how Scaper was used to generate audio stimuli for an audio labeling crowdsourcing experiment, and conclude with a discussion of Scaper's limitations and potential applications.