Artificially Synthesising Data for Audio Classification and Segmentation to Improve Speech and Music Detection in Radio Broadcast

Artificially Synthesising Data for Audio Classification and Segmentation to Improve Speech and Music Detection in Radio Broadcast
复制标题

DOI:
10.1109/icassp39728.2021.9413597
复制
发表时间:
2021-02
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
S. Venkatesh;D. Moffat;Alexis Kirke;Gözel Shakeri;S. Brewster;J. Fachner;Helen Odell-Miller;Alexander J. Street;Nicolas Farina;Sube Banerjee;E. Miranda
S. Venkatesh;D. Moffat;Alexis Kirke;Gözel Shakeri;S. Brewster;J. Fachner;Helen Odell-Miller;Alexander J. Street;Nicolas Farina;Sube Banerjee;E. Miranda
中科院分区:
其他
文献类型:
--
作者:
S. Venkatesh;D. Moffat;Alexis Kirke;Gözel Shakeri;S. Brewster;J. Fachner;Helen Odell-Miller;Alexander J. Street;Nicolas Farina;Sube Banerjee;E. Miranda

文献摘要

被引文献

相似文献

将音频分割成音乐和语音等同质部分有助于我们理解音频的内容。它是一个有用的预处理步骤,索引,存储和修改音频记录,广播和电视节目。用于分割的深度学习模型通常是在有版权的材料上训练的,这些材料是不能共享的。对这些数据集进行注释既耗时又昂贵,因此,它大大减缓了研究进展。在这项研究中,我们提出了一种新的程序,人工合成类似无线电信号的数据。我们复制了一个电台DJ的工作流程,在混合音频和调查参数,如衰减曲线和音频躲避。我们在这些合成数据上训练了一个卷积循环神经网络(CRNN),并在音乐语音检测方面优于最先进的算法。本文演示了数据合成过程作为一种高效的技术来生成大型数据集来训练用于音频分割的深度神经网络。
Segmenting audio into homogeneous sections such as music and speech helps us understand the content of audio. It is useful as a pre-processing step to index, store, and modify audio recordings, radio broadcasts and TV programmes. Deep learning models for segmentation are generally trained on copyrighted material, which cannot be shared. Annotating these datasets is time-consuming and expensive and therefore, it significantly slows down research progress. In this study, we present a novel procedure that artificially synthesises data that resembles radio signals. We replicate the workflow of a radio DJ in mixing audio and investigate parameters like fade curves and audio ducking. We trained a Convolutional Recurrent Neural Network (CRNN) on this synthesised data and outperformed state-of-the-art algorithms for music-speech detection. This paper demonstrates the data synthesis procedure as a highly effective technique to generate large datasets to train deep neural networks for audio segmentation.