Multi-channel Environmental Sound Segmentation utilizing Sound Source Localization and Separation U-Net

Multi-channel Environmental Sound Segmentation utilizing Sound Source Localization and Separation U-Net
复制标题

DOI:
10.1109/ieeeconf49454.2021.9382730
复制
发表时间:
2021-01
期刊:
2021 IEEE/SICE International Symposium on System Integration (SII)
影响因子:
--
通讯作者:
Yui Sudo;Katsutoshi Itoyama;Kenji Nishida;K. Nakadai
Yui Sudo;Katsutoshi Itoyama;Kenji Nishida;K. Nakadai
中科院分区:
其他
文献类型:
--
作者:
Yui Sudo;Katsutoshi Itoyama;Kenji Nishida;K. Nakadai

文献摘要

相似文献

提出了一种多通道环境声分割方法。环境声音分割是一种集声源定位、声源分离和类别识别于一体的综合方法。当多个麦克风可用时,空间特征可以用于提高来自不同方向的信号的分离精度;然而,常规方法具有两个缺点:(a)由于使用空间特征的声源定位和声源分离以及使用频谱特征的类别识别是在同一个神经网络中训练的,所以它过度拟合到达方向和类别之间的关系。(b)虽然语音识别中使用的排列不变训练可以扩展,但由于最大说话人数量的限制,它对环境声音不实用。本文提出了一种多通道环境声音分割方法,该方法结合了同时进行声源定位和声源分离的U-Net和对分离出的声音进行分类的卷积神经网络。此方法可防止对到达方向和类别之间的关系进行过拟合。使用75类环境声音数据集进行仿真实验,结果表明,该方法的均方根误差低于传统方法。
This paper proposes a multi-channel environmental sound segmentation method. Environmental sound segmentation is an integrated method that deals with sound source localization, sound source separation and class identification. When multiple microphones are available, spatial features can be used to improve the separation accuracy of signals from different directions; however, conventional methods have two drawbacks: (a) Since sound source localization and sound source separation using spatial features and class identification using spectral features are trained in the same neural network, it overfits to the relationship between the direction of arrival and the class. (b) Although the permutation invariant training used in speech recognition could be extended, it is not practical for environmental sounds due to the maximum number of speakers limitation. This paper proposes multi-channel environmental sound segmentation method that combines U-Net which simultaneously performs sound source localization and sound source separation, and convolutional neural network which classifies the separated sounds. This method prevents overfitting to the relationship between the direction of arrival and the class. Simulation experiments using the created datasets including 75-class environmental sounds showed that the root mean squared error of the proposed method was lower than that of the conventional method.