Event-Independent Network for Polyphonic Sound Event Localization and Detection

Event-Independent Network for Polyphonic Sound Event Localization and Detection
复制标题

DOI:
--
复制
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Yin Cao;Turab Iqbal;Qiuqiang Kong;Yue Zhong;Wenwu Wang;M. Plumbley
Yin Cao;Turab Iqbal;Qiuqiang Kong;Yue Zhong;Wenwu Wang;M. Plumbley
中科院分区:
其他
文献类型:
--
作者:
Yin Cao;Turab Iqbal;Qiuqiang Kong;Yue Zhong;Wenwu Wang;M. Plumbley

文献摘要

相似文献

复调声音事件定位与检测不仅仅是检测正在发生的声音事件,而是定位相应的声源。这一系列任务首次在DCASE 2019任务3中引入。在2020年,声音事件定位和检测任务在移动声源和重复事件情况下引入了额外的挑战,其中包括具有两个不同到达方向(DoA)角度的两个相同类型的事件。本文提出了一种新的事件无关网络用于多音事件定位和检测。与我们在DCASE 2019任务3中提出的两阶段方法不同,这个新网络是完全端到端的。网络的输入是一阶Ambisonics(FOA)时域信号,然后将其馈送到1-D卷积层以提取声学特征。然后,网络被分成两个并行的分支。第一分支用于声音事件检测(SED),并且第二分支用于DoA估计。存在来自网络的三种类型的预测,SED预测、DoA预测和事件活动检测(EAD)预测,其用于将SED和DoA特征联合收割机组合以用于开始和偏移估计。所有这些预测都具有两个轨迹的格式,表明最多有两个重叠的事件。在每一个轨迹中,最多只能有一个事件发生。这种架构引入了轨道排列的问题。为了解决这个问题,使用帧级置换不变训练方法。实验结果表明,该方法可以检测出复调声音事件及其对应的DOA。在Task 3数据集上的性能与基线方法相比有了很大的提高。
Polyphonic sound event localization and detection is not only detecting what sound events are happening but localizing corresponding sound sources. This series of tasks was first introduced in DCASE 2019 Task 3. In 2020, the sound event localization and detection task introduces additional challenges in moving sound sources and overlapping-event cases, which include two events of the same type with two different direction-of-arrival (DoA) angles. In this paper, a novel event-independent network for polyphonic sound event localization and detection is proposed. Unlike the two-stage method we proposed in DCASE 2019 Task 3, this new network is fully end-to-end. Inputs to the network are first-order Ambisonics (FOA) time-domain signals, which are then fed into a 1-D convolutional layer to extract acoustic features. The network is then split into two parallel branches. The first branch is for sound event detection (SED), and the second branch is for DoA estimation. There are three types of predictions from the network, SED predictions, DoA predictions, and event activity detection (EAD) predictions that are used to combine the SED and DoA features for on-set and off-set estimation. All of these predictions have the format of two tracks indicating that there are at most two overlapping events. Within each track, there could be at most one event happening. This architecture introduces a problem of track permutation. To address this problem, a frame-level permutation invariant training method is used. Experimental results show that the proposed method can detect polyphonic sound events and their corresponding DoAs. Its performance on the Task 3 dataset is greatly increased as compared with that of the baseline method.