An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection

An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
复制标题

DOI:
10.1109/icassp39728.2021.9413473
复制
发表时间:
2020-10
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yin Cao;Turab Iqbal;Qiuqiang Kong;Y. Zhong;Wenwu Wang;Mark D. Plumbley
Yin Cao;Turab Iqbal;Qiuqiang Kong;Y. Zhong;Wenwu Wang;Mark D. Plumbley
中科院分区:
其他
文献类型:
--
作者:
Yin Cao;Turab Iqbal;Qiuqiang Kong;Y. Zhong;Wenwu Wang;Mark D. Plumbley

文献摘要

相似文献

复调声音事件定位与检测(SELD)联合执行声音事件检测(SED)和到达方向(DoA)估计,同时检测声音事件的类型和发生时间以及它们对应的DoA角度。我们从多任务学习的角度研究SELD任务。本文讨论了两个开放的问题。首先,为了检测相同类型但具有不同DoA的重叠声音事件,我们建议使用逐轨输出格式并使用置换不变训练解决伴随的轨道置换问题。进一步使用多头自注意来分离轨道。其次,之前的发现是,与单独学习子任务相比,通过使用硬参数共享,SELD会遭受性能损失。这是解决了软参数共享计划。我们将所提出的方法称为事件独立网络V2(EINV2),它是我们以前提出的方法的改进版本,也是SELD的端到端网络。我们表明,我们提出的EINV2联合SED和DoA估计优于以前的方法的一个很大的利润率,并具有可比的性能,以国家的最先进的合奏模型。
Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA angles simultaneously. We study the SELD task from a multi-task learning perspective. Two open problems are addressed in this paper. Firstly, to detect overlapping sound events of the same type but with different DoAs, we propose to use a trackwise output format and solve the accompanying track permutation problem with permutation-invariant training. Multi-head self-attention is further used to separate tracks. Secondly, a previous finding is that, by using hard parameter-sharing, SELD suffers from a performance loss compared with learning the subtasks separately. This is solved by a soft parameter-sharing scheme. We term the proposed method as Event Independent Network V2 (EINV2), which is an improved version of our previously-proposed method and an end-to-end network for SELD. We show that our proposed EINV2 for joint SED and DoA estimation outperforms previous methods by a large margin, and has comparable performance to state-of-the-art ensemble models.