Polyphonic Sound Event Tracking Using Linear Dynamical Systems

Polyphonic Sound Event Tracking Using Linear Dynamical Systems
复制标题

DOI:
10.1109/taslp.2017.2690576
复制
发表时间:
2017-06
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Emmanouil Benetos;G. Lafay;M. Lagrange;Mark D. Plumbley
Emmanouil Benetos;G. Lafay;M. Lagrange;Mark D. Plumbley
中科院分区:
其他
文献类型:
--
作者:
Emmanouil Benetos;G. Lafay;M. Lagrange;Mark D. Plumbley

文献摘要

被引文献

相似文献

提出了一种基于谱图分解技术和状态空间模型的多音事件检测与跟踪系统。该系统扩展了概率潜在成分分析(PLCA),并围绕频率、声音事件类别、样本索引和声音状态的四维频谱模板字典进行建模。为了在一段时间内联合跟踪多个重叠的声音事件,提出了在PLCA推理中集成线性动力系统(LDS)。该系统假设PLCA声音事件激活是LDS中的(噪声)观测,其潜在状态对应于真实事件激活。LDS培训是利用充分观察到的数据,利用基于PLCA的模型产生的地面真相信息事件激活来实现的。使用从声学场景模拟器生成的办公室声音的复调数据集以及用于比较目的的真实和合成单音数据集,对几种LDS变体进行了评估。结果表明,与单独使用PLCA模型相比,在基于帧的F度量方面,将LDS跟踪集成到PLCA中导致了+8.5-10.5%的改进。此外,在多音事件检测方面,该系统的性能优于几种最先进的方法。
In this paper, a system for polyphonic sound event detection and tracking is proposed, based on spectrogram factorization techniques and state space models. The system extends probabilistic latent component analysis (PLCA) and is modeled around a four-dimensional spectral template dictionary of frequency, sound event class, exemplar index, and sound state. In order to jointly track multiple overlapping sound events over time, the integration of linear dynamical systems (LDS) within the PLCA inference is proposed. The system assumes that the PLCA sound event activation is the (noisy) observation in an LDS, with the latent states corresponding to the true event activations. LDS training is achieved using fully observed data, making use of ground truth-informed event activations produced by the PLCA-based model. Several LDS variants are evaluated, using polyphonic datasets of office sounds generated from an acoustic scene simulator, as well as real and synthesized monophonic datasets for comparative purposes. Results show that the integration of LDS tracking within PLCA leads to an improvement of +8.5–10.5% in terms of frame-based F-measure as compared to the use of the PLCA model alone. In addition, the proposed system outperforms several state-of-the-art methods for the task of polyphonic sound event detection.