Zero-Mean Convolutions for Level-Invariant Singing Voice Detection

Zero-Mean Convolutions for Level-Invariant Singing Voice Detection
复制标题

用于电平不变歌声检测的零均值卷积

DOI:
--
复制
发表时间:
2018
期刊:
International Society for Music Information Retrieval Conference
影响因子:
--
通讯作者:
Bernhard Lehner
Bernhard Lehner
中科院分区:
--
文献类型:
--
作者:
Jan Schlüter;Bernhard Lehner

文献摘要

参考文献

被引文献

相似文献

最先进的歌唱声音检测器是基于在注释示例上训练的分类器。正如最近所显示的那样,这种检测器有一个重要的弱点:由于唱歌的声音与训练数据中的声级相关,分类器学会对输入的大小变得敏感,并对不同声级的相同信号给出不同的预测。从对数量级mel谱图摘录上训练的卷积神经网络(CNN)开始,我们通过强制每个第一层卷积滤波器为零均值来消除这种依赖关系——也就是说,使其系数之和为零。与我们在大规模公共数据集上评估的其他四种方法(数据增强、实例归一化、频谱δ特征和每通道能量归一化(PCEN))相比,零均值卷积实现了完美的声级不变性,而不会对预测精度或计算需求产生任何影响。我们假设零均值卷积对于其他需要鲁棒性以适应水平变化的机器监听任务是有用的。
State-of-the-art singing voice detectors are based on clas-sifiers trained on annotated examples. As recently shown, such detectors have an important weakness: Since singing voice is correlated with sound level in training data, clas-sifiers learn to become sensitive to input magnitude, and give different predictions for the same signal at different sound levels. Starting from a Convolutional Neural Network (CNN) trained on logarithmic-magnitude mel spectrogram excerpts, we eliminate this dependency by forcing each first-layer convolutional filter to be zero-mean – that is, to have its coefficients sum to zero. In contrast to four other methods – data augmentation, instance normalization, spectral delta features, and per-channel energy normalization (PCEN) – that we evaluated on a large-scale public dataset, zero-mean convolutions achieve perfect sound level invariance without any impact on prediction accuracy or computational requirements. We assume that zero-mean convolutions would be useful for other machine listening tasks requiring robustness to level changes.
MedleyDB:用于注释密集型 MIR 研究的多轨数据集
DOI: --
发表时间: 2014
期刊: 15th International Society for Music Information Retrieval Conference
影响因子: --
作者:
Bittner, R.
通讯作者: Bittner, R.