Zero-Mean Convolutions for Level-Invariant Singing Voice Detection
Zero-Mean Convolutions for Level-Invariant Singing Voice Detection
复制标题
用于电平不变歌声检测的零均值卷积
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Bernhard Lehner
中科院分区:
文献类型:
--
作者:
Jan Schlüter;Bernhard Lehner
State-of-the-art singing voice detectors are based on clas-sifiers trained on annotated examples. As recently shown, such detectors have an important weakness: Since singing voice is correlated with sound level in training data, clas-sifiers learn to become sensitive to input magnitude, and give different predictions for the same signal at different sound levels. Starting from a Convolutional Neural Network (CNN) trained on logarithmic-magnitude mel spectrogram excerpts, we eliminate this dependency by forcing each first-layer convolutional filter to be zero-mean – that is, to have its coefficients sum to zero. In contrast to four other methods – data augmentation, instance normalization, spectral delta features, and per-channel energy normalization (PCEN) – that we evaluated on a large-scale public dataset, zero-mean convolutions achieve perfect sound level invariance without any impact on prediction accuracy or computational requirements. We assume that zero-mean convolutions would be useful for other machine listening tasks requiring robustness to level changes.
DOI:
--
发表时间:
2014
期刊:
15th International Society for Music Information Retrieval Conference
影响因子:
--
作者:
Bittner, R.
通讯作者:
Bittner, R.