Downbeat Tracking with Tempo-Invariant Convolutional Neural Networks

Downbeat Tracking with Tempo-Invariant Convolutional Neural Networks
复制标题

使用速度不变卷积神经网络进行悲观跟踪

DOI:
--
复制
发表时间:
2021
期刊:
International Society for Music Information Retrieval Conference
影响因子:
--
通讯作者:
M. Levy
M. Levy
中科院分区:
--
文献类型:
--
作者:
Bruno Di Giorgi;Matthias Mauch;M. Levy

文献摘要

被引文献

相似文献

人类追踪音乐重拍的能力对克里思的变化是稳健的,并且它延伸到以前从未遇到过的节奏。我们提出了一种确定性的时间扭曲操作,通过允许网络独立于克里思学习节奏模式,在卷积神经网络(CNN)中实现这种技能。与传统的深度学习方法不同,传统的深度学习方法在训练数据集中学习节奏模式,在我们的模型中学习的模式是节奏不变的,从而更好地概括克里思,更有效地利用网络容量。我们在一个合成数据集上测试泛化属性,该合成数据集是通过使用FluidSynth渲染Groove Rounds Dataset创建的,分为一个包含原始性能的训练集和一个包含使用不同SoundFonts渲染的节奏缩放版本的测试集(测试时增强)。所提出的模型几乎完美地概括了看不见的tempi(训练集和测试集的F测量值均为0.89),而可比较的传统CNN仅在训练集(0.89)上达到了类似的精度,在测试集上下降到0.54。所提出的模型的泛化优势扩展到真实的音乐,如GTZAN和舞厅数据集上的结果所示。
The human ability to track musical downbeats is robust to changes in tempo, and it extends to tempi never previously encountered. We propose a deterministic time-warping operation that enables this skill in a convolutional neural network (CNN) by allowing the network to learn rhythmic patterns independently of tempo. Unlike conventional deep learning approaches, which learn rhythmic patterns at the tempi present in the training dataset, the patterns learned in our model are tempo-invariant, leading to better tempo generalisation and more efficient usage of the network capacity. We test the generalisation property on a synthetic dataset created by rendering the Groove MIDI Dataset using FluidSynth, split into a training set containing the original performances and a test set containing tempo-scaled versions rendered with different SoundFonts (test-time augmentation). The proposed model generalises nearly perfectly to unseen tempi (F-measure of 0.89 on both training and test sets), whereas a comparable conventional CNN achieves similar accuracy only for the training set (0.89) and drops to 0.54 on the test set. The generalisation advantage of the proposed model extends to real music, as shown by results on the GTZAN and Ballroom datasets.