Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms

Global Structure-Aware Drum Transcription Based on Self-Attention Mechanisms
复制标题

DOI:
10.3390/signals2030031
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Ryoto Ishizuka;Ryo Nishikimi;Kazuyoshi Yoshii
Ryoto Ishizuka;Ryo Nishikimi;Kazuyoshi Yoshii
中科院分区:
其他
文献类型:
--
作者:
Ryoto Ishizuka;Ryo Nishikimi;Kazuyoshi Yoshii

文献摘要

相似文献

本文介绍了一种自动鼓转录(ADT)方法,直接估计一个tatum级鼓得分从音乐信号相比,大多数传统的ADT方法,估计帧级的开始概率鼓。为了估计一个tatum级别的分数,我们提出了一个深度转录模型,该模型由一个帧级编码器和一个tatum级别解码器组成,帧级编码器用于从音乐信号中提取潜在特征,tatum级别解码器用于从tatum级别汇集的潜在特征中估计鼓分数。为了捕捉鼓乐谱的全局重复结构,这是很难学习的递归神经网络(RNN),我们引入了一个自注意机制与tatum同步位置编码到解码器。为了减轻训练基于自我注意力的模型的困难,从配对数据的数量不足,并提高音乐的自然性估计的分数,我们提出了一种正则化的训练方法,使用一个全局结构感知的掩蔽语言(分数)模型与自我注意力机制预训练从广泛收集的鼓分数。实验结果表明,提出的正则化模型优于传统的基于RNN的模型在tatum级的错误率和帧级的F-测量,即使只有有限数量的配对数据是可用的,使非正则化模型表现不佳的基于RNN的模型。
This paper describes an automatic drum transcription (ADT) method that directly estimates a tatum-level drum score from a music signal in contrast to most conventional ADT methods that estimate the frame-level onset probabilities of drums. To estimate a tatum-level score, we propose a deep transcription model that consists of a frame-level encoder for extracting the latent features from a music signal and a tatum-level decoder for estimating a drum score from the latent features pooled at the tatum level. To capture the global repetitive structure of drum scores, which is difficult to learn with a recurrent neural network (RNN), we introduce a self-attention mechanism with tatum-synchronous positional encoding into the decoder. To mitigate the difficulty of training the self-attention-based model from an insufficient amount of paired data and to improve the musical naturalness of the estimated scores, we propose a regularized training method that uses a global structure-aware masked language (score) model with a self-attention mechanism pretrained from an extensive collection of drum scores. The experimental results showed that the proposed regularized model outperformed the conventional RNN-based model in terms of the tatum-level error rate and the frame-level F-measure, even when only a limited amount of paired data was available so that the non-regularized model underperformed the RNN-based model.