Audio-to-score alignment of piano music using RNN-based automatic music transcription

Audio-to-score alignment of piano music using RNN-based automatic music transcription
复制标题

使用基于 RNN 的自动音乐转录进行钢琴音乐的音频与乐谱对齐

DOI:
--
复制
发表时间:
2017
期刊:
ArXiv
影响因子:
--
通讯作者:
Juhan Nam
Juhan Nam
中科院分区:
--
文献类型:
--
作者:
Taegyun Kwon;Dasaem Jeong;Juhan Nam

文献摘要

被引文献

相似文献

我们提出了一个使用神经网络的自动音乐转录(AMT)的钢琴演奏音频与乐谱对齐框架。尽管AMT结果可能包含一些错误,但音符预测输出可以被视为与MIDI音符或色度表示直接比较的学习特征表示。为此,我们使用两个递归神经网络作为基于amt的特征提取器来对齐算法。一个预测帧级88个音符或12色度的存在,另一个检测12色度的音符开始。我们将两种类型的学习特征结合起来进行音频-乐谱对齐。为了便于比较,我们采用动态时间翘曲作为对齐算法,无需任何额外的后处理。我们在MAPS数据集上评估了所提出的框架,并将其与以前的工作进行了比较。结果表明,基于学习特征的对准框架显著提高了对准精度,平均起始误差小于10 ms。
We propose a framework for audio-to-score alignment on piano performance that employs automatic music transcription (AMT) using neural networks. Even though the AMT result may contain some errors, the note prediction output can be regarded as a learned feature representation that is directly comparable to MIDI note or chroma representation. To this end, we employ two recurrent neural networks that work as the AMT-based feature extractors to the alignment algorithm. One predicts the presence of 88 notes or 12 chroma in frame-level and the other detects note onsets in 12 chroma. We combine the two types of learned features for the audio-to-score alignment. For comparability, we apply dynamic time warping as an alignment algorithm without any additional post-processing. We evaluate the proposed framework on the MAPS dataset and compare it to previous work. The result shows that the alignment framework with the learned features significantly improves the accuracy, achieving less than 10 ms in mean onset error.