Automatic Piano Fingering from Partially Annotated Scores using Autoregressive Neural Networks

Automatic Piano Fingering from Partially Annotated Scores using Autoregressive Neural Networks
复制标题

DOI:
10.1145/3503161.3548372
复制
发表时间:
2022-10
期刊:
Proceedings of the 30th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Pedro Ramoneda;Dasaem Jeong;Eita Nakamura;Xavier Serra;M. Miron
Pedro Ramoneda;Dasaem Jeong;Eita Nakamura;Xavier Serra;M. Miron
中科院分区:
其他
文献类型:
--
作者:
Pedro Ramoneda;Dasaem Jeong;Eita Nakamura;Xavier Serra;M. Miron

文献摘要

相似文献

钢琴指法是一项创造性和高度个性化的任务,音乐家在他们最初的音乐教育中逐渐获得。钢琴家必须学会选择手指的顺序来弹奏钢琴键,因为乐谱不像其他技术元素那样刻有手指和手的动作。基于由多个专家注释者完全注释的150个乐谱摘录组成的先前数据集,已经对自动钢琴指法进行了许多研究工作。然而,大多数钢琴乐谱包括部分注释的问题手指和手的动作。我们为该任务引入了一个新的数据集,即ThumbSet数据集,其中包含2523个片段,其中包含来自非专业注释者的部分和嘈杂的钢琴指法注释。作为我们方法的一部分,我们提出了两个自回归神经网络与波束搜索解码建模自动钢琴指法作为一个序列到序列的学习问题,考虑到输出手指标签之间的相关性。我们设计了第一个模型,与之前的建议的精确音高表示。第二个模型使用图神经网络来更有效地表示复调,复调的处理在以前的研究中一直是一个常见的问题。最后,在已有的专家标注数据集上对模型进行微调。评估表明:(1)我们能够在ThumbSet数据集上训练时获得高性能;(2)所提出的模型优于最先进的隐马尔可夫模型和递归神经网络基线。提供代码、数据集、模型和结果,以增强任务的可重复性,包括用于评估的新框架。
Piano fingering is a creative and highly individualised task acquired by musicians progressively in their first music education years. Pianists must learn to choose the order of fingers to play the piano keys because scores do not have engraved finger and hand movements as other technique elements. Numerous research efforts have been conducted for automatic piano fingering based on a previous dataset composed of 150 score excerpts fully annotated by multiple expert annotators. However, most piano sheets include partial annotations for problematic finger and hand movements. We introduce a novel dataset for the task, the ThumbSet dataset, containing 2523 pieces with partial and noisy annotations of piano fingering crowdsourced from non-expert annotators. As part of our methodology, we propose two autoregressive neural networks with beam search decoding for modelling automatic piano fingering as a sequence-to-sequence learning problem, considering the correlation between output finger labels. We design the first model with the exact pitch representation of previous proposals. The second model uses graph neural networks to more effectively represent polyphony, whose treatment has been a common issue across previous studies. Finally, we finetune the models on the existing expert annotations dataset. The evaluation shows that (1) we are able to achieve high performance when training on the ThumbSet dataset and that (2) the proposed models outperform the state-of-the-art hidden Markov models and recurrent neural network baselines. Code, dataset, models, and results are made available to enhance the task reproducibility, including a new framework for evaluation.