Joint Estimation of Note Values and Voices for Audio-to-Score Piano Transcription

Joint Estimation of Note Values and Voices for Audio-to-Score Piano Transcription
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Yuki Hiramatsu;Eita Nakamura;Kazuyoshi Yoshii
Yuki Hiramatsu;Eita Nakamura;Kazuyoshi Yoshii
中科院分区:
其他
文献类型:
--
作者:
Yuki Hiramatsu;Eita Nakamura;Kazuyoshi Yoshii

文献摘要

相似文献

本文介绍了一个国家的最先进的自动钢琴转录(APT)系统,可以转录一个人类可读的符号乐谱从钢琴录音的一个重要改进。虽然由于深度学习的最新进展,音符的音高和起始时间的估计已经得到了极大的改善,但作为APT系统的关键组成部分的音符值和语音标签的估计仍然是一项具有挑战性的任务。先前的研究已经揭示了(i)音符的音高和起始时间是有用的,但是所执行的音符持续时间对于估计音符值而言信息较少,并且(ii)音符值和声音具有相互依赖性。因此,我们提出了一个双向的长短期记忆网络,共同估计注意到的值和语音标签从注意音高和发病时间提前估计。为了提高对输入数据中包含的克里思错误、额外音符和缺失音符的鲁棒性,我们研究了数据增强。实验结果表明了多任务学习和数据增强的有效性,所提出的方法比现有方法获得了更好的准确性。
This paper describes an essential improvement of a state-of-the-art automatic piano transcription (APT) system that can transcribe a human-readable symbolic musical score from a piano recording. Whereas estimation of the pitches and onset times of musical notes has been improved drastically thanks to the recent advances of deep learning, estimation of note values and voice labels, which is a crucial component of the APT system, still remains a challenging task. A previous study has revealed that (i) the pitches and onset times of notes are useful but the performed note durations are less informative for estimating the note values and that (ii) the note values and voices have mutual dependency. We thus propose a bidirectional long short-term memory network that jointly estimates note values and voice labels from note pitches and onset times estimated in advance. To improve the robustness against tempo errors, extra notes, and missing notes included in the input data, we investigate data augmentation. The experimental results show the efficacy of multi-task learning and data augmentation, and the proposed method achieved better accuracies than existing methods.