Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFs

Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFs
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Yujia Yan;Frank Cwitkowitz;Z. Duan
Yujia Yan;Frank Cwitkowitz;Z. Duan
中科院分区:
其他
文献类型:
--
作者:
Yujia Yan;Frank Cwitkowitz;Z. Duan

文献摘要

相似文献

钢琴转录系统通常被优化以估计音频的每个帧处的音高活动。它们通常遵循精心设计的算法和后处理算法,以根据帧级预测来估计音符事件。最近的方法也将钢琴转录框定为多任务学习问题,其中独立地估计音符事件的不同阶段的激活。这些实践与任务的预期结果并不一致,任务的预期结果是将音符间隔指定为整体事件,而不是将不相交的观察结果聚合在一起。在这项工作中,我们提出了一种新的钢琴转录,这是优化直接预测音符事件的制定。我们的方法是基于半马尔可夫条件随机场(半CRF),它产生的时间间隔,而不是个别帧的分数。当以这种方式制定钢琴转录时,我们消除了依赖于对音符事件的不同阶段的不相交帧级估计的需要。我们在MAESTRO数据集上进行了实验,并证明了所提出的模型超越了目前最先进的钢琴转录。我们的研究结果表明,半CRF输出层,虽然仍然是二次的复杂性,是一个简单,快速和性能良好的解决方案,基于事件的预测,并可能导致类似的成功,在其他领域目前依赖于帧级估计。
Piano transcription systems are typically optimized to estimate pitch activity at each frame of audio. They are often followed by carefully designed heuristics and post-processing algorithms to estimate note events from the frame-level predictions. Recent methods have also framed piano transcription as a multi-task learning problem, where the activation of different stages of a note event are estimated independently. These practices are not well aligned with the desired outcome of the task, which is the specification of note intervals as holistic events, rather than the aggregation of disjoint observations. In this work, we propose a novel formulation of piano transcription, which is optimized to directly predict note events. Our method is based on Semi-Markov Conditional Random Fields (semi-CRF), which produce scores for intervals rather than individual frames. When formulating piano transcription in this way, we eliminate the need to rely on disjoint frame-level estimates for different stages of a note event. We conduct experiments on the MAESTRO dataset and demonstrate that the proposed model surpasses the current state-of-the-art for piano transcription. Our results suggest that the semi-CRF output layer, while still quadratic in complexity, is a simple, fast and well-performing solution for event-based prediction, and may lead to similar success in other areas which currently rely on frame-level estimates.