Audio-to-score singing transcription based on a CRNN-HSMM hybrid model

Audio-to-score singing transcription based on a CRNN-HSMM hybrid model
复制标题

DOI:
10.1017/atsip.2021.4
复制
发表时间:
2021-04
影响因子:
3.2
通讯作者:
Ryo Nishikimi;Eita Nakamura;Masataka Goto;Kazuyoshi Yoshii
Ryo Nishikimi;Eita Nakamura;Masataka Goto;Kazuyoshi Yoshii
中科院分区:
--
文献类型:
--
作者:
Ryo Nishikimi;Eita Nakamura;Masataka Goto;Kazuyoshi Yoshii

文献摘要

被引文献

相似文献

本文描述了一种自动歌唱转录(AST)方法,该方法从输入的音乐信号中估计出歌唱旋律的人类可读乐谱。由于歌唱声音具有相当大的音高和时间变化,一种简单的级联方法无法避免许多音高和节奏错误,该方法估计F0轮廓并使用估计的tatum时间对其进行量化。为了解决这个问题,我们制定了一个音乐信号的统一生成模型,该模型由一个半马尔可夫语言模型和基于卷积递归神经网络(CRNN)的声学模型组成,该模型代表了基于音键的潜在音符的生成过程,而声学模型则代表了从音符中观察到的音乐信号的生成过程。由此产生的CRNN- hsmm混合模型使我们能够使用Viterbi算法从音乐信号中估计最可能的音符,同时利用关于音符的语法知识和CRNN的表达能力。实验结果表明,该方法优于传统的最先进的方法,并且音乐语言模型与声学模型的集成对AST的性能有积极的影响。
This paper describes an automatic singing transcription (AST) method that estimates a human-readable musical score of a sung melody from an input music signal. Because of the considerable pitch and temporal variation of a singing voice, a naive cascading approach that estimates an F0 contour and quantizes it with estimated tatum times cannot avoid many pitch and rhythm errors. To solve this problem, we formulate a unified generative model of a music signal that consists of a semi-Markov language model representing the generative process of latent musical notes conditioned on musical keys and an acoustic model based on a convolutional recurrent neural network (CRNN) representing the generative process of an observed music signal from the notes. The resulting CRNN-HSMM hybrid model enables us to estimate the most-likely musical notes from a music signal with the Viterbi algorithm, while leveraging both the grammatical knowledge about musical notes and the expressive power of the CRNN. The experimental results showed that the proposed method outperformed the conventional state-of-the-art method and the integration of the musical language model with the acoustic model has a positive effect on the AST performance.