Bayesian Singing Transcription Based on a Hierarchical Generative Model of Keys, Musical Notes, and F0 Trajectories

Bayesian Singing Transcription Based on a Hierarchical Generative Model of Keys, Musical Notes, and F0 Trajectories
复制标题

DOI:
10.1109/taslp.2020.2996095
复制
发表时间:
2020
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Ryo Nishikimi;Eita Nakamura;Masataka Goto;Katsutoshi Itoyama;Kazuyoshi Yoshii
Ryo Nishikimi;Eita Nakamura;Masataka Goto;Katsutoshi Itoyama;Kazuyoshi Yoshii
中科院分区:
其他
文献类型:
--
作者:
Ryo Nishikimi;Eita Nakamura;Masataka Goto;Katsutoshi Itoyama;Kazuyoshi Yoshii

文献摘要

被引文献

相似文献

本文描述了自动歌唱转录(AST),其从给定的音乐音频信号估计用量化的音高和持续时间表示的歌唱旋律的人类可读的乐谱。为了实现这一目标,我们提出了一种统计方法,通过量化的轨迹声乐基频(F0)的时间和频率方向估计的乐谱。由于人声F0轨迹与乐谱中指定的音符的音高和开始时间有很大的偏差,因此应该考虑音符的本地键和节奏。在这篇文章中,我们提出了一个贝叶斯分层隐半马尔可夫模型(HHSMM),它集成了一个乐谱模型描述的本地键和节奏的音符与F0轨迹模型描述的时间和频率偏差的F0轨迹。给定一个F0轨迹,一个音符序列,本地键,时间和频率偏差可以联合估计通过使用马尔可夫链蒙特卡罗(MCMC)方法。我们研究了所提出的模型的每个组件的效果,并表明乐谱模型提高了AST的性能。
This article describes automatic singing transcription (AST) that estimates a human-readable musical score of a sung melody represented with quantized pitches and durations from a given music audio signal. To achieve the goal, we propose a statistical method for estimating the musical score by quantizing a trajectory of vocal fundamental frequencies (F0s) in the time and frequency directions. Since vocal F0 trajectories considerably deviate from the pitches and onset times of musical notes specified in musical scores, the local keys and rhythms of musical notes should be taken into account. In this article we propose a Bayesian hierarchical hidden semi-Markov model (HHSMM) that integrates a musical score model describing the local keys and rhythms of musical notes with an F0 trajectory model describing the temporal and frequency deviations of an F0 trajectory. Given an F0 trajectory, a sequence of musical notes, that of local keys, and the temporal and frequency deviations can be estimated jointly by using a Markov chain Monte Carlo (MCMC) method. We investigated the effect of each component of the proposed model and showed that the musical score model improves the performance of AST.