Musical Rhythm Transcription Based on Bayesian Piece-Specific Score Models Capturing Repetitions

Musical Rhythm Transcription Based on Bayesian Piece-Specific Score Models Capturing Repetitions
复制标题

基于贝叶斯片段特定乐谱模型捕获重复的音乐节奏转录

DOI:
10.1016/j.ins.2021.04.100
复制
发表时间:
2021
影响因子:
8.1
通讯作者:
Kazuyoshi Yoshii
Kazuyoshi Yoshii
中科院分区:
计算机科学1区
文献类型:
--
作者:
Eita Nakamura;Kazuyoshi Yoshii

文献摘要

相似文献

大多数用于音乐转录的乐谱模型(也称为音乐语言模型)的工作都集中在描述乐谱中音符的局部顺序依赖性,而未能捕获它们的全局重复结构,这可以为转录音乐提供有用的指导。专注于节奏,我们制定了几个类的贝叶斯马尔可夫模型的乐谱,间接使用稀疏的转移概率的注意到或注意到的模式来描述重复。这使我们能够构建一个不固定的重复结构,看不见的分数,并推导出易于处理的推理算法的具体作品的模型。此外,为了描述近似的重复,我们明确地将修改重复的音符/音符模式的过程。我们将这些模型作为节奏转录的先验乐谱模型,其中特定于作品的乐谱模型是通过贝叶斯学习从执行的数据中推断出来的,与传统的乐谱模型的监督构建相反。使用流行音乐的声乐旋律的评估表明,贝叶斯模型提高了大多数测试模型类型的转录准确性,表明所提出的方法的普遍功效。此外,我们发现了一个有效的数据表示建模的节奏,最大限度地提高转录的准确性和计算效率。
Most work on musical score models (a.k.a. musical language models) for music transcription has focused on describing the local sequential dependence of notes in musical scores and failed to capture their global repetitive structure, which can be a useful guide for transcribing music. Focusing on rhythm, we formulate several classes of Bayesian Markov models of musical scores that describe repetitions indirectly using the sparse transition probabilities of notes or note patterns. This enables us to construct piece-specific models for unseen scores with an unfixed repetitive structure and to derive tractable inference algorithms. Moreover, to describe approximate repetitions, we explicitly incorporate a process for modifying the repeated notes/note patterns. We apply these models as prior musical score models for rhythm transcription, where piece-specific score models are inferred from performed MIDI data by Bayesian learning, in contrast to the conventional supervised construction of score models. Evaluations using the vocal melodies of popular music showed that the Bayesian models improved the transcription accuracy for most of the tested model types, indicating the universal efficacy of the proposed approach. Moreover, we found an effective data representation for modelling rhythms that maximizes the transcription accuracy and computational efficiency.