A generative model for music transcription

A generative model for music transcription
复制标题

音乐转录的生成模型

DOI:
--
复制
发表时间:
2006
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
D. Barber
D. Barber
中科院分区:
--
文献类型:
--
作者:
A. Cemgil;H. Kappen;D. Barber

文献摘要

被引文献

相似文献

在本文中,我们提出了一个复调音乐转录的图形模型。我们的模型被描述为一个动态贝叶斯网络,体现了一种透明和易于计算的方法来解决这个声学分析问题。我们方法的一个优点是它强调对声音生成过程进行显式建模。它提供了一个清晰的框架,在这个框架中,关于音乐结构的高级(认知)先验信息可以与低级(声学物理)信息以原则性的方式相结合来执行分析。该模型是一般难以处理的切换卡尔曼滤波模型的特例。在可能的情况下,我们推导出精确的多项式时间推断过程,以及其他有效的近似。我们认为,我们的基于生成模型的方法在计算上对许多音乐应用是可行的,并且很容易扩展到更一般的听觉场景分析场景。
In this paper, we present a graphical model for polyphonic music transcription. Our model, formulated as a dynamical Bayesian network, embodies a transparent and computationally tractable approach to this acoustic analysis problem. An advantage of our approach is that it places emphasis on explicitly modeling the sound generation procedure. It provides a clear framework in which both high level (cognitive) prior information on music structure can be coupled with low level (acoustic physical) information in a principled manner to perform the analysis. The model is a special case of the, generally intractable, switching Kalman filter model. Where possible, we derive, exact polynomial time inference procedures, and otherwise efficient approximations. We argue that our generative model based approach is computationally feasible for many music applications and is readily extensible to more general auditory scene analysis scenarios.