A Mid-level Melody-based Representation for Calculating Audio Similarity

A Mid-level Melody-based Representation for Calculating Audio Similarity
复制标题

DOI:
--
复制
发表时间:
2006
影响因子:
6.8
通讯作者:
M. Marolt
M. Marolt
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Marolt

文献摘要

被引文献

相似文献

我们提出了一个中等水平的基于旋律的表示,它结合了音乐信号的旋律,节奏和结构方面,并用于计算音频相似性的措施。目前大多数音乐相似性的方法要么使用低级别的信号特征,如MFCC,主要捕捉音乐的音色特征,包含很少的语义信息,或需要符号表示,这是很难从音频信号中获得。建议的中级表示是我们试图通过提供一个集成的旋律,节奏和音乐信号的结构表示弥合音频和符号域之间的差距。该表示基于从突出的旋律线中提取的一组旋律片段,它是节拍同步的,这使得它独立于克里思变化,并包含分析作品中短旋律短语的重复信息。我们展示了如何可以自动计算从复调音频信号,并展示其用于发现歌曲之间的旋律相似性。我们目前的结果,通过使用的表示找到不同的音乐收藏中的歌曲的解释。
We propose a mid-level melody-based representation that incorporates melodic, rhythmic and structural aspects of a music signal and is useful for calculating audio similarity measures. Most current approaches to music similarity use either low-level signal features, such as MFCCs that mostly capture timbral characteristics of music and contain little semantic information, or require symbolic representations, which are difficult to obtain from audio signals. The proposed mid-level representation is our attempt to bridge the gap between audio and symbolic domains by providing an integrated melodic, rhythmic and structural representation of music signals. The representation is based on a set of melodic fragments extracted from prominent melodic lines, it is beat-synchronous, which makes it independent of tempo variations and contains information on repetitions of short melodic phrases within the analyzed piece. We show how it can be calculated automatically from polyphonic audio signals and demonstrate its use for discovering melodic similarities between songs. We present results obtained by using the representation for finding different interpretations of songs in a music collection.