Semantic annotation and retrieval of music and sound effects

Semantic annotation and retrieval of music and sound effects
复制标题

DOI:
10.1109/tasl.2007.913750
复制
发表时间:
2008-02-01
影响因子:
--
通讯作者:
Lanckriet, Gert
Lanckriet, Gert
中科院分区:
其他
文献类型:
--
作者:
Turnbull, Douglas;Barrington, Luke;Lanckriet, Gert

文献摘要

被引文献

相似文献

我们提出了一个计算机试听系统,既可以注释新的音轨与语义上有意义的话,并检索相关的轨道从数据库中的未标记的音频内容给出了一个基于文本的查询。我们认为相关的任务,基于内容的音频注释和检索作为一个监督的多类,多标签的问题,我们模型的声学特征和单词的联合概率。我们收集了1700个人类生成的注释,描述了500个西方流行音乐曲目的数据集。对于词汇表中的每个单词,我们使用这些数据在音频特征空间上训练高斯混合模型(GMM)。我们使用加权混合层次期望最大化算法估计模型的参数。该算法是更大的数据集的可扩展性,并产生更好的密度估计比标准的参数估计技术。我们的系统产生的音乐注释的质量与人类在相同任务上的表现相当。我们的“文本查询”系统可以检索到大量的音乐相关的词适当的歌曲。我们还表明,我们的试听系统是一般的学习模型,可以注释和检索声音效果。
We present a computer audition system that can both annotate novel audio tracks with semantically meaningful words and retrieve relevant tracks from a database of unlabeled audio content given a text-based query. We consider the related tasks of content-based audio annotation and retrieval as one supervised multiclass, multilabel problem in which we model the joint probability of acoustic features and words. We collect a data set of 1700 human-generated annotations that describe 500 Western popular music tracks. For each word in a vocabulary, we use this data to train a Gaussian mixture model (GMM) over an audio feature space. We estimate the parameters of the model using the weighted mixture hierarchies expectation maximization algorithm. This algorithm is more scalable to large data sets and produces better density estimates than standard parameter estimation techniques. The quality of the music annotations produced by our system is comparable with the performance of humans on the same task. Our "query-by-text" system can retrieve appropriate songs for a large number of musically relevant words. We also show that our audition system is general by learning a model that can annotate and retrieve sound effects.