You Call That Singing? Ensemble Classification for Multi-Cultural Collections of Music Recordings

You Call That Singing? Ensemble Classification for Multi-Cultural Collections of Music Recordings
复制标题

你管那叫唱歌吗?

DOI:
--
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
M. Casey
M. Casey
中科院分区:
--
文献类型:
--
作者:
Polina Proutskova;M. Casey

文献摘要

被引文献

相似文献

在民族音乐学领域录音中发现的各种声乐风格,音乐纹理和录音技术使我们考虑自动标记内容以了解录音是歌曲还是乐器作品的问题。此外,如果是一首歌,我们感兴趣的是标记声乐结构的各个方面:例如独奏,合唱,无伴奏合唱或乐器演唱。我们提出的证据表明,自动注释是可行的记录收藏展示了广泛的记录技术和代表来自世界各地的音乐文化。我们的实验使用了Alan Lomax Cantometrics训练磁带数据集,以鼓励未来的比较评估。实验进行了标记的子集组成的几百个轨道,注释在轨道和帧的水平,作为无伴奏合唱,唱歌加乐器或乐器。我们训练了一帧接一帧的SVM分类器使用MFCC功能的积极和消极的样本两个任务:每帧标签的歌唱和无伴奏合唱。在进一步的实验中,逐帧分类器输出被整合以估计整个轨道的主要内容。我们的研究结果表明,逐帧分类器实现了71%的帧精度和整个轨道分类器集成实现了88%的精度。最后,我们的分类器错误的分析,建议开发更强大的功能和分类器策略的大型人种学不同的集合的途径。
The wide range of vocal styles, musical textures and recording techniques found in ethnomusicological field recordings leads us to consider the problem of automatically labeling the content to know whether a recording is a song or instrumental work. Furthermore, if it is a song, we are interested in labeling aspects of the vocal texture: e.g. solo, choral, acapella or singing with instruments. We present evidence to suggest that automatic annotation is feasible for recorded collections exhibiting a wide range of recording techniques and representing musical cultures from around the world. Our experiments used the Alan Lomax Cantometrics training tapes data set, to encourage future comparative evaluations. Experiments were conducted with a labeled subset consisting of several hundred tracks, annotated at the track and frame levels, as acapella singing, singing plus instruments or instruments only. We trained frame-by-frame SVM classifiers using MFCC features on positive and negative exemplars for two tasks: per-frame labeling of singing and acapella singing. In a further experiment, the frame-by-frame classifier outputs were integrated to estimate the predominant content of whole tracks. Our results show that frame-byframe classifiers achieved 71% frame accuracy and whole track classifier integration achieved 88% accuracy. We conclude with an analysis of classifier errors suggesting avenues for developing more robust features and classifier strategies for large ethnographically diverse collections.