MusicTM-Dataset for Joint Representation Learning Among Sheet Music, Lyrics, and Musical Audio

MusicTM-Dataset for Joint Representation Learning Among Sheet Music, Lyrics, and Musical Audio
复制标题

DOI:
10.1007/978-981-16-1649-5_7
复制
发表时间:
2020-11
期刊:
Proceedings of the 8th Conference on Sound and Music Technology
影响因子:
--
通讯作者:
Donghuo Zeng;Yi Yu;K. Oyama
Donghuo Zeng;Yi Yu;K. Oyama
中科院分区:
其他
文献类型:
--
作者:
Donghuo Zeng;Yi Yu;K. Oyama

文献摘要

相似文献

本文提出了一个名为MusicTM-DataSet的音乐数据集,用于提高不同类型跨模式检索的表征学习能力。包括三种模式的小的大型音乐数据集可用于CMR的学习表示。为了收集音乐数据集,我们扩展原始的乐谱符号来合成音频和生成的乐谱图像,并建立基于乐谱图像、音频片段和音节表示文本的乐谱符号作为细粒度对齐,以便利用MusicTM数据集来接收多模式数据点的共享表示。MusicTM-DataSet提供了乐谱图像、歌词文本和合成音频三种形态,它们的表示由一些高级模型提取。在本文中,我们介绍了音乐数据集的背景,并阐述了我们的数据收集过程。基于我们的数据集,我们实现了一些针对CMR任务的基本方法。可在https://github.com/dddzeng/MusicTM-Dataset中访问MusicTM-DataSet。
This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is available for learning representations for CMR. To collect a music dataset, we expand the original musical notation to synthesize audio and generated sheet-music image, and build musical notation based sheet-music image, audio clip and syllable-denotation text as fine-grained alignment, such that the MusicTM-Dataset can be exploited to receive shared representation for multi-modal data points. The MusicTM-Dataset presents 3 kinds of modalities, which consists of the image of sheet-music, the text of lyrics and synthesized audio, their representations are extracted by some advanced models. In this paper, we introduce the background of music dataset and express the process of our data collection. Based on our dataset, we achieve some basic methods for CMR tasks. The MusicTM-Dataset are accessible in  https://github.com/dddzeng/MusicTM-Dataset .