Atypical Lyrics Completion Considering Musical Audio Signals

Atypical Lyrics Completion Considering Musical Audio Signals
复制标题

DOI:
10.1007/978-3-030-67832-6_15
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Kento Watanabe;Masataka Goto
Kento Watanabe;Masataka Goto
中科院分区:
其他
文献类型:
--
作者:
Kento Watanabe;Masataka Goto

文献摘要

相似文献

本文讨论了歌词完成的新任务,以提供创造性支持。我们提出的任务旨在建议(1)非典型但(2)适合音乐音频信号的单词。之前的方法侧重于使用倾向于生成频繁短语(例如,“我爱你”)的语言模型的全自动歌词生成任务,尽管非典型化对创造性支持很重要。在这项研究中,我们提出了一种具有负采样策略的新颖向量空间模型,并假设在统一的向量空间中嵌入多模态方面(单词,句子草稿和音乐音频信号)有助于捕获(1)单词的非典型性和(2)单词与音乐音频情绪之间的关系。为了验证我们的假设,我们使用了一个大规模的数据集来研究所提出的多模态向量空间模型是否建议非典型词。从实验结果中得出了几个结论。一是消极抽样策略有助于暗示非典型词。另一个是,嵌入音频信号有助于建议适合所提供音乐音频的情绪的单词。
This paper addresses the novel task of lyrics completion for creative support. Our proposed task aims to suggest words that are (1) atypical but (2) suitable for musical audio signals. Previous approaches focused on fully automatic lyrics generation tasks using language models that tend to generate frequent phrases (e.g., “I love you”), despite the importance of atypicality for creative support. In this study, we propose a novel vector space model with negative sampling strategy and hypothesize that embedding multimodal aspects (words, draft sentences, and musical audio signals) in a unified vector space contributes to capturing (1) the atypicality of words and (2) the relationships between words and the moods of music audio. To test our hypothesis, we used a large-scale dataset to investigate whether the proposed multimodal vector space model suggests atypical words. Several findings were obtained from experiment results. One is that the negative sampling strategy contributes to suggesting atypical words. Another is that embedding audio signals contributes to suggesting words suitable for the mood of the provided music audio.