Music Transformer: Generating Music with Long-Term Structure

Music Transformer: Generating Music with Long-Term Structure
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Ian Simon;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;Monica Dinculescu;D. Eck
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Ian Simon;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;Monica Dinculescu;D. Eck
中科院分区:
其他
文献类型:
--
作者:
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Ian Simon;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;Monica Dinculescu;D. Eck

文献摘要

被引文献

相似文献

音乐在很大程度上依赖于重复来构建结构和意义。自我参照发生在多个时间尺度上,从主题到短语再到整个音乐片段的重复使用,例如阿坝结构的作品。Transformer(Vaswani等人,2017),一个基于自我注意的序列模型,在许多需要保持长期一致性的生成任务中取得了令人信服的结果。这表明,自我注意力也可能非常适合于模仿音乐。然而,在音乐创作和表演中,相对时间是至关重要的。用于表示Transformer中的相对位置信息的现有方法基于成对距离来调制注意力(Shaw等人,2018年)。这对于诸如音乐作品的长序列是不切实际的,因为它们对于中间相对信息的存储器复杂度在序列长度上是二次的。我们提出了一种算法,减少了他们的中间内存需求的线性序列长度。这使我们能够证明,具有我们修改的相对注意力机制的Transformer可以生成分钟长的作品(数千个步骤,四倍于Oore等人,2018)具有引人注目的结构,生成连贯地阐述给定主题的延续,并在seq2seq设置中生成以旋律为条件的重复。我们评估了Transformer与我们的相对注意力机制上的两个数据集,JSB合唱团和钢琴电子竞争,并获得国家的最先进的结果对后者。
Music relies heavily on repetition to build structure and meaning. Self-reference occurs on multiple timescales, from motifs to phrases to reusing of entire sections of music, such as in pieces with ABA structure. The Transformer (Vaswani et al., 2017), a sequence model based on self-attention, has achieved compelling results in many generation tasks that require maintaining long-range coherence. This suggests that self-attention might also be well-suited to modeling music. In musical composition and performance, however, relative timing is critically important. Existing approaches for representing relative positional information in the Transformer modulate attention based on pairwise distance (Shaw et al., 2018). This is impractical for long sequences such as musical compositions since their memory complexity for intermediate relative information is quadratic in the sequence length. We propose an algorithm that reduces their intermediate memory requirement to linear in the sequence length. This enables us to demonstrate that a Transformer with our modified relative attention mechanism can generate minute-long compositions (thousands of steps, four times the length modeled in Oore et al., 2018) with compelling structure, generate continuations that coherently elaborate on a given motif, and in a seq2seq setup generate accompaniments conditioned on melodies. We evaluate the Transformer with our relative attention mechanism on two datasets, JSB Chorales and Piano-e-Competition, and obtain state-of-the-art results on the latter.