An Improved Relative Self-Attention Mechanism for Transformer with Application to Music Generation

An Improved Relative Self-Attention Mechanism for Transformer with Application to Music Generation
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;D. Eck
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;D. Eck
中科院分区:
其他
文献类型:
--
作者:
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;D. Eck

文献摘要

被引文献

相似文献

音乐在很大程度上依赖于自我参照来构建结构和意义。我们探索Transformer架构[27]作为音乐的生成模型,因为自我注意力在需要长期结构的任务上显示出令人信服的结果,例如维基百科摘要生成[18]。然而,时间信息对于复调音乐是至关重要的,并且Transformer在其结构中没有明确地建模绝对或相对时间。为了应对这一挑战,Shaw等人[22]将相对位置表示引入自我注意力,以改进机器翻译。然而,该制剂不能扩展到更长的序列。我们提出了一种改进的配方,减少了相对位置计算的内存需求从O(ld)到O(ld),其中l是序列的长度和d是隐藏的大小,使它能够训练更长的序列,并实现更快的收敛。在符号音乐的实验中,我们发现相对自我注意大大提高了无条件生成的样本质量,并且能够生成长度比训练集更长的序列。当用初始序列启动时,该模型生成连贯地发展素数并表现出长期结构的延续。相对的自我注意力有助于捕捉音乐作品中更丰富的关系2 3。
Music relies heavily on self-reference to build structure and meaning. We explore the Transformer architecture [27] as a generative model for music, as self-attention has shown compelling results on tasks that require long-term structure such as Wikipedia summary generation [18]. However, timing information is critical for polyphonic music, and Transformer does not explicitly model absolute or relative timing in its structure. To address this challenge, Shaw et al. [22] introduced relative position representations to self-attention to improve machine translation. However, the formulation was not scalable to longer sequences. We propose an improved formulation which reduces the memory requirements of the relative position computation from O(ld) to O(ld), where l is the length of sequences and d is the hidden size, making it possible to train much longer sequences and achieve faster convergence. In experiments on symbolic music we find that relative selfattention substantially improves sample quality for unconditioned generation and is able to generate sequences of lengths longer than those from the training set. When primed with an initial sequence, the model generates continuations that develop the prime coherently and exhibit long-term structure. Relative self-attention can be instrumental in capturing richer relationships within a musical piece 2 3.