ABC: Attention with Bounded-memory Control

ABC: Attention with Bounded-memory Control
复制标题

DOI:
10.18653/v1/2022.acl-long.515
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Hao Peng;Jungo Kasai;Nikolaos Pappas;Dani Yogatama;Zhaofeng Wu;Lingpeng Kong;Roy Schwartz;Noah A. Smith
Hao Peng;Jungo Kasai;Nikolaos Pappas;Dani Yogatama;Zhaofeng Wu;Lingpeng Kong;Roy Schwartz;Noah A. Smith
中科院分区:
其他
文献类型:
--
作者:
Hao Peng;Jungo Kasai;Nikolaos Pappas;Dani Yogatama;Zhaofeng Wu;Lingpeng Kong;Roy Schwartz;Noah A. Smith

文献摘要

被引文献

相似文献

变压器的构造在各种自然语言处理(NLP)任务上取得了最新的结果,但是,他们的注意力机制在序列长度上具有二次复杂性,使计算上的间接费用使人可以将其视为一个随机的记忆。读取效率的一种方式是,我们可以将不同的方法归为一个抽象,并以界限控制(ABC),它们在ABC的组织中有所不同。 Al。,2020b)以前不适用于灾难性的关注,我们提出了ABC的新实例,它从现有的ABC方法中汲取了灵感,但是用学习的型号将我们的启发式记忆函数替换为启发式记忆力。基线,它可以显着提高推理时间和空间效率,而没有可忽略的准确性损失。
Transformer architectures have achieved state- of-the-art results on a variety of natural language processing (NLP) tasks. However, their attention mechanism comes with a quadratic complexity in sequence lengths, making the computational overhead prohibitive, especially for long sequences. Attention context can be seen as a random-access memory with each token taking a slot. Under this perspective, the memory size grows linearly with the sequence length, and so does the overhead of reading from it. One way to improve the efficiency is to bound the memory size. We show that disparate approaches can be subsumed into one abstraction, attention with bounded-memory control (ABC), and they vary in their organization of the memory. ABC reveals new, unexplored possibilities. First, it connects several efficient attention variants that would otherwise seem apart. Second, this abstraction gives new insights—an established approach (Wang et al., 2020b) previously thought to not be applicable in causal attention, actually is. Last, we present a new instance of ABC, which draws inspiration from existing ABC approaches, but replaces their heuristic memory-organizing functions with a learned, contextualized one. Our experiments on language modeling, machine translation, and masked language model finetuning show that our approach outperforms previous efficient attention models; compared to the strong transformer baselines, it significantly improves the inference time and space efficiency with no or negligible accuracy loss.