Choose a Transformer: Fourier or Galerkin

Choose a Transformer: Fourier or Galerkin
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Shuhao Cao
Shuhao Cao
中科院分区:
其他
文献类型:
--
作者:
Shuhao Cao

文献摘要

相似文献

在本文中,我们首次将Attention Is All You Need中最先进的Transformer中的自注意力应用于与偏微分方程相关的数据驱动算子学习问题。努力放在一起,解释注意力机制的原理,并提高其功效。利用Hilbert空间中的算子逼近理论,首次证明了在标度点积注意力中的softmax归一化是充分的,但不是必要的。在没有softmax的情况下,可以证明线性化的Transformer变体的近似能力与逐层的Petrov-Galerkin投影相当,并且估计与序列长度无关。提出了一种新的层归一化方案,模仿Petrov-Galerkin投影,允许缩放通过注意层传播,这有助于模型在具有未归一化数据的算子学习任务中实现显着的准确性。最后,我们提出了三个操作学习实验,包括粘性Burgers方程,界面达西流,和反界面系数识别问题。新提出的简单的基于注意力的算子学习器Galerkin Transformer,在训练成本和评估精度方面都比其softmax归一化的同类算法有了显着的改进。
In this paper, we apply the self-attention from the state-of-the-art Transformer in Attention Is All You Need for the first time to a data-driven operator learning problem related to partial differential equations. An effort is put together to explain the heuristics of, and to improve the efficacy of the attention mechanism. By employing the operator approximation theory in Hilbert spaces, it is demonstrated for the first time that the softmax normalization in the scaled dot-product attention is sufficient but not necessary. Without softmax, the approximation capacity of a linearized Transformer variant can be proved to be comparable to a Petrov-Galerkin projection layer-wise, and the estimate is independent with respect to the sequence length. A new layer normalization scheme mimicking the Petrov-Galerkin projection is proposed to allow a scaling to propagate through attention layers, which helps the model achieve remarkable accuracy in operator learning tasks with unnormalized data. Finally, we present three operator learning experiments, including the viscid Burgers' equation, an interface Darcy flow, and an inverse interface coefficient identification problem. The newly proposed simple attention-based operator learner, Galerkin Transformer, shows significant improvements in both training cost and evaluation accuracy over its softmax-normalized counterparts.