A Primal-Dual Framework for Transformers and Neural Networks

A Primal-Dual Framework for Transformers and Neural Networks
复制标题

DOI:
--
复制
发表时间:
2024-06
期刊:
Exploration of Immunology
影响因子:
--
通讯作者:
T. Nguyen;Tam Nguyen;Nhat Ho;A. Bertozzi;Richard Baraniuk;S. Osher
T. Nguyen;Tam Nguyen;Nhat Ho;A. Bertozzi;Richard Baraniuk;S. Osher
中科院分区:
其他
文献类型:
--
作者:
T. Nguyen;Tam Nguyen;Nhat Ho;A. Bertozzi;Richard Baraniuk;S. Osher

文献摘要

相似文献

自我注意力是transformers在序列建模任务中取得显著成功的关键,包括自然语言处理和计算机视觉中的许多应用。像神经网络层一样,这些注意力机制通常是由知识和经验开发的。为了提供一个原则性的框架,在变压器中构建注意层,我们表明,自注意力对应于来自支持向量回归问题,其原始配方的形式的神经网络层的支持向量扩展。使用我们的框架,我们推导出在实践中使用的流行的注意力层,并提出了两个新的注意力:1)批归一化注意力(Attention-BN)来自批归一化层和2)注意力与缩放头(Attention-SH)来自使用较少的训练数据来拟合SVR模型。我们经验证明了Attention-BN和Attention-SH在减少头部冗余,提高模型的准确性,并提高模型在各种实际应用中的效率,包括图像和时间序列分类方面的优势。
Self-attention is key to the remarkable success of transformers in sequence modeling tasks including many applications in natural language processing and computer vision. Like neural network layers, these attention mechanisms are often developed by heuristics and experience. To provide a principled framework for constructing attention layers in transformers, we show that the self-attention corresponds to the support vector expansion derived from a support vector regression problem, whose primal formulation has the form of a neural network layer. Using our framework, we derive popular attention layers used in practice and propose two new attentions: 1) the Batch Normalized Attention (Attention-BN) derived from the batch normalization layer and 2) the Attention with Scaled Head (Attention-SH) derived from using less training data to fit the SVR model. We empirically demonstrate the advantages of the Attention-BN and Attention-SH in reducing head redundancy, increasing the model's accuracy, and improving the model's efficiency in a variety of practical applications including image and time-series classification.