Addressing Token Uniformity in Transformers via Singular Value Transformation

Addressing Token Uniformity in Transformers via Singular Value Transformation
复制标题

DOI:
10.48550/arxiv.2208.11790
复制
发表时间:
2022-08
期刊:
--
影响因子:
--
通讯作者:
Hanqi Yan;Lin Gui;Wenjie Li;Yulan He
Hanqi Yan;Lin Gui;Wenjie Li;Yulan He
中科院分区:
其他
文献类型:
--
作者:
Hanqi Yan;Lin Gui;Wenjie Li;Yulan He

文献摘要

相似文献

在基于变压器的模型中,通常观察到令牌一致性,其中不同的令牌在经历了变压器中堆叠的多个自我关注层后,共享了很大比例的相似信息。本文提出用各变压器层输出的奇异值分布来刻画令牌一致性现象,并通过实验证明较小偏斜的奇异值分布可以缓解令牌一致性问题。基于我们的观察,我们定义了奇异值分布的几个理想性质,并提出了一种新的奇异值更新变换函数。我们证明了该变换函数除了可以减少令牌的一致性外,还应该保持原嵌入空间中的局部邻域结构。我们提出的奇异值变换函数被应用于一系列基于转换器的语言模型,如BERT、ALBERT、Roberta和DistilBERT,在语义文本相似度评估和一系列粘合任务中都观察到了改进的性能。我们的源代码可以在https://github.com/hanqi-qi/tokenUni.git.上找到
Token uniformity is commonly observed in transformer-based models, in which different tokens share a large proportion of similar information after going through stacked multiple self-attention layers in a transformer. In this paper, we propose to use the distribution of singular values of outputs of each transformer layer to characterise the phenomenon of token uniformity and empirically illustrate that a less skewed singular value distribution can alleviate the `token uniformity' problem. Base on our observations, we define several desirable properties of singular value distributions and propose a novel transformation function for updating the singular values. We show that apart from alleviating token uniformity, the transformation function should preserve the local neighbourhood structure in the original embedding space. Our proposed singular value transformation function is applied to a range of transformer-based language models such as BERT, ALBERT, RoBERTa and DistilBERT, and improved performance is observed in semantic textual similarity evaluation and a range of GLUE tasks. Our source code is available at https://github.com/hanqi-qi/tokenUni.git.