Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars

Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars
复制标题

DOI:
10.48550/arxiv.2312.01429
复制
发表时间:
2023-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Kaiyue Wen;Yuchen Li;Bing Liu;Andrej Risteski
Kaiyue Wen;Yuchen Li;Bing Liu;Andrej Risteski
中科院分区:
其他
文献类型:
--
作者:
Kaiyue Wen;Yuchen Li;Bing Liu;Andrej Risteski

文献摘要

相似文献

可解释性方法旨在通过检查模型的各个方面(例如权重矩阵或注意力模式)来理解训练模型(例如 Transofmer)实现的算法。在这项工作中,通过结合理论结果和对合成数据进行仔细控制的实验,我们对只关注模型的各个部分而不是将网络视为整体的方法持批判态度。我们考虑学习(有界)Dyck 语言的简单综合设置。从理论上讲,我们表明(精确或近似)解决此任务的模型集满足从形式语言中的思想(泵引理)衍生的结构特征。我们使用这个特征来表明最优集在质量上是丰富的;特别是,单层的注意力模式可以“接近随机”,同时保留网络的功能。我们还通过大量的实验表明,这些结构不仅仅是理论上的产物:即使在严格限制模型的架构之后,也可以通过标准训练获得截然不同的解决方案。因此,基于检查 Transformer 中的各个头或权重矩阵的可解释性声明可能会产生误导。
Interpretability methods aim to understand the algorithm implemented by a trained model (e.g., a Transofmer) by examining various aspects of the model, such as the weight matrices or the attention patterns. In this work, through a combination of theoretical results and carefully controlled experiments on synthetic data, we take a critical view of methods that exclusively focus on individual parts of the model, rather than consider the network as a whole. We consider a simple synthetic setup of learning a (bounded) Dyck language. Theoretically, we show that the set of models that (exactly or approximately) solve this task satisfy a structural characterization derived from ideas in formal languages (the pumping lemma). We use this characterization to show that the set of optima is qualitatively rich; in particular, the attention pattern of a single layer can be ``nearly randomized'', while preserving the functionality of the network. We also show via extensive experiments that these constructions are not merely a theoretical artifact: even after severely constraining the architecture of the model, vastly different solutions can be reached via standard training. Thus, interpretability claims based on inspecting individual heads or weight matrices in the Transformer can be misleading.