Visualizing Attention in Transformer-Based Language models

Visualizing Attention in Transformer-Based Language models
复制标题

基于 Transformer 的语言模型中的注意力可视化

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Jesse Vig
Jesse Vig
中科院分区:
--
文献类型:
--
作者:
Jesse Vig

文献摘要

参考文献

被引文献

相似文献

我们提出了一个开源的工具,用于可视化基于Transformer的语言模型中的多头自我注意。该工具通过在三个粒度级别可视化注意力来扩展早期的工作:注意力头部级别、模型级别和神经元级别。我们描述了每个视图如何帮助解释模型,并在OpenAI GPT-2预训练语言模型上演示了该工具。我们还提供了三个用例,展示了该工具如何提供有关如何调整或改进模型的见解。
We present an open-source tool for visualizing multi-head self-attention in Transformer-based language models. The tool extends earlier work by visualizing attention at three levels of granularity: the attention-head level, the model level, and the neuron level. We describe how each of these views can help to interpret the model, and we demonstrate the tool on the OpenAI GPT-2 pretrained language model. We also present three use cases showing how the tool might provide insights on how to adapt or improve the model.
DOI: 10.18653/v1/n18-2003
发表时间: 2018-04
期刊: ArXiv
影响因子: --
作者:
Jieyu Zhao;Tianlu Wang;Mark Yatskar;Vicente Ordonez;Kai-Wei Chang
通讯作者: Jieyu Zhao;Tianlu Wang;Mark Yatskar;Vicente Ordonez;Kai-Wei Chang