Open Sesame: Getting inside BERT’s Linguistic Knowledge

Open Sesame: Getting inside BERT’s Linguistic Knowledge
复制标题

DOI:
10.18653/v1/w19-4825
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Yongjie Lin;Y. Tan;R. Frank
Yongjie Lin;Y. Tan;R. Frank
中科院分区:
其他
文献类型:
--
作者:
Yongjie Lin;Y. Tan;R. Frank

文献摘要

被引文献

相似文献

BERT如何以及在多大程度上编码语法敏感的层次信息或位置敏感的线性信息?最近的研究表明,像BERT这样的语境表征在需要对语言结构敏感的任务中表现良好。我们在这里提出了两项研究,旨在提供一个更好的理解BERT的表征的性质。第一个重点是使用诊断分类器识别结构定义的元素,而第二个探索BERT的表示主语-动词协议和照应词-先行词的依赖关系,通过自我注意向量的定量评估。在这两种情况下,我们发现BERT在其较低层上对单词标记的位置信息进行编码,但在较高层上切换到面向层次的编码。然后,我们得出结论,BERT的表示确实模型的语言相关方面的层次结构,虽然他们似乎并没有表现出尖锐的敏感性,层次结构,发现在人类处理的反身回指。
How and to what extent does BERT encode syntactically-sensitive hierarchical information or positionally-sensitive linear information? Recent work has shown that contextual representations like BERT perform well on tasks that require sensitivity to linguistic structure. We present here two studies which aim to provide a better understanding of the nature of BERT’s representations. The first of these focuses on the identification of structurally-defined elements using diagnostic classifiers, while the second explores BERT’s representation of subject-verb agreement and anaphor-antecedent dependencies through a quantitative assessment of self-attention vectors. In both cases, we find that BERT encodes positional information about word tokens well on its lower layers, but switches to a hierarchically-oriented encoding on higher layers. We conclude then that BERT’s representations do indeed model linguistically relevant aspects of hierarchical structure, though they do not appear to show the sharp sensitivity to hierarchical structure that is found in human processing of reflexive anaphora.