Emergent linguistic structure in artificial neural networks trained by self-supervision

Emergent linguistic structure in artificial neural networks trained by self-supervision
复制标题

DOI:
10.1073/pnas.1907367117
复制
发表时间:
2020-12-01
影响因子:
11.1
通讯作者:
Levy, Omer
Levy, Omer
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Manning, Christopher D.;Clark, Kevin;Levy, Omer

文献摘要

被引文献

相似文献

本文探讨了通过自我监督训练的大型人工神经网络学习的语言结构知识,该模型只是试图在给定的上下文中预测一个被屏蔽的单词。人类语言交流是通过单词序列进行的,但语言理解需要构建丰富的层次结构,而这些结构从未被明确观察到。这一机制一直是人类语言习得的主要谜团,而工程工作主要是通过对这种潜在结构标记的句子树库进行监督学习。然而,我们证明了现代深度上下文语言模型在没有任何明确监督的情况下学习这种结构的主要方面。我们开发的方法,识别语言的层次结构出现在人工神经网络,并证明在这些模型中的组件侧重于句法语法关系和照应共指。事实上,我们表明,在这些模型中学习嵌入的线性变换捕获解析树距离到一个令人惊讶的程度,允许近似重建的句子树结构通常假设的语言学家。这些结果有助于解释为什么这些模型在许多语言理解任务中带来了如此大的改进。
This paper explores the knowledge of linguistic structure learned by large artificial neural networks, trained via self-supervision, whereby the model simply tries to predict a masked word in a given context. Human language communication is via sequences of words, but language understanding requires constructing rich hierarchical structures that are never observed explicitly. The mechanisms for this have been a prime mystery of human language acquisition, while engineering work has mainly proceeded by supervised learning on treebanks of sentences hand labeled for this latent structure. However, we demonstrate that modern deep contextual language models learn major aspects of this structure, without any explicit supervision. We develop methods for identifying linguistic hierarchical structure emergent in artificial neural networks and demonstrate that components in these models focus on syntactic grammatical relationships and anaphoric coreference. Indeed, we show that a linear transformation of learned embeddings in these models captures parse tree distances to a surprising degree, allowing approximate reconstruction of the sentence tree structures normally assumed by linguists. These results help explain why these models have brought such large improvements across many language-understanding tasks.