Scalable Syntax-Aware Language Models Using Knowledge Distillation

Scalable Syntax-Aware Language Models Using Knowledge Distillation
复制标题

DOI:
10.18653/v1/p19-1337
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
A. Kuncoro;Chris Dyer;Laura Rimell;S. Clark;Phil Blunsom
A. Kuncoro;Chris Dyer;Laura Rimell;S. Clark;Phil Blunsom
中科院分区:
其他
文献类型:
--
作者:
A. Kuncoro;Chris Dyer;Laura Rimell;S. Clark;Phil Blunsom

文献摘要

被引文献

相似文献

先前的工作表明,在少量的训练数据上,句法神经语言模型比顺序语言模型更成功地学习结构敏感的概括。然而,它们的计算复杂性使得缩放困难,并且当序列模型可以访问更大量的训练数据时,结构偏差是否仍然是必要的仍然是一个悬而未决的问题。为了回答这个问题,我们引入了一种有效的知识蒸馏(KD)技术,该技术将知识从在小语料库上训练的语法语言模型转移到LSTM语言模型,从而使LSTM能够为其学习的更大的训练数据开发一种结构上更敏感的表示。在有针对性的语法评估,我们发现,虽然顺序LSTM执行比以前报道的要好得多,我们提出的技术大大提高了这个基线,产生一个新的艺术状态。我们的研究结果和分析肯定了结构偏差的重要性,即使在模型中,从大量的数据学习。
Prior work has shown that, on small amounts of training data, syntactic neural language models learn structurally sensitive generalisations more successfully than sequential language models. However, their computational complexity renders scaling difficult, and it remains an open question whether structural biases are still necessary when sequential models have access to ever larger amounts of training data. To answer this question, we introduce an efficient knowledge distillation (KD) technique that transfers knowledge from a syntactic language model trained on a small corpus to an LSTM language model, hence enabling the LSTM to develop a more structurally sensitive representation of the larger training data it learns from. On targeted syntactic evaluations, we find that, while sequential LSTMs perform much better than previously reported, our proposed technique substantially improves on this baseline, yielding a new state of the art. Our findings and analysis affirm the importance of structural biases, even in models that learn from large amounts of data.