LSTMs Can Learn Syntax-Sensitive Dependencies Well, But Modeling Structure Makes Them Better

LSTMs Can Learn Syntax-Sensitive Dependencies Well, But Modeling Structure Makes Them Better
复制标题

DOI:
10.18653/v1/p18-1132
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
A. Kuncoro;Chris Dyer;John Hale;Dani Yogatama;S. Clark;Phil Blunsom
A. Kuncoro;Chris Dyer;John Hale;Dani Yogatama;S. Clark;Phil Blunsom
中科院分区:
其他
文献类型:
--
作者:
A. Kuncoro;Chris Dyer;John Hale;Dani Yogatama;S. Clark;Phil Blunsom

文献摘要

被引文献

相似文献

语言具有层次结构,但最近的工作使用主语-动词一致性诊断认为,最先进的语言模型,LSTM,无法学习长距离的语法敏感的依赖关系。使用相同的诊断,我们表明,事实上,LSTM确实成功地学习这种依赖性,只要他们有足够的能力。然后,我们探讨是否有访问显式的句法信息的模型更有效地学习协议,以及如何将这种结构信息纳入模型的影响性能。我们发现,仅仅存在的句法信息并不能提高准确性,但当模型的架构是由语法,数量协议得到改善。此外,我们发现句法结构构建方式的选择会影响数字一致性的学习效果:自上而下的构建在捕获非局部结构依赖性方面优于左角和自下而上的变体。
Language exhibits hierarchical structure, but recent work using a subject-verb agreement diagnostic argued that state-of-the-art language models, LSTMs, fail to learn long-range syntax sensitive dependencies. Using the same diagnostic, we show that, in fact, LSTMs do succeed in learning such dependencies—provided they have enough capacity. We then explore whether models that have access to explicit syntactic information learn agreement more effectively, and how the way in which this structural information is incorporated into the model impacts performance. We find that the mere presence of syntactic information does not improve accuracy, but when model architecture is determined by syntax, number agreement is improved. Further, we find that the choice of how syntactic structure is built affects how well number agreement is learned: top-down construction outperforms left-corner and bottom-up variants in capturing non-local structural dependencies.