Learning the Dyck Language with Attention-based Seq2Seq Models

Learning the Dyck Language with Attention-based Seq2Seq Models
复制标题

DOI:
10.18653/v1/w19-4815
复制
发表时间:
2019-08
期刊:
--
影响因子:
--
通讯作者:
Xiang Yu;Ngoc Thang Vu;Jonas Kuhn
Xiang Yu;Ngoc Thang Vu;Jonas Kuhn
中科院分区:
其他
文献类型:
--
作者:
Xiang Yu;Ngoc Thang Vu;Jonas Kuhn

文献摘要

被引文献

相似文献

广义Dyck语言已被用于分析递归神经网络(RNN)学习上下文无关语法(CFG)的能力。最近的研究得出相互矛盾的结论,他们的表现,特别是关于递归的深度模型的概括性。在本文中,我们回顾了几个常见的模型和实验设置,讨论了任务和分析的潜在问题。此外,我们还探索了在seq2seq框架内使用注意力机制来学习Dyck语言,这可以弥补RNN有限的编码能力。我们的研究结果表明,注意机制仍然不能真正概括的递归深度,虽然他们表现得比其他模型的闭括号标记任务。此外,这也表明,这种常用的任务是不够的,以测试模型的理解CFGs。
The generalized Dyck language has been used to analyze the ability of Recurrent Neural Networks (RNNs) to learn context-free grammars (CFGs). Recent studies draw conflicting conclusions on their performance, especially regarding the generalizability of the models with respect to the depth of recursion. In this paper, we revisit several common models and experimental settings, discuss the potential problems of the tasks and analyses. Furthermore, we explore the use of attention mechanisms within the seq2seq framework to learn the Dyck language, which could compensate for the limited encoding ability of RNNs. Our findings reveal that attention mechanisms still cannot truly generalize over the recursion depth, although they perform much better than other models on the closing bracket tagging task. Moreover, this also suggests that this commonly used task is not sufficient to test a model’s understanding of CFGs.