Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks

Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks
复制标题

DOI:
10.1162/tacl_a_00304
复制
发表时间:
2020-01-01
影响因子:
10.9
通讯作者:
Linzen, Tal
Linzen, Tal
中科院分区:
人文科学1区
文献类型:
--
作者:
McCoy, R. Thomas;Frank, Robert;Linzen, Tal

文献摘要

被引文献

相似文献

暴露于相同训练数据的学习者可能会由于不同的归纳偏差而产生不同的概括。在神经网络模型中,归纳偏差理论上可能来自模型架构的任何方面。我们研究了哪些结构因素会影响神经序列到序列模型在两个句法任务(英语疑问句形成和英语时态反射)上训练的泛化行为。对于这两个任务,训练集与基于层次结构的泛化和基于线性顺序的泛化一致。我们研究的所有架构因素都定性地影响模型的泛化,包括与层次结构没有明确联系的因素。例如,LSTM和GRU表现出质的不同的归纳偏见。然而,唯一的因素,一贯有助于跨任务的层次偏见是使用一个树结构模型,而不是一个模型与顺序递归,这表明类人的句法概括需要架构的句法结构。
Learners that are exposed to the same training data might generalize differently due to differing inductive biases. In neural network models, inductive biases could in theory arise from any aspect of the model architecture. We investigate which architectural factors affect the generalization behavior of neural sequence-to-sequence models trained on two syntactic tasks, English question formation and English tense reinflection. For both tasks, the training set is consistent with a generalization based on hierarchical structure and a generalization based on linear order. All architectural factors that we investigated qualitatively affected how models generalized, including factors with no clear connection to hierarchical structure. For example, LSTMs and GRUs displayed qualitatively different inductive biases. However, the only factor that consistently contributed a hierarchical bias across tasks was the use of a treestructured model rather than a model with sequential recurrence, suggesting that humanlike syntactic generalization requires architectural syntactic structure.