Targeted Syntactic Evaluation of Language Models

Targeted Syntactic Evaluation of Language Models
复制标题

语言模型的针对性句法评估

DOI:
--
复制
发表时间:
2018
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Tal Linzen
Tal Linzen
中科院分区:
--
文献类型:
--
作者:
Rebecca Marvin;Tal Linzen

文献摘要

参考文献

被引文献

相似文献

我们提出了一个用于评估语言模型预测的语法性的数据集。我们自动构建大量不同的英语句子对,每个句子都包含一个语法句子和一个不语法句子。句子对代表结构敏感现象的不同变化:主谓一致、反身照应和负极性项目。我们期望语言模型为符合语法的句子分配比不符合语法的句子更高的概率。在使用该数据集的实验中,LSTM 语言模型在许多结构上表现不佳。具有句法目标(CCG 超级标签)的多任务训练提高了 LSTM 的准确性,但其性能与在线招募的人类参与者的准确性之间仍然存在很大差距。这表明在捕获语言模型中的语法方面,与 LSTM 相比还有相当大的改进空间。
We present a data set for evaluating the grammaticality of the predictions of a language model. We automatically construct a large number of minimally different pairs of English sentences, each consisting of a grammatical and an ungrammatical sentence. The sentence pairs represent different variations of structure-sensitive phenomena: subject-verb agreement, reflexive anaphora and negative polarity items. We expect a language model to assign a higher probability to the grammatical sentence than the ungrammatical one. In an experiment using this data set, an LSTM language model performed poorly on many of the constructions. Multi-task training with a syntactic objective (CCG supertagging) improved the LSTM’s accuracy, but a large gap remained between its performance and the accuracy of human participants recruited online. This suggests that there is considerable room for improvement over LSTMs in capturing syntax in a language model.
DOI: 10.1162/tacl_a_00290
发表时间: 2019-01-01
影响因子: 10.9
作者:
Warstadt, Alex;Singh, Amanpreet;Bowman, Samuel R.
通讯作者: Bowman, Samuel R.
DOI: --
发表时间: 2010-08
期刊: --
影响因子: --
作者:
Joakim Nivre;Laura Rimell;Ryan T. McDonald;Carlos Gómez-Rodríguez
通讯作者: Joakim Nivre;Laura Rimell;Ryan T. McDonald;Carlos Gómez-Rodríguez
DOI: 10.3115/1699571.1699619
发表时间: 2009-08
期刊: --
影响因子: --
作者:
Laura Rimell;S. Clark;Mark Steedman
通讯作者: Laura Rimell;S. Clark;Mark Steedman