COGS: A Compositional Generalization Challenge Based on Semantic Interpretation

COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
复制标题

DOI:
10.18653/v1/2020.emnlp-main.731
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Najoung Kim;Tal Linzen
Najoung Kim;Tal Linzen
中科院分区:
其他
文献类型:
--
作者:
Najoung Kim;Tal Linzen

文献摘要

相似文献

自然语言的特点是组合性:一个复杂表达的意义是由其组成部分的意义构成的。为了便于评估语言处理体系结构的组合能力,我们引入了基于英语片段的语义分析数据集COGS。COGS的评价部分包含多个系统缺口,只能通过成分概化来解决;这些包括熟悉的句法结构的新组合,或者熟悉的单词和熟悉的结构的新组合。在变压器和lstm的实验中,我们发现COGS测试集上的分布精度接近完美(96—99%),但泛化精度明显较低(16—35%),并且对随机种子($\pm$6—8%)表现出很高的敏感性。这些发现表明,当代标准NLP模型的成分泛化能力有限,而COGS是衡量进展的好方法。
Natural language is characterized by compositionality: the meaning of a complex expression is constructed from the meanings of its constituent parts. To facilitate the evaluation of the compositional abilities of language processing architectures, we introduce COGS, a semantic parsing dataset based on a fragment of English. The evaluation portion of COGS contains multiple systematic gaps that can only be addressed by compositional generalization; these include new combinations of familiar syntactic structures, or new combinations of familiar words and familiar structures. In experiments with Transformers and LSTMs, we found that in-distribution accuracy on the COGS test set was near-perfect (96--99%), but generalization accuracy was substantially lower (16--35%) and showed high sensitivity to random seed ($\pm$6--8%). These findings indicate that contemporary standard NLP models are limited in their compositional generalization capacity, and position COGS as a good way to measure progress.