The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
复制标题

DOI:
10.18653/v1/2021.gem-1.10
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Sebastian Gehrmann;Tosin P. Adewumi;Karmanya Aggarwal;Pawan Sasanka Ammanamanchi;Aremu Anuoluwapo;Antoine Bos
Sebastian Gehrmann;Tosin P. Adewumi;Karmanya Aggarwal;Pawan Sasanka Ammanamanchi;Aremu Anuoluwapo;Antoine Bos
中科院分区:
其他
文献类型:
--
作者:
Sebastian Gehrmann;Tosin P. Adewumi;Karmanya Aggarwal;Pawan Sasanka Ammanamanchi;Aremu Anuoluwapo;Antoine Bos

文献摘要

被引文献

相似文献

我们介绍了GEM,一个自然语言生成(NLG)的活基准,它的评估,和测试。衡量NLG的进展依赖于不断发展的自动化指标、数据集和人工评估标准的生态系统。由于这个移动的目标,新模型通常仍然使用建立良好但有缺陷的指标来评估以英语为中心的语料库。这种脱节使得确定当前模式的局限性和取得进展的机会变得困难。针对这一局限性,GEM提供了一个环境,在这个环境中,模型可以很容易地应用于一系列广泛的任务,并在其中可以测试评估策略。定期更新基准将有助于NLG研究变得更加多语言,并与模型一起发展挑战。本文作为相关GEM研讨会2021年共享任务的数据描述。
We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of automated metrics, datasets, and human evaluation standards. Due to this moving target, new models often still evaluate on divergent anglo-centric corpora with well-established, but flawed, metrics. This disconnect makes it challenging to identify the limitations of current models and opportunities for progress. Addressing this limitation, GEM provides an environment in which models can easily be applied to a wide set of tasks and in which evaluation strategies can be tested. Regular updates to the benchmark will help NLG research become more multilingual and evolve the challenge alongside models. This paper serves as the description of the data for the 2021 shared task at the associated GEM Workshop.