Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions

Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions
复制标题

DOI:
10.18653/v1/2020.inlg-1.23
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
David M. Howcroft;Anya Belz;Miruna Clinciu;Dimitra Gkatzia;Sadid A. Hasan;Saad Mahamood;Simon Mille;Emiel van Miltenburg;Sashank Santhanam;Verena Rieser
David M. Howcroft;Anya Belz;Miruna Clinciu;Dimitra Gkatzia;Sadid A. Hasan;Saad Mahamood;Simon Mille;Emiel van Miltenburg;Sashank Santhanam;Verena Rieser
中科院分区:
其他
文献类型:
--
作者:
David M. Howcroft;Anya Belz;Miruna Clinciu;Dimitra Gkatzia;Sadid A. Hasan;Saad Mahamood;Simon Mille;Emiel van Miltenburg;Sashank Santhanam;Verena Rieser

文献摘要

被引文献

相似文献

人类评估仍然是NLG中最值得信赖的评估形式,但是研究人员使用的不同质量标准的高度多样化方法使得很难比较跨论文的结果并得出结论,对元评估和可重复性产生了不利影响。在本文中,我们介绍了(i)通过人类评估的165份NLG论文的数据集,(ii)我们开发的注释方案为评估不同方面的论文标记,(iii)注释的定量分析,以及(iv)一组建议改善评估报告中标准的建议。我们使用注释作为检查评估报告中包含的信息的基础,以及方法,实验设计和术语的一致性水平,尤其是针对200多种不同术语,这些术语已用于评估质量方面。我们得出的结论是,由于报道中缺乏明确性和方法的极端多样性,NLG中的人类评估在2020年非常困惑,并且该领域迫切需要标准方法和术语。
Human assessment remains the most trusted form of evaluation in NLG, but highly diverse approaches and a proliferation of different quality criteria used by researchers make it difficult to compare results and draw conclusions across papers, with adverse implications for meta-evaluation and reproducibility. In this paper, we present (i) our dataset of 165 NLG papers with human evaluations, (ii) the annotation scheme we developed to label the papers for different aspects of evaluations, (iii) quantitative analyses of the annotations, and (iv) a set of recommendations for improving standards in evaluation reporting. We use the annotations as a basis for examining information included in evaluation reports, and levels of consistency in approaches, experimental design and terminology, focusing in particular on the 200+ different terms that have been used for evaluated aspects of quality. We conclude that due to a pervasive lack of clarity in reports and extreme diversity in approaches, human evaluation in NLG presents as extremely confused in 2020, and that the field is in urgent need of standard methods and terminology.