Quantified Reproducibility Assessment of NLP Results

Quantified Reproducibility Assessment of NLP Results
复制标题

NLP 结果的量化再现性评估

DOI:
--
复制
发表时间:
2022
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Simon Mille
Simon Mille
中科院分区:
--
文献类型:
--
作者:
Anya Belz;Maja Popovi'c;Simon Mille

文献摘要

被引文献

相似文献

本文介绍并测试了一种基于计量学概念和定义的量化可重复评估(QRA)的方法。 QRA产生一个单个分数,估计给定系统的可重复性程度和评估度量,基于不同再现之间的分数和差异。我们在18种不同的系统和评估度量组合(涉及不同的NLP任务和评估类型)上测试QRA,为此我们具有原始结果和一到七个复制结果。所提出的QRA方法产生的可重复可复制度得分不仅在多个复制品中相当,而且还具有不同的原始研究。我们发现所提出的方法有助于洞察复制之间变化的原因,因此,可以对系统的哪些方面和/或评估设计的哪些方面得出结论,以提高重复性。
This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology. QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions. We test QRA on 18 different system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results. The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but also of different, original studies. We find that the proposed method facilitates insights into causes of variation between reproductions, and as a result, allows conclusions to be drawn about what aspects of system and/or evaluation design need to be changed in order to improve reproducibility.