Comparing Rating Scales and Preference Judgements in Language Evaluation

Comparing Rating Scales and Preference Judgements in Language Evaluation
复制标题

DOI:
--
复制
发表时间:
2010-07
影响因子:
9.9
通讯作者:
A. Belz;Eric Kow
A. Belz;Eric Kow
中科院分区:
化学1区
文献类型:
--
作者:
A. Belz;Eric Kow

文献摘要

被引文献

相似文献

评级量表评估在 NLP 中很常见,但由于多种原因存在问题,例如它们对于评估者来说可能不直观,评估者之间的一致性和自我一致性往往较低,并且通常应用于结果的参数统计通常被认为不适用于序数数据。在本文中,我们将评级量表与另一种评估范式——偏好强度判断实验(PJE)进行比较,其中评估者的任务更简单,即根据给定的质量标准决定两个文本中哪一个更好。我们提出了三对评估实验,评估不同数据集的文本流畅性和清晰度,其中每对实验之一是评级量表实验,另一个是 PJE。我们发现 PJE 版本的实验具有更好的评估者自我一致性和评估者间一致性,并且系统差异所占的变异比例更大,从而导致发现更多的显着差异。
Rating-scale evaluations are common in NLP, but are problematic for a range of reasons, e.g. they can be unintuitive for evaluators, inter-evaluator agreement and self-consistency tend to be low, and the parametric statistics commonly applied to the results are not generally considered appropriate for ordinal data. In this paper, we compare rating scales with an alternative evaluation paradigm, preference-strength judgement experiments (PJEs), where evaluators have the simpler task of deciding which of two texts is better in terms of a given quality criterion. We present three pairs of evaluation experiments assessing text fluency and clarity for different data sets, where one of each pair of experiments is a rating-scale experiment, and the other is a PJE. We find the PJE versions of the experiments have better evaluator self-consistency and inter-evaluator agreement, and a larger proportion of variation accounted for by system differences, resulting in a larger number of significant differences being found.