Explaining the unsuitability of the kappa coefficient in the assessment and comparison of the accuracy of thematic maps obtained by image classification

Explaining the unsuitability of the kappa coefficient in the assessment and comparison of the accuracy of thematic maps obtained by image classification
复制标题

DOI:
10.1016/j.rse.2019.111630
复制
发表时间:
2020-03-15
影响因子:
13.5
通讯作者:
Foody, Giles M.
Foody, Giles M.
中科院分区:
工程技术1区
文献类型:
--
作者:
Foody, Giles M.

文献摘要

被引文献

相似文献

卡帕系数不是衡量准确度的指标,事实上,它也不是总体一致的指标,而是绝无仅有的一致的指标。然而,偶然性协定在精度评估中是无关紧要的,而且无论如何在计算典型遥感应用的卡帕系数时都是不适当的。卡帕系数的大小也很难解释。可以从满足要求精度目标的分类中获得跨越所有广泛使用的解释比例尺的值,该值表明一致性水平等同于仅由偶然性直到几乎完全一致的水平(例如,对于总体精度为95%的分类,卡帕系数的可能值的范围是-0.026至0.900)。如果不同类别的丰度(即流行率)不同,卡帕系数的比较就特别具有挑战性,因为卡帕系数的大小不仅反映了标签上的一致性,而且反映了所研究人口的性质。结果表明,在精度评估中使用卡帕系数的所有论点都是有缺陷的和/或不相关的,因为它们同样适用于其他有时更容易计算的准确度衡量标准。在评估和比较分类准确度时,应最终注意到要求从准确度评估中摒弃卡帕系数的呼吁,并鼓励研究人员提供一套简单的衡量标准和相关产出,如每类准确度估计数和混淆矩阵。
The kappa coefficient is not an index of accuracy, indeed it is not an index of overall agreement but one of agreement beyond chance. Chance agreement is, however, irrelevant in an accuracy assessment and is anyway inappropriately modelled in the calculation of a kappa coefficient for typical remote sensing applications. The magnitude of a kappa coefficient is also difficult to interpret. Values that span the full range of widely used interpretation scales, indicating a level of agreement that equates to that estimated to arise from chance alone all the way through to almost perfect agreement, can be obtained from classifications that satisfy demanding accuracy targets (e.g. for a classification with overall accuracy of 95% the range of possible values of the kappa coefficient is -0.026 to 0.900). Comparisons of kappa coefficients are particularly challenging if the classes vary in their abundance (i.e. prevalence) as the magnitude of a kappa coefficient reflects not only agreement in labelling but also properties of the populations under study. It is shown that all of the arguments put forward for the use of the kappa coefficient in accuracy assessment are flawed and/or irrelevant as they apply equally to other, sometimes easier to calculate, measures of accuracy. Calls for the kappa coefficient to be abandoned from accuracy assessments should finally be heeded and researchers are encouraged to provide a set of simple measures and associated outputs such as estimates of per-class accuracy and the confusion matrix when assessing and comparing classification accuracy.