Computing inter-rater reliability and its variance in the presence of high agreement

Computing inter-rater reliability and its variance in the presence of high agreement
复制标题

DOI:
10.1348/000711006x126600
复制
发表时间:
2008-05-01
影响因子:
2.6
通讯作者:
Gwet, Kilem Li
Gwet, Kilem Li
中科院分区:
心理学3区
文献类型:
--
作者:
Gwet, Kilem Li

文献摘要

被引文献

相似文献

Pi(pi)和kappa(kappa)统计量被广泛用于精神病学和心理测试领域,以计算评分者对标称标度数据的一致程度。事实上,这些系数偶尔会在被称为卡帕悖论的情况下产生意想不到的结果。本文探讨了这些局限性的根源,并介绍了一种替代的和更稳定的协议系数称为AC(1)系数。本文还提出了多评分者广义pi和AC(1)统计量的新方差估计,其有效性不依赖于评分者之间的独立性假设。这是对现有替代方差的改进,现有替代方差取决于独立性假设。一个蒙特-卡罗模拟研究证明了这些方差估计的有效性置信区间的建设,并确认AC(1)的值作为一个改进的替代现有的评分者间的可靠性统计。
Pi (pi) and kappa (kappa) statistics are widely used in the areas of psychiatry and psychological testing to compute the extent of agreement between raters on nominally scaled data. It is a fact that these coefficients occasionally yield unexpected results in situations known as the paradoxes of kappa. This paper explores the origin of these limitations, and introduces an alternative and more stable agreement coefficient referred to as the AC(1) coefficient. Also proposed are new variance estimators for the multiple-rater generalized pi and AC(1) statistics, whose validity does not depend upon the hypothesis of independence between raters. This is an improvement over existing alternative variances, which depend on the independence assumption. A Monte-Carlo simulation study demonstrates the validity of these variance estimators for confidence interval construction, and confirms the value of AC(1) as an improved alternative to existing inter-rater reliability statistics.