The meaning of kappa: Probabilistic concepts of reliability and validity revisited

The meaning of kappa: Probabilistic concepts of reliability and validity revisited
复制标题

DOI:
10.1016/0895-4356(96)00011-x
复制
发表时间:
1996-07-01
影响因子:
7.2
通讯作者:
GuggenmoosHolzmann, I
GuggenmoosHolzmann, I
中科院分区:
医学2区
文献类型:
--
作者:
GuggenmoosHolzmann, I

文献摘要

被引文献

相似文献

一个框架——“协议概念”——被开发出来,以统一的方式研究科恩kappa的使用,以及机会修正协议的替代措施。着重于内部一致性,证明了对于2 x 2表,只有在考虑到观测环境的特征时,才能在机会校正的一致性的不同度量之间做出适当的选择。特别是,天真地使用科恩的kappa可能会导致对机会修正后的一致性的过分乐观的估计。这种偏差可以通过更精细的研究设计来克服,这些研究设计允许对所讨论的概率进行不受限制的估计。当科恩的kappa被恰当地应用于衡量机会修正后的一致性时,它的值被证明是真实流行的线性函数,而不是抛物线函数。它进一步显示了评级的有效性是如何受到缺乏一致性的影响。根据效度研究的设计,这可能导致基于纯粹形式理由的敏感性和特异性依赖于流行率的估计。所提出的“机会修正”效度指标公式未能对这一现象进行调整。
A framework-the ''agreement concept''-is developed to study the use of Cohen's kappa as well as alternative measures of chance-corrected agreement in a unified manner. Focusing on intrarater consistency it is demonstrated that for 2 x 2 tables an adequate choice between different measures of chance-corrected agreement can be made only if the characteristics of the observational setting are taken into account. In particular, a naive use of Cohen's kappa may lead to strinkingly overoptimistic estimates of chance-corrected agreement. Such bias can be overcome by more elaborate study designs that allow for an unrestricted estimation of the probabilities at issue. When Cohen's kappa is appropriately applied as a measure of chance-corrected agreement, its values prove to be a linear-and not a parabolic-function of true prevalence. It is further shown how the validity of ratings is influenced by lack of consistency. Depending on the design of a validity study, this may lead, on purely formal grounds, to prevalence dependent estimates of sensitivity and specificity. Proposed formulas for ''chance-corrected'' validity indexes fail to adjust for this phenomenon.