International variation in histologic grading is large, and persistent feedback does not improve reproducibility

International variation in histologic grading is large, and persistent feedback does not improve reproducibility
复制标题

DOI:
10.1097/00000478-200306000-00012
复制
发表时间:
2003-06-01
影响因子:
5.6
通讯作者:
Rosenthal, R
Rosenthal, R
中科院分区:
医学1区
文献类型:
--
作者:
Furness, PN;Taub, N;Rosenthal, R

文献摘要

被引文献

相似文献

组织学分级系统用于指导诊断、治疗和国际审计。分级系统的可重复性通常在以前一起工作或培训的病理学家小组中进行测试。这可能低估了评分系统的国际差异。因此,我们评估了一个既定的系统,肾移植病理学的班夫分类,在整个欧洲的再现性。我们还试图通过在14个小组案例中的每个案例后提供个人反馈来提高可重复性。研究的所有特征的Kappa值均低于先前发表的任何文献,证实了国际变异大于先前评估的观察者间变异。使用数字或图形反馈来提高再现性的长期尝试未能产生任何可检测的改进。然后,我们要求参与者对选定的照片进行分级,以消除病理学家查看幻灯片不同区域引起的差异。这仅对某些功能产生了改进的Kappa值。改进受到职等定义性质的影响。基于受某一进程“影响的地区”的定义没有得到改进。结果表明,基于评分系统的决定是危险的,在不同的机构中,评分系统的应用可能非常不同。
Histologic grading systems are used to guide diagnosis, therapy, and audit on an international basis. The reproducibility of grading systems is usually tested within small groups of pathologists who have previously worked or trained together. This may underestimate the international variation of scoring systems. We therefore evaluated the reproducibility of an established system, the Banff classification of renal allograft pathology, throughout Europe. We also sought to improve reproducibility by providing individual feedback after each of 14 small groups of cases. Kappa values for all features studied were lower than any previously published, confirming that international variation is greater than interobserver variation as previously assessed. A prolonged attempt to improve reproducibility, using numeric or graphical feedback, failed to produce any detectable improvement. We then asked participants to grade selected photographs, to eliminate variation induced by pathologists viewing different areas of the slide. This produced improved kappa values only for some features. Improvement was influenced by the nature of the grade definitions. Definitions based on "area affected" by a process were not improved. The results indicate the danger of basing decisions on grading systems that may be applied very differently in different institutions.