Application of a population-based severity scoring system to individual patients results in frequent misclassification

Application of a population-based severity scoring system to individual patients results in frequent misclassification
复制标题

DOI:
10.1186/cc3790
复制
发表时间:
2005-10-01
期刊:
影响因子:
15.1
通讯作者:
Levy, H
Levy, H
中科院分区:
医学1区
文献类型:
--
作者:
Booth, FV;Short, M;Levy, H

文献摘要

被引文献

相似文献

简介 APACHE II (AP2) 的开发是为了以风险调整的方式对重症监护病房的结果进行系统检查。 AP2 已在临床试验中广泛采用,以确保不同组之间的广泛一致性。尽管计算真实 AP2 分数的误差可能无法降低到 15% 以下,但当应用于大量人群时,随机误差的自我抵消效应会降低此类误差的重要性。有人建议在个体患者的临床决策中使用阈值 AP2 评分。本研究报告了参与大型脓毒症试验的研究人员的 AP2 评分错误,并模拟了这种错误率对个别严重脓毒症患者的后果。方法 56 名接受过数据提取和完成 AP2 评分明确培训的研究人员接受了由真实患者病史组成的场景。针对每个场景计算了描述性统计数据。与判定的分数进行比较来计算标准偏差。使用 Shrout-Fleiss 方法进行观察者间可靠性的组内相关性。使用 6、9 和 12 的标准差计算大范围 AP2 评分的理论分布曲线。对于每条曲线,使用 >= 25 的 AP2 评分截止值确定错误分类率。然后将每个真实 AP2 评分的错误分类百分比应用于从 PROGRESS 严重脓毒症登记处获得的相应 AP2 评分。结果 AP2 总评分的错误率为 86%(各个变量的范围为 10% 到 10%) 87%)。观察者间可靠性的组内相关性为 0.51。来自 PROGRESS 登记处的患者。 50% 的 AP2 评分在 17 到 28 之间。在这个四分位数范围内,所有错误分类的患者中有 70% 到 85% 属于该范围内。结论 个别患者被错误评分的可能性比正确评分的可能性更大。从场景中获得的数据表明,当真实 AP2 分数接近任意截止点 25 时,观察到的错误分类率增加。将我们对 AP2 评分误差的研究与已发表的文献相结合,我们得出结论:AP2 是为个体患者做出资源分配决策的唯一不合适的工具。
Introduction APACHE II (AP2) was developed to allow a systematic examination of intensive care unit outcomes in a risk adjusted manner. AP2 has been widely adopted in clinical trials to assure broad consistency amongst different groups. Although errors in calculating the true AP2 score may not be reducible below 15%, the self-canceling effect of random errors reduces the importance of such errors when applied to large populations. It has been suggested that a threshold AP2 score be used in clinical decision making for individual patients. This study reports the AP2 scoring errors of researchers involved in a large sepsis trial and models the consequences of such an error rate for individual severe sepsis patients.Methods Fifty-six researchers with explicit training in data abstraction and completion of the AP2 score received scenarios consisting of composites of real patient histories. Descriptive statistics were calculated for each scenario. The standard deviations were calculated compared with an adjudicated score. Intraclass correlations for inter-observer reliability were performed using Shrout-Fleiss methodology. Theoretical distribution curves were calculated for a broad range of AP2 scores using standard deviations of 6, 9 and 12. For each curve, the misclassification rate was determined using an AP2 score cut-off of >= 25. The percentage of misclassifications for each true AP2 score was then applied to the corresponding AP2 score obtained from the PROGRESS severe sepsis registry.Results The error rate for the total AP2 score was 86% ( individual variables were in the range 10% to 87%). Intraclass correlation for the inter-observer reliability was 0.51. Of the patients from the PROGRESS registry. 50% had AP2 scores in the range 17 to 28. Within this interquartile range, 70% to 85% of all misclassified patients would reside.Conclusion It is more likely that an individual patient will be scored incorrectly than correctly. The data obtained from the scenarios indicated that as the true AP2 score approached an arbitrary cut-off point of 25, the observed misclassification rate increased. Integrating our study of AP2 score errors with the published literature leads us to conclude that the AP2 is an inappropriate sole tool for resource allocation decisions for individual patients.