Reliability of Surgical Risk Calculator Performance Assessment in Single-Institution Data.

Reliability of Surgical Risk Calculator Performance Assessment in Single-Institution Data.
复制标题

单一机构数据中手术风险计算器性能评估的可靠性。

DOI:
10.1016/j.jamcollsurg.2020.09.013
复制
发表时间:
2020
影响因子:
5.2
通讯作者:
Merkow,RyanP
Merkow,RyanP
中科院分区:
医学2区
文献类型:
--
作者:
Fischer,ChelseaP;Cohen,MarkE;Merkow,RyanP

文献摘要

相似文献

手术风险计算器(SRC)是临床医生常用的工具,最初是出于在手术前准确预测患者风险的需要而开发的,用于指导临床医生决策并帮助患者咨询。SRC允许临床医生输入患者因素,例如人口统计学特征和合并症,并基于患者的个体化风险状况生成结局预测。第一个大规模SRC由美国外科医生学会根据NSQIP数据登记处开发,包括1,500多种不同手术的风险概况。1 SRC在临床和研究环境中的使用随着时间的推移而扩大,其准确性已在多种外科手术中进行了评估。Vos及其同事2检查了美国外科医师学会NSQIP SRC在胃癌全胃切除术患者中的表现。作者通过验证机构收集的数据库的估计值,检查了SRC对12种不良结局的预测准确性。SRC的性能被发现是不一致的,在接受全胃切除术的胃癌患者,与整体的并发症预测不足。发现SRC在死亡、肾衰竭、心脏并发症和出院至康复机构/疗养院方面表现良好。重要的是要确定SRC具有良好预测准确性的条件,以及它可能不足以指导提供者和患者的条件。然而,在将这些结果转化为现实世界的实践时,需要承认这项研究存在重要的局限性。本研究在一家机构进行,纳入了少数接受单一复杂手术的患者。小样本量的事件率估计存在固有的不可靠性,这使得难以验证SRC。3只有作者的任何并发症结局、手术部位感染和住院时间结局有足够数量的事件病例(超过100例)才能提供充分的检验。4 SRC旨在对在NSQIP平均医院接受治疗的患者进行预测,当应用于与平均水平相差很大的医院时,性能将下降。4这可能与本研究特别相关,因为所有病例均在高度专业化的医院进行,仅治疗癌症患者。因此,在肯定不能代表美国几乎所有医院的环境中验证SRC可能是不合适的。最后,将SRC的外部验证限制在特定程序中会减少观察到的歧视,因为接受相同程序的患者相对同质,具有相似的风险特征。4,5这些限制并不意味着SRC的外部验证不值得奋进。对手术特定变量的研究可以提高计算器的性能,并为临床医生增加临床相关性。理想情况下,外部验证应该基于大型的多机构数据集,因为来自单个机构的低容量率估计往往不稳定。重要的是,SRC的设计目的不是为了产生完美的预测,它们不应该取代机构经验,周到的患者选择和以患者为中心的决策。SRC是纳入临床实践并指导以患者为中心的决策的有价值的工具。使用这些工具的执业临床医生应该了解它们的局限性和有临床意义的使用场景。
Surgical risk calculators (SRCs) are tools commonly used by clinicians that were initially developed out of the need for accurate patient risk prediction before operations, to both guide clinician decision-making and aid in patient counseling. SRCs allow clinicians to input patient factors, such as demographic characteristics and comorbidities, and generate outcomes predications based on the patient’s individualized risk profile. The first large-scale SRC, developed by the American College of Surgeons from the NSQIP data registry, includes risk profiles for more than 1,500 different procedures. 1 SRC use in clinical and research settings has expanded over time and its accuracy has been evaluated for multiple surgical procedures. Vos and colleagues 2 examined the performance of the American College of Surgeons NSQIP SRC in patients undergoing total gastrectomy for gastric cancer. The authors examined the predictive accuracy of the SRC for 12 adverse outcomes by validating the estimates against an institutionally collected database. Performance of the SRC was found to be inconsistent in patients undergoing total gastrectomy for gastric cancer, with underprediction of complications overall. The SRC was noted to perform well for death, renal failure, cardiac complication, and discharge to rehabilitation facility/nursing home. It is important to determine the conditions under which the SRC has good predictive accuracy and where it might not be sufficiently reliable to guide providers and patients. However, there are important limitations in this study that need to be acknowledged when translating these results to real-world practice. This study was performed in a single institution and included a small number of patients undergoing a single complex procedure. There is inherent unreliability of event rate estimates from small sample sizes, which makes it difficult to validate the SRC. 3 Only the authors’ outcomes of any complication, surgical site infection, and length of stay outcomes have a sufficient number of cases with events (more than 100) to provide an adequate test. 4 The SRC is designed to make predictions for patients treated at the average NSQIP hospital and performance will decline when applied to hospitals that diverge substantially from the average. 4 This might be particularly relevant in this study, as all cases were performed at a highly specialized hospital treating cancer patients only. Therefore, it might be inappropriate to validate an SRC in a setting that is certainly not representative of almost all hospitals in the US. Finally, limiting an external validation of the SRC to a specific procedure reduces observed discrimination, as patients undergoing the same procedure are relatively homogeneous with similar risk profiles. 4, 5These limitations are not to imply external validation of the SRC is not a worthwhile endeavor. Investigation of procedure-specific variables can improve the performance of calculators and add clinical relevance for clinicians. Ideally, external validation should be based on large, multi-institution data sets, as rate estimates from single institutions with low volume tend to be unstable. Importantly, SRCs are not designed to generate perfect predictions, and they should not be a replacement for institution experience, thoughtful patient selection, and patientcentered decision-making. SRCs are valuable tools to incorporate into clinical practice and guide patient-centered decision-making. Practicing clinicians who use such tools should understand both their limitations and the scenarios for clinically meaningful use.