The Subgroup Imperative: Chest Radiograph Classifier Generalization Gaps in Patient, Setting, and Pathology Subgroups

The Subgroup Imperative: Chest Radiograph Classifier Generalization Gaps in Patient, Setting, and Pathology Subgroups
复制标题

DOI:
10.1148/ryai.220270
复制
发表时间:
2023-09-01
期刊:
RADIOLOGY-ARTIFICIAL INTELLIGENCE
影响因子:
--
通讯作者:
Fine, Benjamin
Fine, Benjamin
中科院分区:
其他
文献类型:
--
作者:
Ahluwalia, Monish;Abdalla, Mohamed;Fine, Benjamin

文献摘要

被引文献

相似文献

目的:在一个大的、多样化的、真实世界的数据集上,用稳健的亚组分析对四个胸片分类器进行外部测试。材料和方法:在这项回顾性研究中,提取了加拿大安大略省Trillium Health Partners的成人后前位胸片(2016-01-2020-12)和相关的放射学报告。一个开放源码的自然语言处理工具在当地得到验证,并用于根据相关的放射学报告为197 540个图像数据集生成地面真实标签。四个分类器在每张胸片上产生预测。使用准确性、阳性预测值、阴性预测值、敏感度、特异度、F1评分和Matthews相关系数对整个数据集以及患者、环境和病理亚组进行评估。结果:分类器对外部测试数据集的准确率为68%~77%,为%~75%,特异度为82%~94%。算法显示,对于孤立性发现(43%-65%)、40岁以下患者(27%-39%)和急诊科患者(38%-60%)的敏感度降低,对带有支持装置的正常胸片的特异性降低(59%-85%)。性别和血统的差异代表着沿着算法的接收器操作特征曲线的移动。结论:深度学习胸片分类器的性能受到患者、环境和病理因素的影响,这表明亚组分析是必要的,以便为实施提供信息并监控持续性能,以确保最佳质量、安全性和公平性。
Purpose: To externally test four chest radiograph classifiers on a large, diverse, real-world dataset with robust subgroup analysis. Materials and Methods: In this retrospective study, adult posteroanterior chest radiographs (January 2016-December 2020) and associ-ated radiology reports from Trillium Health Partners in Ontario, Canada, were extracted and de-identified. An open-source natural language processing tool was locally validated and used to generate ground truth labels for the 197 540-image dataset based on the associated radiology report. Four classifiers generated predictions on each chest radiograph. Performance was evaluated using accuracy, positive predictive value, negative predictive value, sensitivity, specificity, F1 score, and Matthews correlation coefficient for the overall dataset and for patient, setting, and pathology subgroups. Results: Classifiers demonstrated 68%-77% accuracy, 64%-75% sensitivity, and 82%-94% specificity on the external testing dataset. Algorithms showed decreased sensitivity for solitary findings (43%-65%), patients younger than 40 years (27%-39%), and patients in the emergency department (38%-60%) and decreased specificity on normal chest radiographs with support devices (59%-85%). Differences in sex and ancestry represented movements along an algorithm's receiver operating characteristic curve. Conclusion: Performance of deep learning chest radiograph classifiers was subject to patient, setting, and pathology factors, demonstrat-ing that subgroup analysis is necessary to inform implementation and monitor ongoing performance to ensure optimal quality, safety, and equity.