Using artificial intelligence to read chest radiographs for tuberculosis detection: A multi-site evaluation of the diagnostic accuracy of three deep learning systems

Using artificial intelligence to read chest radiographs for tuberculosis detection: A multi-site evaluation of the diagnostic accuracy of three deep learning systems
复制标题

DOI:
10.1038/s41598-019-51503-3
复制
发表时间:
2019-10-18
期刊:
影响因子:
4.6
通讯作者:
Creswell, Jacob
Creswell, Jacob
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Qin, Zhi Zhen;Sander, Melissa S.;Creswell, Jacob

文献摘要

被引文献

相似文献

深度学习 (DL) 神经网络最近才被用来解读胸部 X 光片 (CXR),以筛查和分类肺结核 (TB)。目前还没有已发表的研究比较多个深度学习系统和人群。我们对三个 DL 系统(CAD4TB、Lunit INSIGHT 和 qXR)进行了回顾性评估,用于检测尼泊尔和喀麦隆门诊患者胸片中与结核病相关的异常情况。所有 1196 名受试者均接受了 Xpert MTB/RIF 检测,并由两组放射科医生和 DL 系统读取 CXR。 Xpert 被用作参考标准。三个系统的曲线下面积相似:Lunit(0.94,95% CI:0.93-0.96)、qXR(0.94,95% CI:0.92-0.97)和 CAD4TB(0.92,95% CI:0.90-0.95)。当与放射科医生的灵敏度相匹配时,DL 系统的特异性明显更高,只有一个系统除外。使用 DL 系统读取 CXR 可以将所需的 Xpert MTB/RIF 测试数量减少 66%,同时保持 95% 或更高的灵敏度。使用通用截止分数会导致每个站点的表现不同,这凸显了根据筛选人群选择分数的必要性。在人力资源有限且自动化技术可用的结核病项目中应考虑这些深度学习系统。
Deep learning (DL) neural networks have only recently been employed to interpret chest radiography CXR) to screen and triage people for pulmonary tuberculosis (TB). No published studies have compared multiple DL systems and populations. We conducted a retrospective evaluation of three DL systems (CAD4TB, Lunit INSIGHT, and qXR) for detecting TB-associated abnormalities in chest radiographs from outpatients in Nepal and Cameroon. All 1196 individuals received a Xpert MTB/RIF assay and a CXR read by two groups of radiologists and the DL systems. Xpert was used as the reference standard. The area under the curve of the three systems was similar: Lunit (0.94, 95% CI: 0.93-0.96), qXR (0.94, 95% CI: 0.92-0.97) and CAD4TB (0.92, 95% CI: 0.90-0.95). When matching the sensitivity of the radiologists, the specificities of the DL systems were significantly higher except for one. Using DL systems to read CXRs could reduce the number of Xpert MTB/RIF tests needed by 66% while maintaining sensitivity at 95% or better. Using a universal cutoff score resulted different performance in each site, highlighting the need to select scores based on the population screened. These DL systems should be considered by TB programs where human resources are constrained, and automated technology is available.