Artificial intelligence outperforms pulmonologists in the interpretation of pulmonary function tests

Artificial intelligence outperforms pulmonologists in the interpretation of pulmonary function tests
复制标题

DOI:
10.1183/13993003.01660-2018
复制
发表时间:
2019-04-01
影响因子:
24.3
通讯作者:
Vanhauwaert, A.
Vanhauwaert, A.
中科院分区:
医学1区
文献类型:
--
作者:
Topalovic, Marko;Das, Nilakash;Vanhauwaert, A.

文献摘要

被引文献

相似文献

对肺功能测试(pft)诊断呼吸系统疾病的解释是建立在专家意见的基础上的,这种意见依赖于对模式的识别和检测特定疾病的临床背景。在这项研究中,我们旨在探索肺科医生在解释pft时的准确性和解释器可变性,并将其与基于人工智能(AI)的软件进行比较,该软件已在1500多例历史患者中开发和验证。来自欧洲16家医院的120名肺科医生评估了50例PFT病例和临床信息,产生了6000个独立的解释。人工智能软件检查了同样的数据。美国胸科学会/欧洲呼吸学会指南被用作PFT模式解释的金标准。诊断的金标准来源于临床病史、PFT和所有其他检查。肺科医生(高年级73%,低年级27%)对pft的模式识别在74.4 +/- 5.9%的病例(范围56-88%)中符合指南。kappa=0.67的互变率表明了一个共同的共识。肺科医生的诊断正确率为44.6% +/- 8.7%(范围24-62%),具有较大的间变率(kappa=0.35)。基于人工智能的软件完全匹配PFT模式解释(100%),并在82%的所有病例中给出正确的诊断(p
The interpretation of pulmonary function tests (PFTs) to diagnose respiratory diseases is built on expert opinion that relies on the recognition of patterns and the clinical context for detection of specific diseases. In this study, we aimed to explore the accuracy and interrater variability of pulmonologists when interpreting PFTs compared with artificial intelligence (AI)-based software that was developed and validated in more than 1500 historical patient cases.120 pulmonologists from 16 European hospitals evaluated 50 cases with PFT and clinical information, resulting in 6000 independent interpretations. The AI software examined the same data. American Thoracic Society/European Respiratory Society guidelines were used as the gold standard for PFT pattern interpretation. The gold standard for diagnosis was derived from clinical history, PFT and all additional tests.The pattern recognition of PFTs by pulmonologists (senior 73%, junior 27%) matched the guidelines in 74.4 +/- 5.9% of the cases (range 56-88%). The interrater variability of kappa=0.67 pointed to a common agreement. Pulmonologists made correct diagnoses in 44.6 +/- 8.7% of the cases (range 24-62%) with a large interrater variability (kappa=0.35). The AI-based software perfectly matched the PFT pattern interpretations (100%) and assigned a correct diagnosis in 82% of all cases (p