Enhancing thoracic disease detection using chest X-rays from PubMed Central Open Access

Enhancing thoracic disease detection using chest X-rays from PubMed Central Open Access
复制标题

DOI:
10.1016/j.compbiomed.2023.106962
复制
发表时间:
2023-04-23
影响因子:
7.7
通讯作者:
Peng, Yifan
Peng, Yifan
中科院分区:
工程技术2区
文献类型:
--
作者:
Lin, Mingquan;Hou, Bojian;Peng, Yifan

文献摘要

被引文献

相似文献

已经收集了大型胸部X射线(CXR)数据集来训练深度学习模型,以检测CXR上的胸部病理。然而,大多数CXR数据集来自单中心研究,并且收集的病理通常不平衡。本研究的目的是根据PubMed Central Open Access(PMC-OA)中的文章自动构建一个公共的弱标记CXR数据库,并通过使用该数据库作为额外的训练数据来评估CXR病理分类的模型性能。我们的框架包括文本提取,CXR病理验证,子图分离,和图像模态分类。我们已经广泛验证了自动生成的图像数据库在胸部疾病检测任务中的实用性,包括疝、肺部病变、肺炎和气胸。我们选择这些疾病是因为它们在现有数据集中的历史表现不佳:NIH-CXR数据集(112,120 CXR)和MIMIC-CXR数据集(243,324 CXR)。我们发现,使用由所提出的框架提取的额外PMC-CXR进行微调的分类器一致且显着地实现了比没有PMC-CXR的分类器更好的性能(例如,疝气:0.9335 vs. 0.9154;肺部病变:0.7394 vs. 0.7207;肺炎:0.7074 vs. 0.6709;气胸0.8185 vs. 0.7517,均为AUC,p < 0.0001)。与以前手动将医学图像提交到存储库的方法相比,我们的框架可以自动收集数字及其伴随的数字图例。与以前的研究相比,该框架改进了子图分割,并结合了我们先进的自主开发的NLP技术用于CXR病理验证。我们希望它能补充现有的资源,并提高我们的能力,使生物医学图像数据可查找,可访问,可互操作和可重用。
Large chest X-rays (CXR) datasets have been collected to train deep learning models to detect thorax pathology on CXR. However, most CXR datasets are from single-center studies and the collected pathologies are often imbalanced. The aim of this study was to automatically construct a public, weakly-labeled CXR database from articles in PubMed Central Open Access (PMC-OA) and to assess model performance on CXR pathology classification by using this database as additional training data. Our framework includes text extraction, CXR pathology verification, subfigure separation, and image modality classification. We have extensively validated the utility of the automatically generated image database on thoracic disease detection tasks, including Hernia, Lung Lesion, Pneumonia, and pneumothorax. We pick these diseases due to their historically poor performance in existing datasets: the NIH-CXR dataset (112,120 CXR) and the MIMIC-CXR dataset (243,324 CXR). We find that classifiers fine-tuned with additional PMC-CXR extracted by the proposed framework consistently and significantly achieved better performance than those without (e.g., Hernia: 0.9335 vs 0.9154; Lung Lesion: 0.7394 vs. 0.7207; Pneumonia: 0.7074 vs. 0.6709; Pneumothorax 0.8185 vs. 0.7517, all in AUC with p < 0.0001) for CXR pathology detection. In contrast to previous approaches that manually submit the medical images to the repository, our framework can automatically collect figures and their accompanied figure legends. Compared to previous studies, the proposed framework improved subfigure segmentation and incorporates our advanced selfdeveloped NLP technique for CXR pathology verification. We hope it complements existing resources and improves our ability to make biomedical image data findable, accessible, interoperable, and reusable.