Long-Tailed Classification of Thorax Diseases on Chest X-Ray: A New Benchmark Study.

Long-Tailed Classification of Thorax Diseases on Chest X-Ray: A New Benchmark Study.
复制标题

DOI:
10.1007/978-3-031-17027-0_3
复制
发表时间:
2022-09
期刊:
Data augmentation, labelling, and imperfections : second MICCAI workshop, DALI 2022, held in conjunction with MICCAI 2022, Singapore, September 22, 2022, proceedings. DALI (Workshop) (2nd : 2022 : Singapore)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

影像学检查,如胸部x线摄影,会产生少量的常见发现和大量的不常见发现。虽然训练有素的放射科医生可以通过研究一些代表性的例子来学习罕见疾病的视觉表现,但教机器从这种“长尾”分布中学习要困难得多,因为标准方法很容易偏向于最频繁的类别。在本文中,我们提出了一个全面的基准研究的长尾学习问题在特定领域的胸部疾病的胸部x光片。我们专注于从自然分布的胸部x射线数据中学习,优化分类精度,不仅包括常见的“头部”类别,还包括罕见但关键的“尾部”类别。为了实现这一目标,我们引入了一个具有挑战性的新的长尾胸片基准,以促进医学图像分类的长尾学习方法的研究。该基准包括用于19路和20路胸腔疾病分类的两个胸部x射线数据集,包含多达53,000个标记训练图像的类和少至7个标记训练图像的类。我们在这个新的基准上评估了标准和最先进的长尾学习方法,分析了这些方法中哪些方面最有利于长尾医学图像分类,并总结了对未来算法设计的见解。数据集、训练模型和代码可在https://github.com/VITA-Group/LongTailCXR上获得。
Imaging exams, such as chest radiography, will yield a small set of common findings and a much larger set of uncommon findings. While a trained radiologist can learn the visual presentation of rare conditions by studying a few representative examples, teaching a machine to learn from such a “long-tailed” distribution is much more difficult, as standard methods would be easily biased toward the most frequent classes. In this paper, we present a comprehensive benchmark study of the long-tailed learning problem in the specific domain of thorax diseases on chest X-rays. We focus on learning from naturally distributed chest X-ray data, optimizing classification accuracy over not only the common “head” classes, but also the rare yet critical “tail” classes. To accomplish this, we introduce a challenging new long-tailed chest X-ray benchmark to facilitate research on developing long-tailed learning methods for medical image classification. The benchmark consists of two chest X-ray datasets for 19- and 20-way thorax disease classification, containing classes with as many as 53,000 and as few as 7 labeled training images. We evaluate both standard and state-of-the-art long-tailed learning methods on this new benchmark, analyzing which aspects of these methods are most beneficial for long-tailed medical image classification and summarizing insights for future algorithm design. The datasets, trained models, and code are available at https://github.com/VITA-Group/LongTailCXR.