MEDFAIR: Benchmarking Fairness for Medical Imaging

MEDFAIR: Benchmarking Fairness for Medical Imaging
复制标题

DOI:
10.48550/arxiv.2210.01725
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yongshuo Zong;Yongxin Yang;Timothy M. Hospedales
Yongshuo Zong;Yongxin Yang;Timothy M. Hospedales
中科院分区:
其他
文献类型:
--
作者:
Yongshuo Zong;Yongxin Yang;Timothy M. Hospedales

文献摘要

被引文献

相似文献

大量的研究表明,基于机器学习的医疗诊断系统可能会对某些人群产生偏见。这激发了越来越多的偏见缓解算法,旨在解决机器学习中的公平性问题。然而,由于两个原因,很难比较它们在医学成像中的有效性。首先,在评估公平性的标准上几乎没有达成共识。其次,现有的偏见缓解算法是在不同的设置下开发的,例如数据集、模型选择策略、主干和公平性指标,使得基于现有结果的直接比较和评估成为不可能的。在这项工作中,我们引入了MEDFAIR,这是一个框架,用于对医学成像机器学习模型的公平性进行基准测试。MEDFAIR涵盖了来自不同类别的11种算法,来自不同成像模式的9个数据集和3个模型选择标准。通过大量的实验,我们发现模型选择标准问题对公平结果有显著影响;而相比之下,最先进的偏差缓解算法在分布内和分布外设置中都没有显著提高公平结果,而不是经验风险最小化(ERM)。我们从不同的角度对公平性进行评估,并针对不同的医疗应用场景提出不同的伦理原则建议。我们的框架为深度学习中未来偏见缓解算法的开发和评估提供了一个可重复且易于使用的切入点。代码可从https://github.com/ys-zong/MEDFAIR获得。
A multitude of work has shown that machine learning-based medical diagnosis systems can be biased against certain subgroups of people. This has motivated a growing number of bias mitigation algorithms that aim to address fairness issues in machine learning. However, it is difficult to compare their effectiveness in medical imaging for two reasons. First, there is little consensus on the criteria to assess fairness. Second, existing bias mitigation algorithms are developed under different settings, e.g., datasets, model selection strategies, backbones, and fairness metrics, making a direct comparison and evaluation based on existing results impossible. In this work, we introduce MEDFAIR, a framework to benchmark the fairness of machine learning models for medical imaging. MEDFAIR covers eleven algorithms from various categories, nine datasets from different imaging modalities, and three model selection criteria. Through extensive experiments, we find that the under-studied issue of model selection criterion can have a significant impact on fairness outcomes; while in contrast, state-of-the-art bias mitigation algorithms do not significantly improve fairness outcomes over empirical risk minimization (ERM) in both in-distribution and out-of-distribution settings. We evaluate fairness from various perspectives and make recommendations for different medical application scenarios that require different ethical principles. Our framework provides a reproducible and easy-to-use entry point for the development and evaluation of future bias mitigation algorithms in deep learning. Code is available at https://github.com/ys-zong/MEDFAIR.