Reducing bias and increasing utility by federated generative modeling of medical images using a centralized adversary

Reducing bias and increasing utility by federated generative modeling of medical images using a centralized adversary
复制标题

通过使用集中对手对医学图像进行联合生成建模来减少偏差并提高效用

DOI:
--
复制
发表时间:
2021
期刊:
Conference on Information Technology for Social Good
影响因子:
--
通讯作者:
R. Ng
R. Ng
中科院分区:
--
文献类型:
--
作者:
J. Rajotte;S. Mukherjee;Caleb Robinson;Anthony Ortiz;Christopher West;J. Ferres;R. Ng

文献摘要

参考文献

被引文献

相似文献

由于隐私问题,医疗保健机器学习的一个主要障碍是数据无法广泛共享。保护隐私的合成数据生成越来越被视为解决这一问题的方法。然而,由于医疗保健数据通常具有显着的特定于站点的偏差,因此当目标是利用来自多个站点的数据进行机器学习模型训练时,就会促使使用联邦学习。在这里,我们介绍 FELICIA(FEderated Learning with a CentralIzed Adversary),这是一种支持协作学习的生成机制。它是(本地)PrivGAN 机制的广义扩展,允许考虑联合站点的多样性(非 IID)性质。特别是,我们展示了数据有限且存在偏见的网站如何从其他网站中受益,同时保持所有来源数据的私密性。 FELICIA 适用于生成对抗网络 (GAN) 架构的大家族,包括本工作中演示的普通 GAN 和条件 GAN。我们表明,通过使用 FELICIA 机制,图像数量有限的站点可以生成具有改进实用性的高质量合成图像,而没有一个站点需要提供对其真实数据的访问。共享仅通过中央鉴别器进行,其访问权限仅限于合成数据。我们使用基准图像数据集(MNIST、CIFAR-10)以及用于皮肤病变分类任务的医学图像在几个现实的医疗保健场景中展示了这些好处。我们证明 FELICIA 生成的合成图像的效用超过了本地可用数据的效用,并且我们证明它可以纠正类内有偏见的子组的效用降低。
A major roadblock in machine learning for healthcare is the inability of data to be shared broadly, due to privacy concerns. Privacy preserving synthetic data generation is increasingly being seen as a solution to this problem. However, since healthcare data often has significant site-specific biases, it has motivated the use of federated learning when the goal is to utilize data from multiple sites for machine learning model training. Here, we introduce FELICIA (FEderated LearnIng with a CentralIzed Adversary), a generative mechanism enabling collaborative learning. It is a generalized extension of the (local) PrivGAN mechanism allowing to take into account the diversity (non-IID) nature of the federated sites. In particular, we show how a site with limited and biased data could benefit from other sites while keeping data from all the sources private. FELICIA works for a large family of Generative Adversarial Networks (GAN) architectures including vanilla and conditional GANs as demonstrated in this work. We show that by using the FELICIA mechanism, a site with a limited amount of images can generate high-quality synthetic images with improved utility, while none of the sites need to provide access to their real data. The sharing happens solely through a central discriminator with access limited to synthetic data. We demonstrate these benefits on several realistic healthcare scenarios using benchmark image datasets (MNIST, CIFAR-10) as well as on medical images for the task of skin lesion classification. We show that the utility of synthetic images generated by FELICIA surpasses that of the data available locally and we demonstrate that it can correct the reduced utility of a biased subgroup within a class.
观点:学习分享。
DOI: 10.1038/509s68a
发表时间: 2014
期刊: Nature
影响因子: 64.8
作者:
Quackenbush,John
通讯作者: Quackenbush,John