FedPseudo: Privacy-Preserving Pseudo Value-Based Deep Learning Models for Federated Survival Analysis

FedPseudo: Privacy-Preserving Pseudo Value-Based Deep Learning Models for Federated Survival Analysis
复制标题

DOI:
10.1145/3580305.3599348
复制
发表时间:
2023-08
期刊:
Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Md Mahmudur Rahman;S. Purushotham
Md Mahmudur Rahman;S. Purushotham
中科院分区:
其他
文献类型:
--
作者:
Md Mahmudur Rahman;S. Purushotham

文献摘要

相似文献

生存分析,又名时间-事件分析,对患者护理有广泛的影响。联邦生存分析(FSA)是一种新兴的联邦学习(FL)范式,用于对多个医疗机构的分布式分散数据进行生存分析。FSA使个人医疗机构(称为客户)能够在确保隐私的同时改善其生存预测。然而,由于客户之间的非线性和非IID数据分布,以及审查造成的偏见,FSA面临着挑战。虽然最近的研究已经采用了考克斯比例风险(CoxPH)生存模型的FSA,这些挑战的系统性探索是目前缺乏的。在本文中,我们通过引入FedPseudo来解决这些关键挑战,FedPseudo是一个用于FSA的基于伪值的深度学习框架。FedPseudo使用深度学习模型从非线性生存数据中学习鲁棒的表示,利用伪值的力量来处理非均匀删失,并采用FedAvg等FL算法来学习模型参数。我们提出了一种新的和简单的方法来估计FSA的伪值。我们提供了理论证明,估计的伪值,称为联邦伪值,是一致的。此外,我们的实证结果表明,它们可以计算速度比传统的方法推导伪值。为了保证和增强估计伪值和共享模型参数的隐私性,系统地研究了差分隐私(DP)在联邦伪值和局部模型更新中的应用.此外,我们采用V -可用信息度量来量化客户端数据的信息量,用于训练生存模型,并利用该度量来显示参与FSA的优势。我们对合成和真实世界的生存数据集进行了广泛的实验,以证明我们的FedPseudo框架比其他FSA方法具有更好的性能,并且与最佳集中训练的深度生存模型相似。此外,FedPseudo在不同的删失设置中始终获得上级结果。
Survival analysis, aka time-to-event analysis, has a wide-ranging impact on patient care. Federated Survival Analysis (FSA) is an emerging Federated Learning (FL) paradigm for performing survival analysis on distributed decentralized data available at multiple medical institutions. FSA enables individual medical institutions, referred to as clients, to improve their survival predictions while ensuring privacy. However, FSA faces challenges due to non-linear and non-IID data distributions among clients, as well as bias caused by censoring. Although recent studies have adapted Cox Proportional Hazards (CoxPH) survival models for FSA, a systematic exploration of these challenges is currently lacking. In this paper, we address these critical challenges by introducing FedPseudo, a pseudo value-based deep learning framework for FSA. FedPseudo uses deep learning models to learn robust representations from non-linear survival data, leverages the power of pseudo values to handle non-uniform censoring, and employs FL algorithms such as FedAvg to learn model parameters. We propose a novel and simple approach for estimating pseudo values for FSA. We provide theoretical proof that the estimated pseudo values, referred to as Federated Pseudo Values, are consistent. Moreover, our empirical results demonstrate that they can be computed faster than traditional methods of deriving pseudo values. To ensure and enhance the privacy of both the estimated pseudo values and the shared model parameters, we systematically investigate the application of differential privacy (DP) on both the federated pseudo values and local model updates. Furthermore, we adapt V -Usable Information metric to quantify the informativeness of a client's data for training a survival model and utilize this metric to show the advantages of participating in FSA. We conducted extensive experiments on synthetic and real-world survival datasets to demonstrate that our FedPseudo framework achieves better performance than other FSA approaches and performs similarly to the best centrally trained deep survival model. Moreover, FedPseudo consistently achieves superior results across different censoring settings.