Federated Learning for Sparse Bayesian Models with Applications to Electronic Health Records and Genomics

Federated Learning for Sparse Bayesian Models with Applications to Electronic Health Records and Genomics
复制标题

稀疏贝叶斯模型的联合学习及其在电子健康记录和基因组学中的应用

DOI:
10.1142/9789811270611_0044
复制
发表时间:
2023
期刊:
Pacific Symposium on Biocomputing 2023
影响因子:
--
通讯作者:
Ni, Yang
Ni, Yang
中科院分区:
--
文献类型:
--
作者:
Kidd, Brian;Wang, Kunbo;Xu, Yanxun;Ni, Yang

文献摘要

参考文献

被引文献

相似文献

随着包括生物和生物医学领域在内的各个学科对侵犯隐私的担忧上升,联合学习正变得越来越受欢迎。其主要思想是使用仅对该服务器可用的数据在每台服务器上本地训练模型,并在全局级别聚合模型(而不是数据)信息。虽然联合学习已经在深度神经网络等机器学习方法方面取得了重大进展,但就我们所知,它在稀疏贝叶斯模型中的发展仍然不足。稀疏贝叶斯模型具有很高的可解释性,具有自然的不确定性量化,这是许多科学问题的理想属性。然而,如果没有联合学习算法,它们对来自多个来源的敏感生物/生物医学数据的适用性是有限的。因此,为了填补文献中的这一空白,我们提出了一种新的贝叶斯联邦学习框架,该框架能够在不破坏隐私的情况下汇集来自不同数据源的信息。所提出的方法在概念上易于理解和实施,适应了不同数据源的样本异质性(即非IID观测),并允许原则性不确定性量化。我们用三个具体的稀疏贝叶斯模型,即稀疏回归模型、马尔可夫随机场模型和有向图模型来说明所提出的框架。通过三个真实的数据例子展示了这三个模型的应用,包括多医院新冠肺炎研究、乳腺癌蛋白质-蛋白质相互作用网络和基因调控网络。
Federated learning is becoming increasingly more popular as the concern of privacy breaches rises across disciplines including the biological and biomedical fields. The main idea is to train models locally on each server using data that are only available to that server and aggregate the model (not data) information at the global level. While federated learning has made significant advancements for machine learning methods such as deep neural networks, to the best of our knowledge, its development in sparse Bayesian models is still lacking. Sparse Bayesian models are highly interpretable with natural uncertain quantification, a desirable property for many scientific problems. However, without a federated learning algorithm, their applicability to sensitive biological/biomedical data from multiple sources is limited. Therefore, to fill this gap in the literature, we propose a new Bayesian federated learning framework that is capable of pooling information from different data sources without breaching privacy. The proposed method is conceptually simple to understand and implement, accommodates sampling heterogeneity (i.e., non-iid observations) across data sources, and allows for principled uncertainty quantification. We illustrate the proposed framework with three concrete sparse Bayesian models, namely, sparse regression, Markov random field, and directed graphical models. The application of these three models is demonstrated through three real data examples including a multi-hospital COVID-19 study, breast cancer protein-protein interaction networks, and gene regulatory networks.
实现癌症基因组数据的共同愿景。
DOI: 10.1056/nejmp1607591
发表时间: 2016-09-22
期刊: The New England journal of medicine
影响因子: --
作者:
Grossman RL;Heath AP;Ferretti V;Varmus HE;Lowy DR;Kibbe WA;Staudt LM
通讯作者: Staudt LM
DOI: 10.1126/science.abf3066
发表时间: 2021-10
期刊: SCIENCE
影响因子: 56.9
作者:
Kim, Minkyu;Park, Jisoo;Bouhaddou, Mehdi;Kim, Kyumin;Rojc, Ajda;Modak, Maya;Soucheray, Margaret;McGregor, Michael J.;O'Leary, Patrick;Wolf, Denise;Stevenson, Erica;Foo, Tzeh Keong;Mitchell, Dominique;Herrington, Kari A.;Munoz, Denise P.;Tutuncuoglu, Beril;Chen, Kuei-Ho;Zheng, Fan;Kreisberg, Jason F.;Diolaiti, Morgan E.;Gordan, John D.;Coppe, Jean-Philippe;Swaney, Danielle L.;Xia, Bing;van 't Veer, Laura;Ashworth, Alan;Ideker, Trey;Krogan, Nevan J.
通讯作者: Krogan, Nevan J.
贝叶斯非参数聚类的共识变分和蒙特卡罗算法
DOI: --
发表时间: 2020
期刊: 2020 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者:
Yang Ni;D. Jones;Zeya Wang
通讯作者: Zeya Wang
DOI: 10.1109/lsp.2015.2503725
发表时间: 2016-01-01
影响因子: 3.9
作者:
Makalic, Enes;Schmidt, Daniel F.
通讯作者: Schmidt, Daniel F.
DOI: --
发表时间: 1996
期刊: Conference on Uncertainty in Artificial Intelligence
影响因子: --
作者:
T. Richardson
通讯作者: T. Richardson