Federated learning of predictive models from federated Electronic Health Records.

Federated learning of predictive models from federated Electronic Health Records.
复制标题

DOI:
10.1016/j.ijmedinf.2018.01.007
复制
发表时间:
2018-04
影响因子:
4.9
通讯作者:
Shi W
Shi W
中科院分区:
医学2区
文献类型:
--
作者:
Brisimi TS;Chen R;Mela T;Olshevsky A;Paschalidis IC;Shi W

文献摘要

参考文献

被引文献

相似文献

在大数据时代,针对大规模机器学习问题的计算效率和隐私感知解决方案变得至关重要,特别是在医疗保健领域,大量数据存储在不同的位置,并由不同的实体拥有。过去的研究集中在集中式算法上,这些算法假设存在一个中央数据存储库(数据库),该数据库存储并可以处理来自所有参与者的数据。然而,当数据不在中心位置、不能很好地扩展到非常大的数据集并且引入可能危及数据完整性和隐私的单点故障风险时,这样的体系结构可能是不切实际的。鉴于大量数据广泛分布在医院/个人之间,非常需要一种分散的、可计算可扩展的方法。我们的目标是解决一个二进制监督分类问题,以使用分布式算法预测心脏事件的住院人数。我们寻求开发一个通用的分散优化框架,使多个数据持有者能够协作并收敛到一个共同的预测模型,而不需要明确地交换原始数据。重点研究了软边距L1正则化稀疏支持向量机(SSVM)分类器。提出了一种迭代聚类初始对偶分裂(CPDS)算法,用于分布式求解大规模支持向量机问题。这种分布式学习方案与多机构协作或点对点应用程序相关,允许数据持有者进行协作,同时保持每个参与者的数据隐私。我们根据一年前患者电子健康记录中的信息,对预测一年内因心脏病住院的问题进行了CPD测试。CPD比集中式方法收敛更快,但代价是代理之间的一些通信。与另一种分布式算法相比,它的收敛速度更快,通信开销更小。在这两种情况下,通过分类器的接收器工作特征曲线(AUC)下的面积来衡量,它获得了相似的预测精度。我们提取算法发现的预测未来住院的重要特征,从而提供一种解释分类结果并为预防工作提供信息的方法。
In an era of “big data,” computationally efficient and privacy-aware solutions for large-scale machine learning problems become crucial, especially in the healthcare domain, where large amounts of data are stored in different locations and owned by different entities. Past research has been focused on centralized algorithms, which assume the existence of a central data repository (database) which stores and can process the data from all participants. Such an architecture, however, can be impractical when data are not centrally located, it does not scale well to very large datasets, and introduces single-point of failure risks which could compromise the integrity and privacy of the data. Given scores of data widely spread across hospitals/individuals, a decentralized computationally scalable methodology is very much in need. We aim at solving a binary supervised classification problem to predict hospitalizations for cardiac events using a distributed algorithm. We seek to develop a general decentralized optimization framework enabling multiple data holders to collaborate and converge to a common predictive model, without explicitly exchanging raw data. We focus on the soft-margin l1-regularized sparse Support Vector Machine (sSVM) classifier. We develop an iterative cluster Primal Dual Splitting (cPDS) algorithm for solving the large-scale sSVM problem in a decentralized fashion. Such a distributed learning scheme is relevant for multi-institutional collaborations or peer-to-peer applications, allowing the data holders to collaborate, while keeping every participant’s data private. We test cPDS on the problem of predicting hospitalizations due to heart diseases within a calendar year based on information in the patients Electronic Health Records prior to that year. cPDS converges faster than centralized methods at the cost of some communication between agents. It also converges faster and with less communication overhead compared to an alternative distributed algorithm. In both cases, it achieves similar prediction accuracy measured by the Area Under the Receiver Operating Characteristic Curve (AUC) of the classifier. We extract important features discovered by the algorithm that are predictive of future hospitalizations, thus providing a way to interpret the classification results and inform prevention efforts.
DOI: 10.4258/hir.2010.16.4.253
发表时间: 2010-12
影响因子: 2.9
作者:
Son YJ;Kim HG;Kim EH;Choi S;Lee SK
通讯作者: Lee SK
DOI: 10.1109/titb.2008.2004495
发表时间: 2009-01-01
影响因子: --
作者:
Khandoker, Ahsan H.;Palaniswami, Marimuthu;Karmakar, Chandan K.
通讯作者: Karmakar, Chandan K.
DOI: 10.1186/1472-6947-10-16
发表时间: 2010-03-22
影响因子: 3.5
作者:
Yu W;Liu T;Valdez R;Gwinn M;Khoury MJ
通讯作者: Khoury MJ
DOI: 10.1016/j.ijmedinf.2005.05.002
发表时间: 2005-08-01
影响因子: 4.9
作者:
Statnikov, A;Tsamardinos, I;Aliferis, CF
通讯作者: Aliferis, CF
DOI: 10.1142/s0218488502001648
发表时间: 2002-10-01
影响因子: 1.5
作者:
Sweeney, L
通讯作者: Sweeney, L