Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithm

Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithm
复制标题

DOI:
10.1093/jamia/ocz199
复制
发表时间:
2020-03-01
影响因子:
6.4
通讯作者:
Chen, Yong
Chen, Yong
中科院分区:
管理学2区
文献类型:
--
作者:
Duan, Rui;Boland, Mary Regina;Chen, Yong

文献摘要

被引文献

相似文献

目的:我们提出了一种一次性的、隐私保护的分布式算法,用于跨多个临床站点执行逻辑回归(ODAL)。材料和方法:ODAL有效地利用来自本地站点(患者级数据可访问的站点)的信息,并结合来自其他站点的似然函数的一阶(ODAL1)和二阶(ODAL2)梯度来构建估计器,而无需跨站点迭代通信或传输患者级数据。我们通过广泛的模拟研究和对宾夕法尼亚大学卫生系统数据集的应用来评估ODAL。通过将其与基于组合的个体参与者数据或汇集数据(即金标准)的估计器进行比较来评估估计的准确性。结果:我们的模拟研究表明,ODAL1与合并估计相比的相对估计偏差< 3%,标准误差比为
Objectives: We propose a one-shot, privacy-preserving distributed algorithm to perform logistic regression (ODAL) across multiple clinical sites.Materials and Methods: ODAL effectively utilizes the information from the local site (where the patient-level data are accessible) and incorporates the first-order (ODAL1) and second-order (ODAL2) gradients of the likelihood function from other sites to construct an estimator without requiring iterative communication across sites or transferring patient-level data. We evaluated ODAL via extensive simulation studies and an application to a dataset from the University of Pennsylvania Health System. The estimation accuracy was evaluated by comparing it with the estimator based on the combined individual participant data or pooled data (ie, gold standard).Results: Our simulation studies revealed that the relative estimation bias of ODAL1 compared with the pooled estimates was < 3%, and the ratio of standard errors was