A secure distributed logistic regression protocol for the detection of rare adverse drug events.

A secure distributed logistic regression protocol for the detection of rare adverse drug events.
复制标题

DOI:
10.1136/amiajnl-2011-000735
复制
发表时间:
2013-05-01
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Kantarcioglu M
Kantarcioglu M
中科院分区:
其他
文献类型:
--
作者:
El Emam K;Samet S;Arbuckle L;Tamblyn R;Earle C;Kantarcioglu M

文献摘要

参考文献

被引文献

相似文献

评估药物进入市场后的比较风险的能力有限。对于罕见不良事件,需要汇集来自多个来源的数据,以具有检测遗传、种族和临床定义亚群中安全性和有效性差异的把握度和足够的人群异质性。然而,将来自不同数据托管人或司法管辖区的数据集结合起来对汇总数据进行分析会产生重大的隐私问题,需要加以解决。解决这些问题的现有协议可能会导致分析准确性降低,并可能导致敏感信息泄露。为逻辑回归开发一个安全的分布式多方计算协议,提供强隐私保证。我们开发了一个安全的分布式逻辑回归协议,使用一个单一的分析中心与多个网站提供数据。理论安全性分析表明,该协议是鲁棒的似是而非的共谋攻击,并不允许各方获得新的信息,他们之间交换的数据。在模拟数据集上评估了该协议的计算性能和准确性。计算性能随着数据集大小的增加而线性扩展。站点的增加导致计算时间的指数增长。不过,就最多五个地点而言,时间仍然很短,不会影响实际应用。模型参数与SAS中分析的合并原始数据的结果相同,证明模型准确度较高。拟议的协议和原型系统将允许以安全的方式开发逻辑回归模型,而无需共享个人健康信息。这可以缓解建立大规模上市后监测计划的关键障碍之一。我们扩展了安全协议,通过广义估计方程来考虑站点内患者之间的相关性,并通过将其扩展到广义线性模型来适应其他链接函数。
There is limited capacity to assess the comparative risks of medications after they enter the market. For rare adverse events, the pooling of data from multiple sources is necessary to have the power and sufficient population heterogeneity to detect differences in safety and effectiveness in genetic, ethnic and clinically defined subpopulations. However, combining datasets from different data custodians or jurisdictions to perform an analysis on the pooled data creates significant privacy concerns that would need to be addressed. Existing protocols for addressing these concerns can result in reduced analysis accuracy and can allow sensitive information to leak. To develop a secure distributed multi-party computation protocol for logistic regression that provides strong privacy guarantees. We developed a secure distributed logistic regression protocol using a single analysis center with multiple sites providing data. A theoretical security analysis demonstrates that the protocol is robust to plausible collusion attacks and does not allow the parties to gain new information from the data that are exchanged among them. The computational performance and accuracy of the protocol were evaluated on simulated datasets. The computational performance scales linearly as the dataset sizes increase. The addition of sites results in an exponential growth in computation time. However, for up to five sites, the time is still short and would not affect practical applications. The model parameters are the same as the results on pooled raw data analyzed in SAS, demonstrating high model accuracy. The proposed protocol and prototype system would allow the development of logistic regression models in a secure manner without requiring the sharing of personal health information. This can alleviate one of the key barriers to the establishment of large-scale post-marketing surveillance programs. We extended the secure protocol to account for correlations among patients within sites through generalized estimating equations, and to accommodate other link functions by extending it to generalized linear models.
DOI: 10.1186/1471-2458-11-454
发表时间: 2011-06-09
期刊: BMC public health
影响因子: 4.5
作者:
El Emam K;Mercer J;Moreau K;Grava-Gubins I;Buckeridge D;Jonker E
通讯作者: Jonker E
DOI: 10.1002/pds.2053
发表时间: 2011-01-01
影响因子: 2.6
作者:
Coloma, Preciosa M.;Schuemie, Martijn J.;Sturkenboom, Miriam
通讯作者: Sturkenboom, Miriam
DOI: 10.1097/mlr.0b013e3181d9919f
发表时间: 2010-06-01
期刊: MEDICAL CARE
影响因子: 3
作者:
Brown, Jeffrey S.;Holmes, John H.;Platt, Richard
通讯作者: Platt, Richard
DOI: 10.1056/nejmp0911494
发表时间: 2010-03-11
期刊: The New England journal of medicine
影响因子: --
作者:
Basch E
通讯作者: Basch E
DOI: 10.1001/jama.282.19.1845
发表时间: 1999-11-17
影响因子: 120.7
作者:
Effler, P;Ching-Lee, M;Jernigan, D
通讯作者: Jernigan, D