An efficient and accurate distributed learning algorithm for modeling multi-site zero-inflated count outcomes.

An efficient and accurate distributed learning algorithm for modeling multi-site zero-inflated count outcomes.
复制标题

DOI:
10.1038/s41598-021-99078-2
复制
发表时间:
2021-10-04
期刊:
影响因子:
4.6
通讯作者:
Chen Y
Chen Y
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Edmondson MJ;Luo C;Duan R;Maltenfort M;Chen Z;Locke K Jr;Shults J;Bian J;Ryan PB;Forrest CB;Chen Y

文献摘要

参考文献

被引文献

相似文献

临床研究网络(CRN)由多个医疗保健系统组成,每个系统都有来自多个护理站点的患者数据,有利于研究罕见的结果并提高结果的普遍性。虽然CRN鼓励在医疗保健系统之间共享汇总数据,但由于隐私法规,CRN内的各个系统通常无法共享患者级数据,从而禁止多站点回归,这需要分析师访问汇集在一起的所有个体患者数据。荟萃分析通常用于对存储在CRN内的多个机构中的数据进行建模,但可能导致有偏估计,特别是在罕见事件背景下。我们提出了一种通信效率高,隐私保护算法,用于在CRN内建模多站点零膨胀计数结果。我们的方法是一种用于执行栅栏回归(ODAH)的一次性分布式算法,对存储在多个站点中的零膨胀计数数据进行建模,而无需在站点之间共享患者水平数据,从而使估计值非常接近在合并患者水平数据分析中获得的估计值。我们通过广泛的模拟和两个使用电子健康记录的真实数据应用来评估我们的方法:检查与儿科可避免住院相关的风险因素,并对与结直肠癌治疗相关的严重不良事件频率进行建模。在模拟中,ODAH在所有探索的设置中产生的偏倚小于0.1%,而荟萃分析估计值显示偏倚高达12.7%,荟萃分析在高零膨胀或低事件发生率的设置中表现最差。在两种应用的分析中,ODAH估计的20个系数中有18个的偏倚小于10%,而荟萃分析估计的偏倚明显更高。相对于现有的分布式数据分析方法,ODAH提供了一种高度准确,计算效率高的方法来建模多站点零膨胀计数数据。
Clinical research networks (CRNs), made up of multiple healthcare systems each with patient data from several care sites, are beneficial for studying rare outcomes and increasing generalizability of results. While CRNs encourage sharing aggregate data across healthcare systems, individual systems within CRNs often cannot share patient-level data due to privacy regulations, prohibiting multi-site regression which requires an analyst to access all individual patient data pooled together. Meta-analysis is commonly used to model data stored at multiple institutions within a CRN but can result in biased estimation, most notably in rare-event contexts. We present a communication-efficient, privacy-preserving algorithm for modeling multi-site zero-inflated count outcomes within a CRN. Our method, a one-shot distributed algorithm for performing hurdle regression (ODAH), models zero-inflated count data stored in multiple sites without sharing patient-level data across sites, resulting in estimates closely approximating those that would be obtained in a pooled patient-level data analysis. We evaluate our method through extensive simulations and two real-world data applications using electronic health records: examining risk factors associated with pediatric avoidable hospitalization and modeling serious adverse event frequency associated with a colorectal cancer therapy. In simulations, ODAH produced bias less than 0.1% across all settings explored while meta-analysis estimates exhibited bias up to 12.7%, with meta-analysis performing worst in settings with high zero-inflation or low event rates. Across both applied analyses, ODAH estimates had less than 10% bias for 18 of 20 coefficients estimated, while meta-analysis estimates exhibited substantially higher bias. Relative to existing methods for distributed data analysis, ODAH offers a highly accurate, computationally efficient method for modeling multi-site zero-inflated count data.
DOI: 10.1097/mlr.0b013e31829b1d10
发表时间: 2013-08
期刊: Medical care
影响因子: 3
作者:
Jiang X;Sarwate AD;Ohno-Machado L
通讯作者: Ohno-Machado L
DOI: 10.1093/jamia/ocx068
发表时间: 2017-11-01
期刊: Journal of the American Medical Informatics Association : JAMIA
影响因子: --
作者:
Kuo TT;Kim HE;Ohno-Machado L
通讯作者: Ohno-Machado L
DOI: 10.1093/jamiaopen/ooz050
发表时间: 2019-12-01
期刊: JAMIA OPEN
影响因子: 2.1
作者:
Bian, Jiang;Loiacono, Alexander;Hogan, William
通讯作者: Hogan, William
DOI: 10.1080/01621459.2018.1429274
发表时间: 2019-04-03
影响因子: 3.7
作者:
Jordan, Michael I.;Lee, Jason D.;Yang, Yun
通讯作者: Yang, Yun
DOI: 10.1177/0962280214527079
发表时间: 2016-12-01
影响因子: 2.3
作者:
Neelon, Brian;Chang, Howard H.;Hastings, Nicole S.
通讯作者: Hastings, Nicole S.